<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Chit's Programming Blog]]></title><description><![CDATA[Welcome to Chit's programming blog, I'm Chit, a Computer Science student sharing my coding journey with you. I write articles and tutorials about technologies.]]></description><link>https://blog.cpbprojects.me</link><generator>RSS for Node</generator><lastBuildDate>Wed, 09 Sep 2026 03:58:04 GMT</lastBuildDate><atom:link href="https://blog.cpbprojects.me/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Focuster review: My experience as a student in summer [week 3]]]></title><description><![CDATA[Background
Hi, I'm a student who will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the fifth week, and I'll share my exper...]]></description><link>https://blog.cpbprojects.me/focuster-review-my-experience-as-a-student-in-summer-week-3</link><guid isPermaLink="true">https://blog.cpbprojects.me/focuster-review-my-experience-as-a-student-in-summer-week-3</guid><category><![CDATA[Productivity]]></category><category><![CDATA[review]]></category><category><![CDATA[student]]></category><category><![CDATA[todoapp]]></category><category><![CDATA[#trial]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Mon, 28 Aug 2023 11:31:30 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1693222238382/06452c01-14a3-423a-9311-83b9d82f6fcc.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-background">Background</h2>
<p>Hi, I'm a student who will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the fifth week, and I'll share my experience using <a target="_blank" href="https://www.focuster.com/">Focuster</a>, the focus manager for entrepreneurs.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693133310437/9e87fd38-2acb-49e1-bdac-2a9150b17564.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-effectiveness">Effectiveness</h2>
<h3 id="heading-positives">Positives</h3>
<ul>
<li><strong>Simplistic interface</strong>: Focuster has a very simple interface, which doesn't overwhelm you with information. The main screen has three tabs, a to-do list of all the things you do, a list of all things you plan to do today, and a calendar view of the tasks. You first create to-do items, drag it to the middle for it to be scheduled, and it will show up on the calendar on the right, simple as that.</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693133510278/df202c62-fe80-4fc8-bcab-c3cb3f65d4dc.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>Increase Awareness of time spending</strong>: Same as Motion and Reclaim, you have a clear view of the tasks you will do that day, and how your time is allocated.</p>
</li>
<li><p><strong>Visualization of how much you can do in one day</strong>: The bar just below the date shows how many tasks you have scheduled for the day, so you know how much you will be working and don't overwork yourself.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693219485166/6d005499-26f8-4f29-9626-c297f9c737e9.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>Instant reschedule if not done</strong>: If you scheduled a task and didn't do it, it will be rescheduled for you.</p>
</li>
<li><p><strong>Have different lists for tasks</strong>: You can have tasks on different lists, which makes it easy to manage</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693219690814/23b40c4c-5116-4a95-a961-2bbca3d6e192.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-negatives">Negatives</h3>
<ul>
<li><strong>Limited support on routines</strong>: Having routines is useful when I plan to use this app in a personal context, for tasks such as taking a bath, doing laundry or checking emails. In Focuster we can add routine tasks, but you can't have a preferred time for the task. There also isn't an interface where you can see all of your routines.</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693220356371/84d9b355-769f-4862-aa5a-15e0ac853fb2.png" alt class="image--center mx-auto" /></p>
<ul>
<li><strong>Need to manually add tasks to a day</strong>: Except for routines, Focuster doesn't automatically add tasks for your day, so you have to add them yourself.</li>
</ul>
<h2 id="heading-ease-of-use">Ease of use</h2>
<h3 id="heading-positives-1">Positives</h3>
<ul>
<li><p><strong>Simple induction</strong>: Focuster gives very simple onboarding tasks when you sign up, instead of having an entire onboarding course.</p>
</li>
<li><p><strong>Integrate with Google Calendar</strong>: Once you link it up to your Google Calendar, you can start scheduling.</p>
</li>
<li><p><strong>Quick startup/load</strong>: Compared to Motion or Reclaim, Focuster starts up almost immediately, and loads very fast too, so we don't have to wait when we open the app</p>
</li>
<li><p>Easy shorthand for adding tasks: You can type "wash dish 20m every day" when adding the task, and it will become a task called "wash dish", which takes 20 minutes, and will be scheduled every day. Very handly.</p>
</li>
</ul>
<h3 id="heading-negatives-1">Negatives</h3>
<ul>
<li><strong>No mobile app</strong>: They do not have a mobile app, so it is a bit harder to add tasks on mobile.</li>
</ul>
<h2 id="heading-customizability">Customizability</h2>
<h3 id="heading-positive">Positive</h3>
<ul>
<li><strong>Colours for tasks in different task lists</strong>: So we can easily distinguish tasks for different projects.</li>
</ul>
<h3 id="heading-negative">Negative</h3>
<ul>
<li><strong>No custom time schedule</strong>: We only have a single schedule called work, We cannot have a separate schedule for school or personal tasks, the schedule also starts and ends at the same time every day, and we cannot have a different schedule on weekends.</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693220815590/6d993aa1-c7ec-479b-8f2a-499d3b3035be.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-cost-value-proposition"><strong>Cost-Value Proposition</strong></h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1693218631398/046a8c8b-435c-4559-b780-c2130faba223.png" alt class="image--center mx-auto" /></p>
<p>The price for the basic plan is basically $8/£6.36 per month for only the most basic features, and for the other ones, they cost $15/£11.92 per month. This is for the annual pricing, for monthly payment, the price goes up to $10 per month for basic, and $20 per month for pro.</p>
<p>This app provides a solid set of functionalities for a scheduling app, and the price is definitely not unaffordable. This price for the basic plan is similar to the starter plan for Reclaim, but Reclaim seems to come up on top with more features.</p>
<h2 id="heading-experience">Experience</h2>
<p>Using this app feels very different from Motion or Reclaim, the one feels much simpler, adding to-do items, then dragging them to the calendar, and marking them as finished. While this one doesn't have the more advanced options such as auto-scheduling, it is still a very solid way of managing your time.</p>
<h2 id="heading-next-week">Next week</h2>
<p>Next week I will be trailing a vastly different productivity tool, <a target="_blank" href="https://habitica.com/static/home">Habitica</a> - Gamify your life. Instead of scheduling things for you, this productivity tool makes tasks into quests you do, like in a game, to motivate you to do your tasks. I'm excited how this will turn out.</p>
<p><a target="_blank" href="https://wssdb.cpbprojects.me/">An app I'm working on</a></p>
]]></content:encoded></item><item><title><![CDATA[Reclaim.AI review: My experience as a student in internship [week 2]]]></title><description><![CDATA[Background
Hi, I'm a student currently doing a summer internship and will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the...]]></description><link>https://blog.cpbprojects.me/reclaimai-review-my-experience-as-a-student-in-internship-week-2</link><guid isPermaLink="true">https://blog.cpbprojects.me/reclaimai-review-my-experience-as-a-student-in-internship-week-2</guid><category><![CDATA[review]]></category><category><![CDATA[Productivity]]></category><category><![CDATA[student]]></category><category><![CDATA[task management]]></category><category><![CDATA[#ai-tools]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Fri, 18 Aug 2023 19:07:10 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1692385558997/5cd2253e-b50d-4d47-9aad-72807848b1a9.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-background"><strong>Background</strong></h2>
<p>Hi, I'm a student currently doing a summer internship and will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the third week, and I'll share my experience using <a target="_blank" href="https://reclaim.ai/r/s/fXHXC">Reclaim</a>, a Smart AI scheduler for busy teams.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692128956502/951f82e9-f970-44a7-8fd6-b7f95b375398.png" alt class="image--center mx-auto" /></p>
<p>I reviewed Motion in <a target="_blank" href="https://blog.cpbprojects.me/motion-review-my-experience-as-a-student-in-internship-week-1">my last blog post</a>, so this blog post will use <a target="_blank" href="https://www.usemotion.com/">Motion</a> as a reference.</p>
<h2 id="heading-effectiveness">Effectiveness</h2>
<h3 id="heading-positives">Positives</h3>
<ul>
<li><p><strong>Increase awareness of time spent</strong>: Similar to Motion, Reclaim lays out all tasks and timetables, providing a clear view of my daily agenda.</p>
</li>
<li><p><strong>Auto-scheduling with AI</strong>: When we have something to do, we click on add tasks, and it will be automatically scheduled on our timetable.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692221494649/cc1b6678-5a20-48df-b9f1-be85ee1dffd3.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>Optional auto rescheduling</strong>: If the task is not marked as done on that day, the tasks will be scheduled again the next day so the task will not be forgotten. The difference between Reclaim and Motion is that in Motion, you have to mark the task as done as soon as it is finished, for Reclaim, you can mark them as done on the same day. (This is an optional feature, the default is to mark everything as done after it happens)</p>
</li>
<li><p><strong>Statistics</strong>: You can see how you are going on the statistics page, which helps you see any improvement or setbacks</p>
</li>
</ul>
<h3 id="heading-negatives">Negatives</h3>
<ul>
<li><strong>Tasks cannot be organized into projects</strong>: Some tasks can be combined into bigger tasks, and being into the same project. However, Reclaim does not support organizing tasks into projects.</li>
</ul>
<h2 id="heading-ease-of-use"><strong>Ease of use</strong></h2>
<h3 id="heading-positives-1">Positives</h3>
<ul>
<li><strong>Good onboarding</strong>: There is a comprehensive setup guide, it introduces all the features of the app, so you can learn how to fully use this tool. One bad thing is that the startup screens take very long to load, more than 30 seconds for me.</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692129985178/d87711f2-7e38-4807-8f96-1c81880f9c00.png" alt class="image--center mx-auto" /></p>
<ul>
<li><strong>Integrated with Google Calendar</strong>: Tasks scheduled will automatically show up on google calendar, they also have a google calendar plugin to quickly add tasks</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692130284192/b2dbc185-e6fa-4643-a07d-636727b39ea1.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-negatives-1">Negatives</h3>
<ul>
<li><p><strong>Only available on the website</strong>: there are no mobile or desktop apps, therefore you must use the app from the website.</p>
</li>
<li><p><strong>Only supporting Google Calendar</strong>: Microsoft or other types of calendars are not supported, luckily I use google calendar anyways so it doesn't affect me</p>
</li>
<li><p><strong>Slow loading</strong>: when the website starts up, it tends to take quite some time to load, which is a mild inconvenience</p>
</li>
</ul>
<h2 id="heading-customisability">Customisability</h2>
<h3 id="heading-positives-2">Positives</h3>
<ul>
<li><p>Allow multiple schedules: we have a set of time set for work, school or personal life, so we can add tasks to their specific slot</p>
</li>
<li><p><strong>Comprehensive settings page</strong>: Compared to Motion, so much more things can be set, the most helpful for me is default task hours, so when I add a new task, it will be quicker as I have my usual settings in default</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692221681884/36e9cca3-b952-4449-bc89-c2aea1c43a26.png" alt class="image--center mx-auto" /></p>
<ul>
<li><strong>Allow preferred time for habits</strong>: you can set a preferred time to complete a habit, for example, you can prefer writing your diary at 11 pm, which makes much more sense than writing a diary in the morning</li>
</ul>
<h3 id="heading-negatives-2">Negatives</h3>
<p>In terms of customizability, I don't really have anything to complain about.</p>
<h2 id="heading-cost-value-proposition"><strong>Cost-Value Proposition</strong></h2>
<p>The price for <a target="_blank" href="https://reclaim.ai/r/s/fXHXC">Reclaim</a> is £10 per month or £96 per year. This price is almost half of Motion. It is a fair deal if it really helps me increase my productivity. However, since I will only be using the basic functionalities of organizing tasks, but not the more business-centric ones such as meeting scheduling, it may still be a bit expensive.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692222018962/83498978-4590-419a-a4bc-781a675b6cbf.png" alt class="image--center mx-auto" /></p>
<p>The good thing is that they have a free tier, which only allows up to 2 habits, but is still fine for basic stuff. They also have an education discount which allows for half price, so that's worth checking out. They also have a discount for if you switch from other providers, which is kind of savage if you ask me.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1692222039029/e1eb6e22-bc78-4fc0-9191-5c2e19704b6b.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-experience">Experience</h2>
<p>Reclaim feels like a more polished Motion, with a cleaner interface, and some more helpful functionality, aside from that, it pretty much accomplishes the same thing.</p>
<h2 id="heading-next-week">Next week</h2>
<p>Next week I will be trialling another productivity tool: I was going to go with <a target="_blank" href="https://www.getclockwise.com/">GetClockwise.com</a> since it is listed as one of the competitors of Reclaim, but I found out it is only for business Google accounts. I will go for <a target="_blank" href="https://next.focuster.com/">next.focuster.com</a> instead, and see how it helps me.</p>
<p><a target="_blank" href="https://wssdb.cpbprojects.me/">An app I'm working on</a></p>
]]></content:encoded></item><item><title><![CDATA[Motion Review: My experience as a student in internship [week 1]]]></title><description><![CDATA[Background
Hi, I'm a student currently doing a summer internship and will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the...]]></description><link>https://blog.cpbprojects.me/motion-review-my-experience-as-a-student-in-internship-week-1</link><guid isPermaLink="true">https://blog.cpbprojects.me/motion-review-my-experience-as-a-student-in-internship-week-1</guid><category><![CDATA[review]]></category><category><![CDATA[Productivity]]></category><category><![CDATA[Time management]]></category><category><![CDATA[#ai-tools]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Mon, 31 Jul 2023 21:51:57 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1690840123756/5bdf844a-d728-430c-8d9b-b88b7b77c726.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-background">Background</h2>
<p>Hi, I'm a student currently doing a summer internship and will soon enter my final year. I'm reviewing different productivity apps to see which one can help me waste less time and have more time to work on my personal projects. This is the first week, and I'll share my experience using <a target="_blank" href="https://www.usemotion.com/">Motion</a>, an AI-based productivity tool for managing tasks and schedules.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690839030903/9f5bfa82-cce9-482d-8115-e1d69c3030f5.png" alt="motion landing page" class="image--center mx-auto" /></p>
<h2 id="heading-effectiveness">Effectiveness</h2>
<h3 id="heading-positives">Positives</h3>
<ul>
<li><p><strong>Increase Awareness of time spending</strong>: Having a visible timetable and task list in front of me, and the obligation to mark tasks as complete after executing them, offered a real-time view of my accomplishments and pending tasks. This kept me aware and accountable for my day.</p>
</li>
<li><p><strong>Seeing how little time I actually have</strong>: Once I saw the recurring tasks scheduled on my calendar, I realized how scarce my 'free' time was, making me appreciate and utilize it more consciously.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690839236159/9f373a0f-0a4b-436f-b4a5-315a462fdfd3.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>AI auto-scheduling</strong>: Letting an AI manage my schedule freed my mind from juggling 'when to do what'. This helped me concentrate more on 'what' needs to be done, enhancing my focus and effectiveness.</p>
</li>
<li><p><strong>Automated rescheduling</strong>: If I didn't mark a task as complete, Motion would reschedule it automatically. This feature was particularly helpful when unplanned events (like a lengthy meal with friends) derailed my schedule, saving me from the tedious task of manually rearranging my calendar.</p>
</li>
<li><p><strong>Project-based planning</strong>: The tool allows the creation of projects and their associated tasks. This made scheduling and tracking tasks related to different personal projects more straightforward.</p>
</li>
<li><p><strong>Scheduled Breaks</strong>: Motion places breaks after every 'x' minute. This helped instil discipline in my break patterns, preventing me from extending breaks excessively.</p>
</li>
</ul>
<h2 id="heading-negatives">Negatives</h2>
<ul>
<li><p><strong>Not so intelligent intelligent-scheduling</strong>: While Motion claims to offer intelligent scheduling, my experience suggested a 'first come, first serve' approach. The tool would sometimes schedule routine tasks, like writing a diary or bathing, as the first activities of the day, which doesn't make sense. I had to manually rearrange these tasks.</p>
</li>
<li><p><strong>Inefficient time segmentation</strong>: Motion mandates tasks be scheduled in multiples of 15 minutes, which doesn't always align with real-world scenarios. Some tasks may take less time, causing potential inefficiencies in the schedule.</p>
</li>
</ul>
<h2 id="heading-ease-of-use">Ease of use</h2>
<h3 id="heading-positives-1">Positives</h3>
<ul>
<li><p><strong>Comprehensive Onboarding</strong>: Motion provides a clear and detailed introduction upon first use. It helped me understand how to set up basic functionalities, making the initial interaction smooth and enjoyable.</p>
</li>
<li><p><strong>Robust Help Page</strong>: The help page on <a target="_blank" href="help.usemotion.com">help.usemotion.com</a> is well-organized and informative. It offers guidance on how to use various features, which is a great resource when you're new to the tool.</p>
</li>
<li><p><strong>Available on multiple platforms</strong>: With a Chrome plugin and a mobile app, Motion allows users to schedule tasks on different devices according to their convenience.</p>
</li>
<li><p><strong>Integrated with calendar</strong>: The ability to click and drag on the calendar to create fixed-time events is a significant feature. These events also sync with Google Calendar automatically, reducing the redundancy of managing multiple calendars.</p>
</li>
</ul>
<h3 id="heading-negatives-1">Negatives</h3>
<ul>
<li><p><strong>Limited mobile interactions</strong>: The inability to drag and create fixed-time tasks on the mobile app can be somewhat frustrating, especially when on the go.</p>
</li>
<li><p><strong>Very slow startup</strong>: The loading time for the app, whether on mobile, web, or the Windows app, is notably slow, typically around 6 seconds.</p>
</li>
<li><p><strong>Have to manually set tasks as completed</strong>: You have to manually mark a task as complete within 30 minutes of finishing it. Although this feature ensures that unfinished tasks are rescheduled, it can be a bit annoying if you're quickly moving between tasks.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690839431920/0731f44a-07f6-45ec-bf86-77e9692027a3.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>Slow scheduling/rescheduling</strong>: Shifting tasks around can be sluggish, taking around 10 to 20 seconds for the system to reschedule. Even ChatGPT replies faster than this.</p>
</li>
<li><p><strong>Inflexible schedule order</strong>: Motion offers different "schedules", which are timeframes where you do stuff. Unfortunately, it doesn't allow editing the order of them, and the work schedule always comes first, I find myself having to switch to 'life' every time I create a new task.</p>
</li>
<li><p><strong>Bugs</strong>: I encountered a bug where completed tasks wouldn't display even when I set the preference to 'show'. Though the issue was fixed later, it was a minor annoyance.</p>
</li>
</ul>
<h2 id="heading-customisability">Customisability</h2>
<h3 id="heading-positives-2">Positives</h3>
<ul>
<li><p><strong>Definable schedules</strong>: Motion allows users to define their schedules, such as 'work' or 'life'. For example, since I use Motion for activities after work, I have set up a "life" schedule for post-work hours on weekdays and the entire duration of weekends.</p>
</li>
<li><p><strong>Custom task timing</strong>: Tasks in Motion can have specific time frames. For instance, you might prefer to take a bath only between 8 pm and 11 pm, and Motion allows for such specific scheduling. It also lets you add preferred times for tasks, making the scheduling more personalized.</p>
</li>
</ul>
<h3 id="heading-negatives-2">Negatives</h3>
<ul>
<li>No colour customization: Tasks scheduled automatically by Motion cannot be colour-coded. They are all displayed in grey. The lack of colour-coding for tasks can make the calendar look monotonous and limit visual differentiation between different tasks or task categories.</li>
</ul>
<h2 id="heading-cost-value-proposition">Cost-Value Proposition</h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690839094108/1ae204b3-1675-49c3-afa0-39b7424f7dd7.png" alt class="image--center mx-auto" /></p>
<p>The pricing for Motion is $228/£178 per year or $34/£27 per month, it is definitely not cheap. Breaking down the annual cost, Motion comes to approximately 50p GBP per day. The question then becomes: does Motion provide value that exceeds or at least matches this daily cost?</p>
<p>If used diligently, Motion could potentially help you achieve your personal project goals. There's also a chance that these personal projects might generate income, indirectly making up for the cost of the tool. In this way, Motion can be seen as an investment in your productivity and potential success.</p>
<p>However, the other question is, can I achieve the same goal with another tool? there are so many productivity tools out there, and most of them are cheaper while offering functionality that isn't that far off, especially when I don't use the work-based functionality, such as meeting scheduling, kanban boards, or meeting time matches, then the cost might seem steep for the value I'm receiving. Other tools could potentially serve the same purpose as Motion for a fraction of the cost.</p>
<h2 id="heading-experience">Experience</h2>
<p>Using Motion has been an enlightening experience. It introduced me to many new concepts in time management. I learned about the differences between free, busy, and scheduled tasks, as well as the concept of task blockers. This newfound knowledge enabled me to better categorize my tasks and identify potential obstacles in my schedule.</p>
<p>The most significant impact of Motion has been on my awareness and utilization of time. With its intuitive task and time management features, Motion has encouraged me to stay on track with my tasks and commitments. As a result, I have noticed a significant reduction in time wastage.</p>
<p>Moreover, even if I do deviate from my schedule due to unforeseen circumstances or momentary lapses, Motion automatically reschedules my tasks. This auto-rescheduling feature prevents me from getting overwhelmed by piling tasks and helps me get back on track without any hassle.</p>
<h2 id="heading-next-week">Next week</h2>
<p>As set out in this journey, next week I will be trialling another productivity tool: <a target="_blank" href="reclaim.ai">Reclaim.ai</a>, another AI calendar that automatically plans the day for me, and see how well it does to help me rescue my time.</p>
<p><a target="_blank" href="https://wssdb.cpbprojects.me/">An app I'm working on</a></p>
]]></content:encoded></item><item><title><![CDATA[Trialling Productivity Tools to Rescue My Time [Week 0]]]></title><description><![CDATA[Introduction
I've been seeing a lot of ads on UseMotion lately, it is an AI-powered auto-scheduling tool promising to make your time so efficient it feels like squeezing an extra month into your year. This brave claim reminded me of the hours I’ve al...]]></description><link>https://blog.cpbprojects.me/trialling-productivity-tools-to-rescue-my-time-week-0</link><guid isPermaLink="true">https://blog.cpbprojects.me/trialling-productivity-tools-to-rescue-my-time-week-0</guid><category><![CDATA[Productivity]]></category><category><![CDATA[Time management]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Tue, 25 Jul 2023 21:45:07 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1690321356599/2f404dbe-3c2a-421b-903c-04d72c8febed.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-introduction">Introduction</h2>
<p>I've been seeing a lot of ads on <a target="_blank" href="https://www.usemotion.com/">UseMotion</a> lately, it is an AI-powered auto-scheduling tool promising to make your time so efficient it feels like squeezing an extra month into your year. This brave claim reminded me of the hours I’ve allowed to slip by reproductively, particularly after a long work day when my energy dwindles and motivation wanes. Personal projects, once a source of joy and creativity, now lie dormant (as you can see by me not writing a blog post in months). This realization led me to think - is there a better way to organize my time? Could modern technology lend me a helping hand? To find out, I decided to embark on a journey, trialling different productivity tools, with me as the test subject, to see if they could salvage my most precious asset: time.</p>
<h2 id="heading-a-little-bit-about-myself">A little bit about myself</h2>
<p>To help you understand how I will evaluate the productivity tools, it's crucial to understand who I am and my daily routines. The effectiveness of any tool greatly depends on the person using it and the specific needs it addresses. So let me introduce myself, I'm a soon-to-be final-year student, currently immersed in a summer internship that ends at 6 pm. Once home, a list of everyday chores awaits me, often stretching into my evening. I also have personal projects that I'm passionate about and yearn to devote more time to. Therefore I will use the productivity tools for my personal life instead of work, I want to be able to do my chores, tasks, and personal projects effectively.</p>
<p>Given that I plan to trial these tools sequentially, spending approximately one week with each, my circumstances will change over the trial period. My internship will end, and my final year of university will commence, altering my daily routine and the time I have at my disposal. Nonetheless, I'll try to be objective when I test the different productivity tools.</p>
<h2 id="heading-criteria-that-i-will-use">Criteria that I will use</h2>
<ol>
<li><p>Effectiveness: How much does it help me to complete my tasks quicker</p>
</li>
<li><p>Ease of Use: How easy is it to use the tool</p>
</li>
<li><p>Customizability: Since most productivity tools are for work, I want them to be customizable so that I can use them to manage my personal life</p>
</li>
<li><p>Cost-Value Proposition: I don't have money as I'm a student, so if I were to subscribe to a tool, it better be worth it</p>
</li>
</ol>
<h2 id="heading-our-first-contestant-usemotion">Our first contestant: UseMotion</h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690319277398/f8b07dc1-3c8e-4920-9cf1-cf027999e218.png" alt class="image--center mx-auto" /></p>
<p>Motion is an AI-powered productivity app that helps manage your tasks on your calendar, the advertised point is how it automatically builds your schedule using AI, so if you have some change in your plan, the tasks that don't have to be done in that specific time will be automatically shifted. I'm not going into details on how Motion works as this isn't the point of this blog post, I recommend checking out <a target="_blank" href="https://readcaffeine.com/2022/03/motion-review/">Caffeine's article on it</a> if you want to find out more.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1690320225125/d7f3ff10-e35b-43f3-84fc-97184e363fdc.png" alt class="image--center mx-auto" /></p>
<p>The interface looks something like this, I've added my recurring tasks, as well as my one-time tasks here, if the task doesn't have a fixed time, it will be automatically scheduled on the calendar, that's why the calendar is filled.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>I'll be using Motion during its free trial of a week, and discovering its features and do they benefit me. Stay tuned for the next week, where I will conclude my thoughts on Motion, and trial another productivity tool.</p>
]]></content:encoded></item><item><title><![CDATA[Use cases of my prayer tracker website]]></title><description><![CDATA[In the previous blog, I talked about my plan to build this prayer tracker website. This time, I will talk about the use cases of this website, to better plan the website.

Side note: I will call this a web application from now on. Websites tend to be...]]></description><link>https://blog.cpbprojects.me/use-cases-of-my-prayer-tracker-website</link><guid isPermaLink="true">https://blog.cpbprojects.me/use-cases-of-my-prayer-tracker-website</guid><category><![CDATA[use-cases]]></category><category><![CDATA[Diagram]]></category><category><![CDATA[Web Development]]></category><category><![CDATA[webapps]]></category><category><![CDATA[web app development]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Fri, 10 Feb 2023 12:42:31 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1676031910441/35a8bf81-2c56-4476-bd55-02c05442511d.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In <a target="_blank" href="https://blog.cpbprojects.me/my-plan-for-building-a-prayer-tracker-website">the previous blog</a>, I talked about my plan to build this prayer tracker website. This time, I will talk about the use cases of this website, to better plan the website.</p>
<ul>
<li>Side note: I will call this a web application from now on. Websites tend to be more static and focused on providing information, while web applications are more interactive and focused on providing a specific service or functionality to users. Therefore my idea classifies as a web app.</li>
</ul>
<h2 id="heading-use-case-diagram">Use case diagram</h2>
<p>As you can see, I am to provide a lot of functionality to the user. They can be classified into three major parts</p>
<ol>
<li><p>Interacting with prayer items: Create, edit, and delete them. These are the main functionality of the web app</p>
</li>
<li><p>Authentication: Sign in/up/out, and edit account details.</p>
</li>
<li><p>Payment: Provide payment method via a Payment provider.</p>
</li>
</ol>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672318010624/ffdff354-1a7e-47db-853a-abff522abfa8.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-interacting-with-prayer-items">Interacting with prayer items</h3>
<p>The idea of the app is to remind the user to pray about things, therefore items that the user should pray for will be displayed. There is also a pray button, when clicked, it indicates that the user has prayed for this item today, and the item will only come back to the list after the selected frequency.</p>
<p>There is an options button, when clicked, opens a dropdown menu, allowing the user to edit, mark as prayed, mark as denied, and delete the item. If the user clicks the item content, the item will expand showing more details.</p>
<p>There should also be an add new button, which will create a prompt for users to fill in details of the new prayer item the user wishes to create.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672316717824/9f320f06-9493-494e-8781-1d363dbffa98.png" alt class="image--center mx-auto" /></p>
<p>There should also be two separate areas for answered and denied prayers, with the options of editing and deleting.</p>
<h3 id="heading-authentication">Authentication</h3>
<p>The web app should support signing up with email and asking for the username in the process. It should also support signing up with OAuth using Google or Facebook.</p>
<p>For sign-in, we should allow users to sign in with email or OAuth. If the user forgot their password, we should allow the user to reset their password through email authentication.</p>
<p>For logged-in users, they should be able to change passwords, change email, log out, and even delete their accounts because of GDPR regulations.</p>
<h3 id="heading-payment">Payment</h3>
<p>I want to implement a subscription system so that users that want more functionalities can do so by subscribing to more advanced plans. To do so, I need a page for users to view their current subscriptions, purchase a subscription, and manage their current subscriptions.</p>
<h2 id="heading-detailed-use-case">Detailed use case</h2>
<p>Below I am going to document the key use cases</p>
<h3 id="heading-use-case-1-display-the-list-of-active-prayer-items">Use case 1: Display the list of active prayer items</h3>
<p>Actors: User, System</p>
<p>Preconditions:</p>
<ul>
<li><p>The web app has established a connection to the server</p>
</li>
<li><p>The user is authenticated</p>
</li>
</ul>
<p>The flow of events:</p>
<ol>
<li><p>If the user has encryption enabled, try to read the encryption key from local storage, if it is not there, prompt the user to enter it</p>
</li>
<li><p>the web app makes a request to the server for the list of prayers items the user has</p>
</li>
<li><p>the server queries the database and sends back the list of prayer items</p>
</li>
<li><p>the web app sorts the prayer items by the next_date, then display the ones with a next_date before or equal to today with a for loop</p>
<ol>
<li><p>If the item is encrypted, decrypt it with the secret key, and display it with a lock icon</p>
</li>
<li><p>If the item is not encrypted, display it</p>
</li>
</ol>
</li>
<li><p>for the prayer items with a next_date in the future, they will be displayed below with a grey overlay, to symbolize that they are queued for the future.</p>
</li>
</ol>
<p>Postconditions: The list of active prayer items is displayed on the UI</p>
<h3 id="heading-use-case-2-add-a-prayer-item">Use case 2: Add a prayer item</h3>
<p>Actors: User, system</p>
<p>Preconditions:</p>
<ul>
<li><p>The web app has established a connection to the server</p>
</li>
<li><p>The user is authenticated</p>
</li>
</ul>
<p>The flow of event:</p>
<ol>
<li><p>the user input the prayer item details</p>
</li>
<li><p>the user press add item or press enter on the keyboard</p>
</li>
<li><p>the web app performs client-side input validation. Including the input length</p>
</li>
<li><p>the web app sends a request to the server to add the item</p>
<ol>
<li>If the user enabled encryption, the prayer content will be encrypted before being sent to the server</li>
</ol>
</li>
<li><p>the server performs server-side validation, including prayer item count</p>
</li>
<li><p>the server adds the prayer item to the database</p>
</li>
<li><p>the server increments the total number of prayer items for that user from the database</p>
</li>
<li><p>the server sends a response to the web app</p>
</li>
<li><p>the web app displays the additional prayer item</p>
</li>
</ol>
<p>Postconditions: the newly added prayer item is in the web app as well as the server database</p>
<p>Alternative flow:</p>
<ul>
<li>In steps 5-6, if there is any error causing the insert to be unsuccessful, the server will send back a response about why the insert failed, and the web app will display the reason to the user</li>
</ul>
<h3 id="heading-use-case-3-new-user-signup">Use case 3: New user signup</h3>
<p>Actors: User, system</p>
<p>Preconditions: The web app has established a connection to the server</p>
<p>The flow of event:</p>
<ol>
<li><p>the user completes the sign-up form</p>
</li>
<li><p>the web app sends the sign-up request to the server</p>
</li>
<li><p>the server creates the user</p>
</li>
<li><p>the after-insert trigger function triggers, inserting the user in the other tables</p>
</li>
<li><p>if the user uses email auth, they have to go to their email account and verify the email</p>
</li>
<li><p>the user logs in</p>
</li>
</ol>
<p>Postconditions: the user lands on the dashboard, and sees a basic tutorial</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>After doing this analysis of use cases, I now have a better understanding of the specific needs and goals of users, which helps to ensure that the web app is designed and developed to meet their expectations.</p>
]]></content:encoded></item><item><title><![CDATA[Set up a Svelte todo list on self-hosted Supabase + Email sign up + Google, Facebook Auth + host on GitHub pages]]></title><description><![CDATA[In this tutorial. I will guide you through creating a to-do list web app using Svelte as the front end hosted on GitHub pages, and Supabase as the back end. I will also discuss how to set up an email authentication, as well as google and Facebook thi...]]></description><link>https://blog.cpbprojects.me/set-up-a-svelte-todo-list-on-self-hosted-supabase-email-sign-up-google-facebook-auth-host-on-github-pages</link><guid isPermaLink="true">https://blog.cpbprojects.me/set-up-a-svelte-todo-list-on-self-hosted-supabase-email-sign-up-google-facebook-auth-host-on-github-pages</guid><category><![CDATA[supabase]]></category><category><![CDATA[Svelte]]></category><category><![CDATA[oauth]]></category><category><![CDATA[email]]></category><category><![CDATA[Tutorial]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Fri, 06 Jan 2023 13:21:09 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1673011024664/333788d2-8e29-4612-958f-e70ca932f389.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this tutorial. I will guide you through creating a to-do list web app using Svelte as the front end hosted on GitHub pages, and Supabase as the back end. I will also discuss how to set up an email authentication, as well as google and Facebook third-party OAuth. The outcome would be like <a target="_blank" href="https://sveltesupabasetodoexample.cpbprojects.me/">this</a> from <a target="_blank" href="https://github.com/chit-uob/svelte_supabase_todo_example">this Github repo</a>.</p>
<h2 id="heading-important-information">Important information</h2>
<p>This is my second time making a tutorial, if there is anything unclear, feel free to ask in the comments and I'll try to reply to them.</p>
<p>This tutorial assumes you are using a self-hosted Supabase. If you are using a managed Supabase you should follow the instructions on <a target="_blank" href="https://github.com/supabase/supabase/tree/master/examples/todo-list/sveltejs-todo-list">the GitHub page</a> instead.</p>
<p>If you don't already have a self-hosted Supabase, check out my other tutorial <a target="_blank" href="https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide">How to Self-host Supabase</a> and come back to this tutorial afterwards.</p>
<h2 id="heading-create-a-github-repository-for-the-svelte-app">Create a GitHub repository for the Svelte App</h2>
<p>We will host our svelte app on GitHub pages. It is a static site hosting service that host files straight out of a repository on GitHub. It has the benefit of being free and has pretty good bandwidth since it is hosted by a big company.</p>
<p>We first need to create a GitHub repository. if you <strong>don't</strong> have GitHub Pro, you must create a public repository in order to share your site with GitHub pages, but with GitHub Pro, both private and public repositories support GitHub pages.</p>
<p>For the repository name, you can choose any name you want unless you want to page to be on your main GitHub page (Main GitHub page means the website on your {github_username}.github.io). If so, you need to name your repository as <code>{github_username}.github.io</code>.</p>
<p>If you name your repository other things, your website would have a link of <code>{github_username}.github.io/{project_name}</code> . In my case, my website would be on <a target="_blank" href="https://chit-uob.github.io/svelte_supabase_todo_example">https://chit-uob.github.io/svelte_supabase_todo_example</a>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672932802794/1a862415-a227-4e0f-acdb-be28cdd52ea6.png" alt class="image--center mx-auto" /></p>
<p>You can then click <strong>Create Repository</strong>.</p>
<h2 id="heading-clone-to-do-list-example">Clone To-do list example</h2>
<p>Then we clone the Svelte to-do list example from the Supabase GitHub. We choose a folder to put the project, and then we clone it with the sparse setting, this clones the repository without downloading every single file. Then we can go inside the repository and spare-checkout the to-do list example.</p>
<ul>
<li><p>Make sure you have Git installed, install from <a target="_blank" href="https://git-scm.com/">here</a> if you don't already</p>
</li>
<li><p>To open the command prompt on windows, you can right-click an empty space on the folder, and choose <strong>Open in Windows Terminal</strong></p>
</li>
</ul>
<pre><code class="lang-bash">git <span class="hljs-built_in">clone</span> --depth 1 --filter=blob:none --sparse https://github.com/supabase/supabase
<span class="hljs-built_in">cd</span> supabase
git sparse-checkout <span class="hljs-built_in">set</span> examples/todo-list/sveltejs-todo-list
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672933152316/1329594a-c258-4bc1-82ef-b88009300041.png" alt class="image--center mx-auto" /></p>
<p>Now we can go in <code>supabase -&gt; examples -&gt; todo-list</code> , and copy the <code>sveltejs-todo-list</code> folder to another folder you use for the project. I am going to use the same example folder. Then we open the terminal inside the <code>sveltejs-todo-list</code> folder, and run the following commands.</p>
<pre><code class="lang-bash">git init
git add .
git commit -m <span class="hljs-string">"first commit"</span>
git branch -M main
git remote add origin https://github.com/{github_username}/{repo_name}.git
git push -u origin main
</code></pre>
<ul>
<li>If you haven't already logged in to GitHub, it may prompt you to log in with your credentials</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672933722199/90c9d40f-1268-41a2-9d35-68fa9d7611e3.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-configure-front-end">Configure front end</h2>
<p>Then we install the GitHub pages npm package</p>
<ul>
<li>You need Node.js installed for this, install it from <a target="_blank" href="https://nodejs.org/en/download/">here</a> if not already</li>
</ul>
<pre><code class="lang-bash">npm install gh-pages --save-dev
</code></pre>
<h3 id="heading-set-up-packagejson">Set up package.json</h3>
<p>Inside the file <code>package.json</code>, we add a "homepage" property and set it as <code>https://{github_username}.github.io/{project_name}</code>, and a "deploy" inside the "script" property and set it as <code>gh-pages -d dist</code>, so lines 4-10 should look something like this:</p>
<pre><code class="lang-json">  <span class="hljs-string">"version"</span>: <span class="hljs-string">"2.0.0"</span>,
  <span class="hljs-string">"type"</span>: <span class="hljs-string">"module"</span>,
  <span class="hljs-string">"homepage"</span>: <span class="hljs-string">"https://chit-uob.github.io/svelte_supabase_todo_example"</span>,
  <span class="hljs-string">"scripts"</span>: {
    <span class="hljs-attr">"deploy"</span>: <span class="hljs-string">"gh-pages -d dist"</span>,
    <span class="hljs-attr">"dev"</span>: <span class="hljs-string">"concurrently \"npm run dev:css\" \"vite\""</span>,
    <span class="hljs-attr">"dev:css"</span>: <span class="hljs-string">"tailwindcss -w -i ./src/tailwind.css -o src/assets/app.css"</span>,
</code></pre>
<p>Then we copy the <code>.env.example</code> file, and rename it to <code>.env</code> . Inside there we fill in the backend URL and anon key. If you followed <a target="_blank" href="https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide">my last tutorial</a>, <code>VITE_SUPABASE_URL</code> would be your backend domain name. Otherwise, it would be whatever you set your <code>API_EXTERNAL_URL</code> inside the <code>.env</code> file inside the Supabase docker folder. As for <code>VITE_SUPABASE_ANON_KEY</code> , it is the value of <code>ANON_KEY</code> in the Supabase docker folder <code>.env</code> file.</p>
<pre><code class="lang-apache">...
<span class="hljs-attribute">JWT_SECRET</span>=...
<span class="hljs-attribute">ANON_KEY</span>={THIS IS THE ANON KEY}
<span class="hljs-attribute">SERVICE_ROLE_KEY</span>=...
...
<span class="hljs-comment">## General</span>
...
<span class="hljs-attribute">DISABLE_SIGNUP</span>=...
<span class="hljs-attribute">API_EXTERNAL_URL</span>={THIS IS THE SUPABASE URL}
...
</code></pre>
<ul>
<li>You should then add <code>.env</code> to the <code>.gitignore</code> file, in order to not commit your secret keys.</li>
</ul>
<h3 id="heading-if-not-a-user-site-nor-custom-domain-set-up-viteconfigts">If not a user site nor custom domain: set up vite.config.ts</h3>
<p>If you are not using the user site (<a target="_blank" href="http://username.github.io">username.github.io</a>), or you are not using a custom domain, then you would need to configure the <code>vite.config.ts</code> , and set <code>base</code> to your repository name, so lines 5-8 look something like this:</p>
<pre><code class="lang-json"><span class="hljs-comment">// https://vitejs.dev/config/</span>
export default defineConfig({
  base: '/svelte_supabase_todo_example/',
  plugins: [svelte()],
...
</code></pre>
<p>This is because normally Vite looks for files in the main URL, which would be <code>username.github.io/{whatever_file}</code> , but since we want it to look for files inside <code>username.github.io/{project_name}/{whatever_file}</code> , we need to set the base to the project name.</p>
<h3 id="heading-change-front-end-authentication">Change front-end authentication</h3>
<p>The original code uses GitHub and Google auth, but we want to use Facebook and Google auth instead, so inside the <code>src/lib/Auth.svelte</code> file, in line 114, we change it into <code>on:click={() =&gt; handleOAuthLogin("github")}</code> . In line 118, we change it into <code>Facebook</code> .</p>
<h3 id="heading-optional-mapping-a-custom-domain-to-the-github-page">Optional: Mapping a custom domain to the GitHub page</h3>
<p>if you have a custom domain, you have to create a file called <code>CNAME</code> inside the public folder, with the content being the custom domain. And change the homepage inside the package.json to be the custom domain.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1673009446761/77ffe1a2-0c98-4102-8591-aca2abd07364.png" alt class="image--center mx-auto" /></p>
<p>Navigate to your DNS provider and create a <code>CNAME</code> record that points your subdomain to the default domain for your site. For example, if you want to use the subdomain <code>subdomain.example.com</code> for your user site, create a <code>CNAME</code> record that points <code>subdomain.example.com</code> to <code>&lt;user&gt;.github.io</code>. Also, add <code>A Record</code> with host <code>@</code> and value <code>185.199.108.153</code>, <code>185.199.109.153</code>, <code>185.199.110.153</code>, <code>185.199.111.153</code>. So your record looks like this:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1673009762106/cbd5315f-ef86-459f-861c-41e9f0f8e87a.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p>Side note: when I was testing this, GitHub seems to dislike having dashes <code>-</code> or underscore <code>_</code> inside the subdomain</p>
</li>
<li><p>If you used a user site <code>&lt;user&gt;.github.io</code> , every other project site you use afterwards, will start with your custom domain, instead of <code>&lt;user&gt;.github.io</code> .</p>
</li>
</ul>
<h2 id="heading-deploy-with-github-pages">Deploy with GitHub pages</h2>
<p>We can install the necessary packages, build the project, and then deploy it.</p>
<pre><code class="lang-bash">npm install
npm run build
npm run deploy
</code></pre>
<p>We then visit the repository on GitHub and go to Settings -&gt; Pages, select Deploy from a branch, and select <code>gh-pages</code> as the branch. Then we can optionally enable HTTPS too.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672937035801/765242dc-b7b4-405e-ab73-d68df76885b5.png" alt class="image--center mx-auto" /></p>
<p>If you used a custom domain, the page would like something like this:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1673009927498/6e99664e-c5cf-4301-84c6-88c0e7a24c0d.png" alt class="image--center mx-auto" /></p>
<p><strong>Congratulations!</strong> When you go to the URL stated above, you will be able to visit the website.</p>
<p>However, only the front end is working for now, in order to have a full-fletch web application, we need the backend to work as well.</p>
<h2 id="heading-make-the-database-table">Make the database table</h2>
<p>The initial Supabase database given to us only contains the auth tables, we need to create the profiles table and todo table on our own. In the hosted version, the snippets are provided, but since we are self-hosting, we will have to copy them from there.</p>
<p>I created a Supabase account, created a project, and then copied the snippet. For you, you can just copy the following code to the run panel and run it.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672937838265/04a6fe24-b916-4afe-9055-53b0c636b7d0.png" alt class="image--center mx-auto" /></p>
<p>Go to the Supabase dashboard, inside the default project -&gt; SQL editor, we can paste the following code there, and click run to run them.</p>
<div class="gist-block embed-wrapper" data-gist-show-loading="false" data-id="6b73ed146dae74ac17d92d6d92bf13cc"><div class="embed-loading"><div class="loadingRow"></div><div class="loadingRow"></div></div><a href="https://gist.github.com/chit-uob/6b73ed146dae74ac17d92d6d92bf13cc" class="embed-card">https://gist.github.com/chit-uob/6b73ed146dae74ac17d92d6d92bf13cc</a></div><p> </p>
<p>Then for the todo list</p>
<div class="gist-block embed-wrapper" data-gist-show-loading="false" data-id="1257c9e959e92b0f60d98bb9da8b7c72"><div class="embed-loading"><div class="loadingRow"></div><div class="loadingRow"></div></div><a href="https://gist.github.com/chit-uob/1257c9e959e92b0f60d98bb9da8b7c72" class="embed-card">https://gist.github.com/chit-uob/1257c9e959e92b0f60d98bb9da8b7c72</a></div><p> </p>
<h2 id="heading-email-sign-up">Email sign up</h2>
<p>To enable email sign-up, we need our backend to be able to send emails. Email servers are very difficult to host, so we will use an SMTP relay service, and have them send the emails for us. There are many choices, I chose to use Sendinblue because it has the biggest free tier email sending limit. It can send up to 300 emails per day, which means up to 9000 emails per month. You can use whatever service you want, you just need to get the required credentials.</p>
<h3 id="heading-get-the-credentials-from-sendinblue">Get the credentials from Sendinblue</h3>
<p>Create an account in <a target="_blank" href="https://www.sendinblue.com/">Sendinblue</a>. In the dashboard, click on the Profile in the upper right corner -&gt; <strong>SMTP and API</strong>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672938369717/fa00e063-c71d-4352-aa25-aea7b176cbc0.png" alt class="image--center mx-auto" /></p>
<p>Inside that page, we need the SMTP Server, Port, Login. We also need to generate an SMTP key, which would act as the password.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672938519683/1bcd1e44-6849-453f-83ae-1a20bd822ba8.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-fill-in-the-credentials-in-the-self-hosted-supabase">Fill in the credentials in the self-hosted Supabase</h3>
<p>Going back to the Supabase backend, in the <code>supabase/docker/</code> directory inside the <code>.env</code> file, we can paste the respective values inside the respective field.</p>
<pre><code class="lang-apache"><span class="hljs-comment">## Email auth</span>
<span class="hljs-attribute">ENABLE_EMAIL_SIGNUP</span>=true
<span class="hljs-attribute">ENABLE_EMAIL_AUTOCONFIRM</span>=false
<span class="hljs-attribute">SMTP_ADMIN_EMAIL</span>={The email address you want the email to be from}
<span class="hljs-attribute">SMTP_HOST</span>=smtp-relay.sendinblue.com
<span class="hljs-attribute">SMTP_PORT</span>={The port given to you}
<span class="hljs-attribute">SMTP_USER</span>={Your login}
<span class="hljs-attribute">SMTP_PASS</span>={Your SMTP key}
<span class="hljs-attribute">SMTP_SENDER_NAME</span>={The name of the sender of your choosing}
</code></pre>
<h2 id="heading-social-oauth">Social OAuth</h2>
<p>We also want to enable Open Authentication login, because it skips the step of creating an account for the user, so users will more likely be inclined to sign up to your site.</p>
<h3 id="heading-google">Google</h3>
<p>Supabase already has <a target="_blank" href="https://supabase.com/docs/guides/auth/social-login/auth-google">a guide on this</a>, however, that is for the managed version, for the self-hosted version, there is something we need to do differently.</p>
<p>First, go to <a target="_blank" href="https://cloud.google.com/">https://cloud.google.com/</a>, Sign in to google if not already, and then click console. It will prompt you to accept the terms of service.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672948135036/1d054252-abd4-4687-bfa5-b3df9cc0521d.png" alt class="image--center mx-auto" /></p>
<p>Inside there, click on <code>Select a Project</code> at the top left and click <code>new project</code>. Fill in your app information then click <code>Create</code>, this will take you to the dashboard for the new project.</p>
<p>In the search bar at the top labelled <code>Search products and resources</code> type <code>OAuth</code>. Click on <code>OAuth consent screen</code> from the list of results. On the <code>OAuth consent screen</code> page select <code>External</code>. Click <code>Create</code>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672949416610/81da8540-53b0-485e-b879-b394da45d5d3.png" alt class="image--center mx-auto" /></p>
<p>On the <code>Edit app registration</code> page fill out your app information. The Application home page is your Svelte website, and The Authorised domain is the domain of your website. Click <code>Save and continue</code> at the bottom.</p>
<p>Click <code>Credentials</code> at the left to go to the <code>Credentials</code> page on the Google Cloud Platform console. Click <code>Create Credentials</code> near the top then select <code>OAuth client ID .</code> On the <code>Create OAuth client ID</code> page, select your application type. If you're not sure, choose <code>Web application</code>. Fill in your app name. At the bottom, under <code>Authorized redirect URIs</code> click <code>Add URI</code>.</p>
<p>Your URI should be what you set as the <code>API_EXTERNAL_URL</code> inside the supabase docker <code>.env</code> file + <code>/auth/v1/callback</code>. For example, <code>https://mybackend.com/auth/v1/callback</code>. Enter your callback URI under <code>Authorized redirect URIs</code> at the bottom. Enter your callback URI in the <code>Valid OAuth Redirect URIs</code> box. Click <code>Save Changes</code> at the bottom right. Click <code>Create</code>. Then we need to copy and save the values under <code>Your Client ID</code> and <code>Your Client Secret</code>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672949421501/e5c909e4-bcfc-4d42-b2c6-c14a6beaf06f.png" alt class="image--center mx-auto" /></p>
<p>Then we can go back to the supabase docker <code>.env</code> file, and add the following code</p>
<pre><code class="lang-apache"><span class="hljs-comment">## Google auth</span>
<span class="hljs-attribute">ENABLE_GOOGLE_SIGNUP</span>=true
<span class="hljs-attribute">GOOGLE_CLIENT_ID</span>={the_client_id_given_to_you}
<span class="hljs-attribute">GOOGLE_CLIENT_SECRET</span>={the_client_secret_given_to_you}
</code></pre>
<p>And in the supabase docker <code>docker-compose.yml</code> file, under the auth environment section, (around line 87), we fill in the following:</p>
<pre><code class="lang-apache"><span class="hljs-attribute">GOTRUE_EXTERNAL_GOOGLE_ENABLED</span>: <span class="hljs-variable">${ENABLE_GOOGLE_SIGNUP}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_GOOGLE_CLIENT_ID</span>: <span class="hljs-variable">${GOOGLE_CLIENT_ID}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_GOOGLE_SECRET</span>: <span class="hljs-variable">${GOOGLE_CLIENT_SECRET}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_GOOGLE_REDIRECT_URI</span>: {the_callback_uri}
</code></pre>
<p>Around this part is a lot of GOTRUE_..., so it should be pretty easy to find. Keep in mind that the ${} things should be taken literally, you shouldn't replace the string inside ${}, but do replace {the callback uri} with the actual callback uri.</p>
<h3 id="heading-facebook">Facebook</h3>
<p>The Facebook guide can also be found on <a target="_blank" href="https://supabase.com/docs/guides/auth/social-login/auth-facebook">the Supabase guide</a>, I am going to briefly go through what you need to do, and the self-hosted instructions not included in their guide.</p>
<p>We first create and configure a Facebook Application on the <a target="_blank" href="https://developers.facebook.com/"><strong>Facebook Developers Site</strong></a>. We first login. Click on <code>My Apps</code> at the top right. Click <code>Create App</code> near the top right. Select your app type and click <code>Continue</code> (Do NOT choose "business", it complicates things). Fill in your app information, then click <code>Create App</code>. This should bring you to the screen: <code>Add Products to Your App</code>. (Alternatively you can click on <code>Add Product</code> in the left sidebar to get to this screen.)</p>
<p>From the <code>Add Products to your App</code> screen, we can click <code>Setup</code> under <code>Facebook Login</code> . Skip the Quickstart screen, instead, in the left sidebar, click <code>Settings</code> under <code>Facebook Login</code> . Enter your callback URI under <code>Valid OAuth Redirect URIs</code> on the <code>Facebook Login Settings</code> page. Enter this in the <code>Valid OAuth Redirect URIs</code> box. Click <code>Save Changes</code> at the bottom right.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672951160599/9e458a81-86c8-43b3-8e7c-0cae1c7de680.png" alt class="image--center mx-auto" /></p>
<p>Be aware that you have to set the right access levels on your Facebook App to enable 3rd party applications to read the email address. From the <code>App Review -&gt; Permissions and Features</code> screen: Click the button <code>Request Advanced Access</code> on the right side of <code>public_profile</code> and <code>email</code> .</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1672951082358/8f2cd056-255c-4d3c-8ed8-e3d7dee17959.png" alt class="image--center mx-auto" /></p>
<p>Click <code>Settings / Basic</code> in the left sidebar. Copy your App ID from the top of the <code>Basic Settings</code> page. Under <code>App Secret</code> click <code>Show</code> then copy your secret. Make sure all required fields are completed on this screen.</p>
<p>Then similarly, we go to the supabase docker <code>.env</code> file, and add the following code:</p>
<pre><code class="lang-apache"><span class="hljs-comment">## Facebook auth</span>
<span class="hljs-attribute">ENABLE_FACEBOOK_SIGNUP</span>=true
<span class="hljs-attribute">FACEBOOK_CLIENT_ID</span>={the_id}
<span class="hljs-attribute">FACEBOOK_CLIENT_SECRET</span>={the_secret}
</code></pre>
<p>And in the supabase docker <code>docker-compose.yml</code> file, under the google stuff we just added, we add:</p>
<pre><code class="lang-apache"><span class="hljs-attribute">GOTRUE_EXTERNAL_FACEBOOK_ENABLED</span>: <span class="hljs-variable">${ENABLE_FACEBOOK_SIGNUP}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_FACEBOOK_CLIENT_ID</span>: <span class="hljs-variable">${FACEBOOK_CLIENT_ID}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_FACEBOOK_SECRET</span>: <span class="hljs-variable">${FACEBOOK_CLIENT_SECRET}</span>
<span class="hljs-attribute">GOTRUE_EXTERNAL_FACEBOOK_REDIRECT_URI</span>: {the_callback_uri}
</code></pre>
<ul>
<li>Note that both Google and Facebook Auth are in testing mode, to use them in production, you need to change them to production mode.</li>
</ul>
<h3 id="heading-replace-site-url">Replace site URL</h3>
<p>also inside the supabase docker <code>.env</code> file, you can change the <code>SITE_URL=http://localhost:3000</code> , and set the URL to the URL of your front-end website.</p>
<h2 id="heading-restart-supabase">Restart Supabase</h2>
<p>Now that everything is set, you can restart Supabase. If you followed <a target="_blank" href="https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide">my last tutorial</a>, Supabase would be running inside a screen, do <code>screen -r supabase</code> , then <code>ctrl + c</code> to stop the already running docker, then <code>docker compose up</code> to start the docker again. If it is not already running, you can just do <code>docker compose up</code> inside the supabase docker directory to start Supabase.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Now you should get a fully functional to-do list web app, you can try the different logins, from email sign-up and sending an email to confirm your account, to one-click OAuth with Google and Facebook. After logging in, you can also add and remove to-do items, and they will stay there when you log in with different devices.</p>
<p>This is my second touch on making a tutorial, if there are any steps unclear, feel free to leave a comment and I'll try to answer them. hope this tutorial helps.</p>
]]></content:encoded></item><item><title><![CDATA[How to Self-host Supabase: A complete guide]]></title><description><![CDATA[In this guide, I will explain how to self-host Supabase. Including how to purchase and secure a virtual private server (VPS), install and set up Supabase, Set up a reverse proxy using Nginx, and get a secure socket layer (SSL) certificate to accept H...]]></description><link>https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide</link><guid isPermaLink="true">https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide</guid><category><![CDATA[supabase]]></category><category><![CDATA[Tutorial]]></category><category><![CDATA[vps]]></category><category><![CDATA[nginx]]></category><category><![CDATA[SSL]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Sat, 24 Dec 2022 22:35:42 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1671921046034/e46c1532-78f9-43c8-b2a6-5cc11bf64908.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this guide, I will explain how to self-host Supabase. Including how to purchase and secure a virtual private server (VPS), install and set up Supabase, Set up a reverse proxy using Nginx, and get a secure socket layer (SSL) certificate to accept HTTPS traffic.</p>
<p>This is my first time writing a complementary guide, so if there is anything unclear, or you have any feedback, please leave it in the comment below.</p>
<h2 id="heading-what-you-will-achieve">What you will achieve</h2>
<p>If you follow the steps, you will end up having a password-protected Supabase server hosted on a VPS, a reverse proxy that redirects traffic from the domain name to the respective endpoint on the Supabase back end with an SSL certificate.</p>
<p>To continue, you will need:</p>
<ol>
<li><p>A domain name, free or paid, that you can set the A or CNAME record</p>
</li>
<li><p>A VPS running Ubuntu 20+, if you don't already this guild will also introduce one</p>
</li>
</ol>
<h2 id="heading-buy-a-vps-and-set-up">Buy a VPS and set up</h2>
<p>To self-host a Supabase back end, obviously, you need a VPS. If you already have one, you can move on to the section about securing the VPS and see if there is more you can do to <a class="post-section-overview" href="#heading-securing-the-vps">secure your VPS</a>. If not, the following guide will talk about how to purchase one.</p>
<h3 id="heading-purchasing-a-vps">Purchasing a VPS</h3>
<p>I will be using OxideHost VPS because they are relatively cheap. I've been using them and they are pretty reliable. For those interested, this is my <a target="_blank" href="https://billing.oxide.host/aff.php?aff=166">affiliate link</a>.</p>
<p>After choosing the VPS and purchasing it with Ubuntu 22.04 installed, an email is sent to give me information about how to connect to my VPS. The most important part is the hostname, IP, and default password.</p>
<h3 id="heading-connect-to-it">Connect to it</h3>
<pre><code class="lang-bash">ssh {username}@{ip-address}
</code></pre>
<p>To connect to the VPS I just purchased, the simplest way is to use the <code>ssh</code> command in the command line. However, using this method, we will have to type this command and enter the password every time. A better method would be to use SSH clients such as <a target="_blank" href="https://www.putty.org/">PuTTY</a> which is free. I will be using <a target="_blank" href="https://termius.com/">Termius</a> since it has a nice interface and is free for students.</p>
<p><img src="https://billing.oxide.host/images/kb/3_d144a7.png" alt /></p>
<p>I will create a new host, inputting the hostname given by the email, and username which is <code>root</code> by default, and the password was given to me. It will ask you a question about do you want to connect to it, click <code>yes</code>. Then I will have connected to the VPS for the first time.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671622141798/yidZnFtC6.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-securing-the-vps">Securing the VPS</h3>
<p>Now that we are connected to our VPS, we want to make sure we are the only people having access to it, so we need to take measures to secure it.</p>
<h4 id="heading-update-its-software">update its software</h4>
<pre><code class="lang-bash">apt update
apt dist-upgrade
</code></pre>
<p>The first task is to update it since updated software usually has its security vulnerabilities patches. This also upgrades the Linux distribution. There will be a few questions the command prompt asks you, you can type <code>y</code> and then press enters to say yes to them. and if it is a multiple choice, you can press <code>tab</code> to navigate around, and press enters to confirm the selection.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671622242369/4-gFu4AlO.png" alt class="image--center mx-auto" /></p>
<h4 id="heading-add-another-user">Add another user</h4>
<p>Now that you acquired your VPS, hackers also want to get access to it. There are tons of robots on the internet cracking passwords of random VPS, they usually try to log in as root since it is the default configuration. So we will add a new user, pass the root permission to it, and disable root login so those bad robots won't be able to log in to our VPS.</p>
<pre><code class="lang-bash">adduser {username}
usermod -aG sudo {username}
</code></pre>
<p>The username cannot have spaces within, so you can use an underscore like <code>learn_supabase</code> or <code>learnsupabase</code> . Then it will ask us some questions, we can input blank if we don't want to answer them. But we need to fill in a <strong>secure password</strong>, I recommend inputting a randomly generated password using some <a target="_blank" href="https://www.dashlane.com/features/password-generator">secure password generators</a> so that there is no way for anyone to guess it. Then we can give the user Sudo privileges. And <strong>try to log in using the new username and password</strong> by creating a new host and connecting it.</p>
<h4 id="heading-disable-the-root-login">Disable the root login</h4>
<pre><code class="lang-bash">nano /etc/ssh/sshd_config
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671623556148/9wNCAihMU.png" alt class="image--center mx-auto" /></p>
<p>Now we need to change the <code>Port</code> to some random number to make it harder to guess, the number must be smaller than 65535. And we also change the link <code>PermitRootLogin</code> to <code>no</code>, so we can no longer log in as root.</p>
<p>After making the change, we can do <code>ctrl + o</code> to write out the changes, and <code>ctrl + x</code> to exit nano. Then we do the following command to restart the <code>sshd</code> service, so the new configuration will come into effect.</p>
<pre><code class="lang-bash">systemctl restart sshd
</code></pre>
<h2 id="heading-install-supabase">Install Supabase</h2>
<h3 id="heading-install-docker">Install docker</h3>
<p>We need docker for this, a tutorial about how to install docker can be found <a target="_blank" href="https://www.linode.com/docs/guides/installing-and-using-docker-on-ubuntu-and-debian/">here</a>. In this tutorial, I will summarize what commands we need to run. This command assumes you are using Ubuntu, if you are using other distros, please refer to the above tutorial link.</p>
<pre><code class="lang-bash">sudo apt install apt-transport-https ca-certificates curl gnupg lsb-release
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | sudo gpg --dearmor -o /usr/share/keyrings/docker-archive-keyring.gpg
<span class="hljs-built_in">echo</span> <span class="hljs-string">"deb [arch=amd64 signed-by=/usr/share/keyrings/docker-archive-keyring.gpg] https://download.docker.com/linux/ubuntu <span class="hljs-subst">$(lsb_release -cs)</span> stable"</span> | sudo tee /etc/apt/sources.list.d/docker.list &gt; /dev/null
sudo apt update
sudo apt install docker-ce docker-ce-cli containerd.io
sudo apt install docker-compose-plugin
</code></pre>
<p>Then, we also need to enable the user to use docker.</p>
<pre><code class="lang-bash">sudo usermod -aG docker {username}
</code></pre>
<h3 id="heading-clone-supabase">Clone Supabase</h3>
<p>Then we want to clone the Supabase from Github.</p>
<pre><code class="lang-bash">git <span class="hljs-built_in">clone</span> --depth 1 https://github.com/supabase/supabase
</code></pre>
<h2 id="heading-set-up-supabase-secrets">Set up Supabase secrets</h2>
<pre><code class="lang-bash"><span class="hljs-built_in">cd</span> supabase/docker/
cp .env.example .env
</code></pre>
<p>Now we get inside the docker directory and copy the file for the environment variables.</p>
<h3 id="heading-generate-passwords">Generate passwords</h3>
<p>We open the <code>.env</code> file, and we will see a lot of settings.</p>
<pre><code class="lang-bash">nano .env
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671721310426/b16d8cd9-f1d9-4496-b799-fe3712756a21.png" alt class="image--center mx-auto" /></p>
<p>We want to set the passwords for the above fields. For <code>POSTGRES_PASSWORD</code> , you can use a randomly generated password. For the JWT secret, we have to use the <a target="_blank" href="https://supabase.com/docs/guides/self-hosting#api-keys">custom key generator</a>. you can copy the JWT secret and paste it into the file, and then we can select <code>ANON_KEY</code> , hit generate, copy it to the <code>ANON_KEY</code> field, then the <code>SERVICE_KEY</code> field.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671721891892/5901353a-8c78-4fe6-b4ba-251fa0de343e.png" alt class="image--center mx-auto" /></p>
<pre><code class="lang-bash">nano volumes/api/kong.yml
</code></pre>
<p>Then we also copy the same key to the kong config, pasting the new <code>anon</code> key and <code>service</code> key.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671722061768/11c398ed-6533-45ef-bd6c-0c6024e06edb.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-start-docker">Start docker</h3>
<p>Now we open a new screen and name it so that we can connect to it later</p>
<pre><code class="lang-bash">screen
</code></pre>
<p>Here you will see the introduction of the screen, hit <code>enter</code> . Then we can do <code>ctrl + a</code> then <code>ctrl + d</code> to disconnect from the screen. In the console, you will see your screen number <code>[detached from 3775963.pts-0.hostname]</code> , the <code>3775963</code> number is the screen number.</p>
<pre><code class="lang-bash">screen -S {screen_number} -X sessionname supabase
screen -r supabase
</code></pre>
<p>Now we have reconnected to screen <code>supabase</code> , when we disconnect from it, we use <code>ctrl + a</code> then <code>ctrl + d</code> .</p>
<pre><code class="lang-bash"><span class="hljs-built_in">cd</span> supabase/docker/
docker compose up
</code></pre>
<p>Now the supabase docker is started, it will download a bunch of things, and you can connect to your domain in port 3000.</p>
<h2 id="heading-domain-mapping">Domain mapping</h2>
<p>Now I want to point a custom domain to the backend. I used Namecheap, and I create an A record, pointing a subdomain to the IP address of my VPS.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671722678692/01578d23-97a3-4d36-8b61-42bab6955e57.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-reverse-proxy">Reverse proxy</h2>
<p>Reverse proxy points the connections to the correct port. The part is explained very well in the <a target="_blank" href="https://www.linode.com/docs/guides/installing-supabase/#using-a-reverse-proxy">Linode blog post</a>, so I am only going to briefly talk about what commands needed to be run, you can refer to <a target="_blank" href="https://www.linode.com/docs/guides/installing-supabase/#using-a-reverse-proxy">their blog post</a> for a more detailed explanation.</p>
<p>We first install Nginx for reverse proxy.</p>
<pre><code class="lang-bash">sudo apt install nginx
sudo systemctl status nginx
</code></pre>
<p>If we connect to our domain, we will see the Nginx welcome page.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671724075251/f835e84f-cfec-42d6-8f7d-b1087629ddf3.png" alt class="image--center mx-auto" /></p>
<pre><code class="lang-bash">sudo nano /etc/nginx/sites-available/{your_domain_name}
</code></pre>
<p>then we copy the following content (from the <a target="_blank" href="https://www.linode.com/docs/guides/installing-supabase/#using-a-reverse-proxy">Linode guide</a>) into the file, replacing the <code>example.com</code> with your custom domain.</p>
<pre><code class="lang-nginx"><span class="hljs-attribute">map</span> <span class="hljs-variable">$http_upgrade</span> <span class="hljs-variable">$connection_upgrade</span> {
    <span class="hljs-attribute">default</span> upgrade;
    '' close;
}

<span class="hljs-attribute">upstream</span> supabase {
    <span class="hljs-attribute">server</span> localhost:<span class="hljs-number">3000</span>;
}

<span class="hljs-attribute">upstream</span> kong {
    <span class="hljs-attribute">server</span> localhost:<span class="hljs-number">8000</span>;
}

<span class="hljs-section">server</span> {
    <span class="hljs-attribute">listen</span> <span class="hljs-number">80</span>;
    <span class="hljs-attribute">server_name</span> example.com;

    <span class="hljs-comment"># REST</span>
    <span class="hljs-attribute">location</span> <span class="hljs-regexp">~ ^/rest/v1/(.*)$</span> {
        <span class="hljs-attribute">proxy_set_header</span> Host <span class="hljs-variable">$host</span>;
        <span class="hljs-attribute">proxy_pass</span> http://kong;
        <span class="hljs-attribute">proxy_redirect</span> <span class="hljs-literal">off</span>;
    }

    <span class="hljs-comment"># AUTH</span>
    <span class="hljs-attribute">location</span> <span class="hljs-regexp">~ ^/auth/v1/(.*)$</span> {
        <span class="hljs-attribute">proxy_set_header</span> Host <span class="hljs-variable">$host</span>;
        <span class="hljs-attribute">proxy_pass</span> http://kong;
        <span class="hljs-attribute">proxy_redirect</span> <span class="hljs-literal">off</span>;
    }

    <span class="hljs-comment"># REALTIME</span>
    <span class="hljs-attribute">location</span> <span class="hljs-regexp">~ ^/realtime/v1/(.*)$</span> {
        <span class="hljs-attribute">proxy_redirect</span> <span class="hljs-literal">off</span>;
        <span class="hljs-attribute">proxy_pass</span> http://kong;
        <span class="hljs-attribute">proxy_http_version</span> <span class="hljs-number">1</span>.<span class="hljs-number">1</span>;
        <span class="hljs-attribute">proxy_set_header</span> Upgrade <span class="hljs-variable">$http_upgrade</span>;
        <span class="hljs-attribute">proxy_set_header</span> Connection <span class="hljs-variable">$connection_upgrade</span>;
        <span class="hljs-attribute">proxy_set_header</span> Host <span class="hljs-variable">$host</span>;
    }

    <span class="hljs-comment"># STUDIO</span>
    <span class="hljs-attribute">location</span> / {
        <span class="hljs-attribute">proxy_set_header</span> Host <span class="hljs-variable">$host</span>;
        <span class="hljs-attribute">proxy_pass</span> http://supabase;
        <span class="hljs-attribute">proxy_redirect</span> <span class="hljs-literal">off</span>;
        <span class="hljs-attribute">proxy_set_header</span> Upgrade <span class="hljs-variable">$http_upgrade</span>;
    }
}
</code></pre>
<p>After that, we can do <code>ctrl + o</code> and <code>ctrl + x</code> to write and exit. Then we can create a symbolic link of the config from sites-available to sites-enabled, link the default config, and then restart Nginx.</p>
<pre><code class="lang-bash">sudo ln -s /etc/nginx/sites-available/{your_domain_name} /etc/nginx/sites-enabled/{your_domain_name}
sudo unlink /etc/nginx/sites-enabled/default
sudo systemctl restart nginx
</code></pre>
<p>Now, we can connect to Supabase using the domain name.</p>
<h2 id="heading-get-ssl">Get SSL</h2>
<p>In this part, we will use Certbot to give our Supabase an SSL certificate. Their official website have pretty good instructions, so you can follow that. But I'll also let you know what commands I used.</p>
<p>We first install <code>core</code>, then <code>certbot</code>. Then we configure it, and at last try how it auto-renews, and then restart Nginx for it to take effect.</p>
<pre><code class="lang-bash">sudo snap install core; sudo snap refresh core
sudo snap install --classic certbot
sudo ln -s /snap/bin/certbot /usr/bin/certbot
sudo certbot --nginx
sudo certbot renew --dry-run
sudo systemctl restart nginx
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671894627444/62293b4c-8763-4e17-86ea-bc827d8b8504.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-secure-supabase">Secure Supabase</h2>
<p>The Supabase dashboard everyone can connect to, to secure it, we do this to generate the password. It will ask you to type in the password two times, so it is best you generate a password from a secure password generator and then paste it there twice.</p>
<pre><code class="lang-bash">sudo htpasswd -c /etc/apache2/.htpasswd {admin_username}
</code></pre>
<p>Then we can check the password with</p>
<pre><code class="lang-bash">cat /etc/apache2/.htpasswd
</code></pre>
<p>Then we edit the Nginx config to password protect</p>
<pre><code class="lang-bash">sudo nano /etc/nginx/sites-available/{domain name}
</code></pre>
<p>and make it so that in the studio</p>
<pre><code class="lang-nginx">    <span class="hljs-comment"># STUDIO</span>
    <span class="hljs-attribute">location</span> / {
        <span class="hljs-attribute">proxy_set_header</span> Host <span class="hljs-variable">$host</span>;
        <span class="hljs-attribute">proxy_pass</span> http://supabase;
        <span class="hljs-attribute">proxy_redirect</span> <span class="hljs-literal">off</span>;
        <span class="hljs-attribute">proxy_set_header</span> Upgrade <span class="hljs-variable">$http_upgrade</span>;

        <span class="hljs-attribute">auth_basic</span>           <span class="hljs-string">"Administrator’s Area"</span>;
        <span class="hljs-attribute">auth_basic_user_file</span> /etc/apache2/.htpasswd;
    }
</code></pre>
<p>at last restart</p>
<pre><code class="lang-bash">sudo service nginx restart
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1671895136816/2a9eeac2-07cd-4144-8be1-9fca76ed2f11.png" alt class="image--center mx-auto" /></p>
<p>Here you can fill in the username as the <code>{admin_username}</code> you put in, and the password as the password you filled in.</p>
<h2 id="heading-reconfigure-supabase">Reconfigure Supabase</h2>
<p>Now, we want to set the domain for the Supabase backend and external API.</p>
<pre><code class="lang-bash">screen -r supabase
</code></pre>
<p>Then we press <code>ctrl + c</code> to end the Supabase docker instance. Then we edit the <code>.env</code> file.</p>
<pre><code class="lang-bash">nano .env
</code></pre>
<p>We need to change the</p>
<pre><code class="lang-apache">...
<span class="hljs-attribute">API_EXTERNAL_URL</span>=https://{your_domain}
...
<span class="hljs-attribute">SUPABASE_PUBLIC_URL</span>=https://{your_domain} # replace if you intend to use Studio outside of localhost
</code></pre>
<p>Then we can do <code>ctrl + o</code> and <code>ctrl + x</code> to save the changes, then do the following to get Supabase back up.</p>
<pre><code class="lang-bash">docker compose up
</code></pre>
<p>Then we can exit the screen by doing <code>ctrl + a</code> then <code>ctrl + d</code> to disconnect without terminating the process inside.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Congratulations, you now have a working Supabase backend, protected by a password, inside a certified domain. Hopefully, my explanation is clear, is there are any questions or comments, feel free to leave a comment. Your next step would be to develop your front end, which I will cover in <a target="_blank" href="https://blog.cpbprojects.me/set-up-a-svelte-todo-list-on-self-hosted-supabase-email-sign-up-google-facebook-auth-host-on-github-pages">the next part of this guide</a>.</p>
]]></content:encoded></item><item><title><![CDATA[What I learned from self-hosting a Supabase Svelte project: Part 2]]></title><description><![CDATA[In part 1, we went over self-hosting supabase, setting up Svelte, enabling third-party authentication and email sending. In this blog post, I will talk about what I learned about routing, SSL, domain name, local storage, and hosting svelte using GitH...]]></description><link>https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-2</link><guid isPermaLink="true">https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-2</guid><category><![CDATA[Svelte]]></category><category><![CDATA[GitHubPages]]></category><category><![CDATA[nginx]]></category><category><![CDATA[SSL]]></category><category><![CDATA[domain]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Thu, 15 Dec 2022 17:17:03 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1671124518635/F0TVhFRe8.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In <a target="_blank" href="https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-1">part 1</a>, we went over self-hosting supabase, setting up Svelte, enabling third-party authentication and email sending. In this blog post, I will talk about what I learned about routing, SSL, domain name, local storage, and hosting svelte using GitHub pages.</p>
<h2 id="heading-routing-with-nginx">Routing with Nginx</h2>
<p>Currently, there are lots of programs running on various ports of my VPS. Supabase studio on 3000, Svelte on 5173, PostgreSQL running on port 5432, and Kong running on 8000 and 8443 for HTTP and HTTPS respectively. How do I take the users that want to connect to my front end (Svelte) to port 5173, and secure my Supabase studio?</p>
<p>The answer is routing, we can set up a router, or a reverse proxy, to take users that are connecting from HTTP and HTTPS, which is port 80 and port 443, and route them to port 5173, my Svelte front end. To do that, I will use Nginx, because that is the most popular one. In the following, I'm going to summarize what I did to get it working, If you want a more detailed explanation or code examples, I recommend you check out <a target="_blank" href="https://www.linode.com/docs/guides/installing-supabase/">this article from Linode</a>, I find it very helpful.</p>
<p>In Nginx, we create a <code>server</code> to listen to a single port, then we tell it how to redirect based on <code>location</code>. Locations are the URL path after the domain name, the location for <code>example.com</code> is <code>/</code>, and <code>example.com/terms/</code> is <code>^terms/(.*)$</code> . Then in the location, we can ask it to redirect to a certain port, such as <code>proxy_pass http://localhost:5173</code> , this would redirect the users connecting from this port at this location to this address instead.</p>
<p>An important thing for Nginx is that we have to configure which rules are active. There is a directory called <code>sites-available</code>, and one called <code>sites-enabled</code>. What we have to do it to configure the settings in <code>sites-available</code>, then created a symbolic link to the <code>sites-enabled</code> directory. Then we have to restart the Nginx service to apply changes. An important thing to do is to unlink sites you aren't using, I once forgot to unlink the default rules and was wondering how the default rules were applying for so long.</p>
<h2 id="heading-getting-ssl">Getting SSL</h2>
<p>After that, we can redirect HTTP traffic, but we need an SSL certificate if we want HTTPS traffic. They provide credibility and encrypt traffic. There are 3 ways to get a certificate:</p>
<ol>
<li><p>Purchase it from a trusted certificate authority (CA), the cost can vary from £5 a year, up to who knows how much</p>
</li>
<li><p>Let a proxy like <a target="_blank" href="https://www.cloudflare.com/ssl/">Cloudflare</a> do it for you, users will connect to Cloudflare with the certificate, and then Cloudflare redirects the traffic to your site, and the users see the certificate from Cloudflare. This is free for personal or hobby projects.</p>
</li>
<li><p>Generate your own using a service like <a target="_blank" href="https://certbot.eff.org/">Certbot</a>, it generates the certificate for you and you don't have to pay a penny. But the downside is that you will have to configure it yourself.</p>
</li>
</ol>
<p>I went with Certbot since it is free. <a target="_blank" href="https://certbot.eff.org/">The website</a> contains very detailed instructions about installing it. After getting it, I need to configure Nginx to use the certificate, redirect HTTPS traffic to my front end, and ask HTTP traffic to upgrade to HTTPS traffic.</p>
<h2 id="heading-domain-name">Domain name</h2>
<p>Then I needed a domain name for my website, you can think of a domain name as basically a link for the website. I went over the three providers that offer free 1-year in Github student pack, namely <em>Namecheap</em>, <em>Name.com</em> and <em>.TECH</em>. I analyzed what service they provide and how much the domain is going to cost after the 1-year free trial. Namecheap has the cheapest domain name after the free trial, so I decided to use it.</p>
<p>Now I have to configure it to redirect traffic from the domain name to the IP address of my VPS. There are <a target="_blank" href="https://www.namecheap.com/support/knowledgebase/article.aspx/579/2237/which-record-type-option-should-i-choose-for-the-information-im-about-to-enter/">different types of records</a>, the ones that you're most likely to use are an A record or a CNAME record. <em>A records</em> allow you to associate a domain name or a subdomain to an IP address, <em>CNAME</em> records allow you to point your domain name or subdomain to a hostname.</p>
<p>I pointed the subdomain I wanted to the IP address of my VPS, then pointed the www version of the domain to the domain without the www part.</p>
<h2 id="heading-storing-data">Storing data</h2>
<p>For the encryption function of the website, I want people to be able to generate an encryption key locally and have it stored locally. So the data they post to the server will be encrypted. This makes it so that even if the website got hacked, the hackers couldn't read the encrypted content.</p>
<p>Users don't want to enter the key every time they visit the website, I need a way to store the encryption key locally. I was thinking to use cookies for this task, but then I found out that cookies get sent back to the web server for every request, which isn't secure at all. So I searched that there is Local Storage in browsers now, and the information there will not be sent to the web server, so I decided to use it.</p>
<p>It is very easy to use, to write something to local storage, you do <code>localStorage.setItem("name", "Chris");</code>, and to read something, you do <code>let myName = localStorage.getItem("name");</code>.</p>
<h2 id="heading-svelte-can-be-hosted-statically">Svelte can be hosted statically</h2>
<p>I don't know how I didn't know this, I always thought Svelte has to be hosted on a VPS, maybe because I developed Django apps before and they require a dedicated host. I found this out because I was talking to a friend about how I am buying a VPS to host the website, and he asked me why didn't I use Github pages instead. I said the website needed to be hosted, but he said it doesn't. I looked into it, and turns out he was right! Svelte doesn't have to be hosted dynamically. When we run <code>npm run dev</code>, it runs the developer version, which is dynamically hosted because it reacts to changes. But if we do <code>npm run build</code>, the page is built and the <code>dist</code> directory. Then we can deploy through any static site hosts, including Github pages.</p>
<p>GitHub pages can be used for public repositories for free accounts and private repositories for Pro accounts, which I have from the GitHub student pack. Moreover, GitHub is a big website, so I'm sure loading a front end there is very fast, as they would have good content delivery. Security isn't a big concern either, since only the front end is hosted there, and the important data is secured in the back end. Therefore I will be using GitHub pages to host my site on second thoughts.</p>
<h2 id="heading-hosting-on-github-pages">Hosting on GitHub Pages</h2>
<p>To host on Github pages, if we want to use the user GitHub page (<code>username.github.io</code>), instead of <code>username.github.io/reponame</code>, we need to have the website in the repository with the name <code>username.github.io</code> . Then we can control what to deploy, we can either set up a GitHub workflow, which is GitHub doing something for you automatically, or just deploy from a branch.</p>
<p>To automate this process, there is an npm package called <a target="_blank" href="https://www.npmjs.com/package/gh-pages">gh-pages</a>. It enables you to deploy your new version of the build, which is inside the <code>dist</code> folder for Svelte projects, to a new branch. Then if you configure your GitHub page to deploy that branch, this would automatically be deployed.</p>
<p>There are some important things to set up. We have to configure the homepage to the GitHub page website in <code>package.json</code>, as well as set up a new run config called <code>deploy</code>, which runs <code>gh-pages -d dist</code> . If we are using a custom domain, we also need to create a <code>CNAME</code> file inside the <code>public/</code> directory with the content being the custom domain name. Then when we run <code>npm run build</code> then <code>npm run deploy</code>, our GitHub page will be deployed.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>In this blog post, I discussed my experience with self-hosting a Supabase Svelte project. I covered setting up routing with Nginx, obtaining an SSL certificate with Certbot, securing a domain name with Namecheap, and hosting the project on GitHub Pages. Overall, I found that self-hosting my project was a valuable learning experience that allowed me to gain a deeper understanding of web hosting and the tools involved.</p>
]]></content:encoded></item><item><title><![CDATA[What I learned from self-hosting a Supabase Svelte project: Part 1]]></title><description><![CDATA[In this blog post, I will talk about what I learned from making a Supabase Svelte project. The plan for making this project is explained in this blog post. I will go over the obstacles I've overcome when trying the different building blocks of the pr...]]></description><link>https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-1</link><guid isPermaLink="true">https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-1</guid><category><![CDATA[learning]]></category><category><![CDATA[supabase]]></category><category><![CDATA[oauth]]></category><category><![CDATA[email]]></category><category><![CDATA[Git]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Wed, 14 Dec 2022 10:32:46 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1671013792794/_iC_31hGj.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog post, I will talk about what I learned from making a Supabase Svelte project. The plan for making this project is explained in <a target="_blank" href="https://blog.cpbprojects.me/my-plan-for-building-a-prayer-tracker-website">this blog post</a>. I will go over the obstacles I've overcome when trying the different building blocks of the project, namely self-hosting supabase, setting up Svelte, enabling third-party authentication and email sending.</p>
<h2 id="heading-self-hosting-supabase">Self hosting Supabase</h2>
<p>The advantage of Supabase is that it supports self-host, so I don't have to worry about limits and bills. I went ahead and searched for the <a target="_blank" href="https://supabase.com/docs/guides/hosting/overview">official tutorial</a>. I expected a very complete tutorial that covers every step, but what I got was a very high-level overview of how to do it, talking about architecture, configuration, database etc. Then there is another tutorial on <a target="_blank" href="https://supabase.com/docs/guides/hosting/docker">self-hosting with Docker</a>. To their credit, there is a very concise script for getting the docker running, but that's all you get in terms of a hands-on tutorial. The rest of the article is yet another high-level overview of important things to do. This is fine for experienced developers, but as a new developer, I would have appreciated better guidance.</p>
<p>At that time, I didn't know what is docker-compose, how I deploy the docker from the local environment to the production environment (my VPS), and how to set up the secrets, so I kept on searching for more helpful tutorials.</p>
<p>Luckily, I stumbled upon this <a target="_blank" href="https://www.youtube.com/watch?v=0bqxrm4PnMA">Youtube tutorial</a>, huge shout out to <a target="_blank" href="https://www.youtube.com/@silkodyssey">Kelvin Pompey</a>. He demonstrated the entire process of self-hosting Supabase with an Ubuntu Server on Digital Ocean, which I was able to follow step by step, and reproduce the working Supabase backend. I didn't use Digital Ocean, I used a VPS from <a target="_blank" href="https://billing.oxide.host/aff.php?aff=166">Oxide Host</a> since I've been using them for a few years starting from hosting a Discord bot, and the service is reliable.</p>
<p>I learned how to set up supabase, <a target="_blank" href="https://www.digitalocean.com/community/tutorials/how-to-install-and-use-docker-compose-on-ubuntu-20-04">how to install docker-compose</a> and that docker-compose starts the processes inside the <code>docker-compose.yml</code> file, and to update the secrets inside the <code>.env</code> file before deploying the Supabase.</p>
<h2 id="heading-setting-up-svelte-front-end">Setting up Svelte front end</h2>
<p>Now that I've set up the backend, I want some example frontend code to test if the backend is working. I was originally thinking about making a React app as it is pretty popular. But I found an <a target="_blank" href="https://github.com/supabase/supabase/tree/master/examples/todo-list/sveltejs-todo-list">officially maintained to-do list app</a> written in Svelte, since it is very similar to a prayer tracker, I've decided to use that as the skeleton code and build on that.</p>
<p>The code is within the <a target="_blank" href="https://github.com/supabase/supabase">Supabase repo</a>, and the Supabase repo is huge, so I don't want to clone the entire thing, just the part I need.</p>
<h3 id="heading-only-cloning-a-part-of-a-repo">Only cloning a part of a repo</h3>
<p>I've learned this from <a target="_blank" href="https://stackoverflow.com/questions/600079/how-do-i-clone-a-subdirectory-only-of-a-git-repository">this StackOverflow post</a>. We need the latest version of Git, I failed at first because I was using an older version of git that doesn't support sparse checkout.</p>
<p>I am going to use the Supabase repo as an example, the <code>--depth 1</code> flag stops git from cloning any files inside folders. and the <code>--filter=blob:none</code> will filter out all blobs (file contents) until needed by Git, and <code>--sparse</code> just tells git you plan to git checkout later.</p>
<p>Then we <code>cd</code> into the directory, and sparse checkout, with <code>set</code> as the path of the content we want to check out.</p>
<pre><code class="lang-bash">git <span class="hljs-built_in">clone</span> --depth 1 --filter=blob:none --sparse https://github.com/supabase/supabase
<span class="hljs-built_in">cd</span> supabase
git sparse-checkout <span class="hljs-built_in">set</span> examples/todo-list/sveltejs-todo-list
</code></pre>
<p>Then I just needed to run <code>npm install</code> to install the dependencies and <code>npm run dev</code> to run the Svelte project. I also installed <code>nvm</code> to manage the version of Node js.</p>
<p>I also needed to update the <code>.env</code> file to paste the Supabase host address, and the <code>anon_key</code> used in Supabase. Side note: I used the default credentials the first time I used Supabase, when I changed the credentials, the database no longer recognised me, I had to do <code>docker-compose down --volumes</code> to reset everything after I updated the <code>.env</code> credentials.</p>
<h2 id="heading-enabling-third-party-authentication">Enabling third-party authentication</h2>
<p><a target="_blank" href="https://supabase.com/docs/guides/auth">Supabase supports third-party authentication</a>, and it does that by recognising the email address. So let's say you used the same email address <code>example@example.com</code> on Google and Facebook, signing in with either will lead you to log in to the same account.</p>
<p>These third parties use an authentication standard Open Authentication (OAuth). OAuth is a protocol that allows applications to access user information without requiring their password. This allows users to securely share their information with third-party applications, minimizing the risk of a security breach.</p>
<p>Basically, instead of identifying a user using a password, the app asks Google, "do you trust this guy?". If Google trusts them, the app trusts them too.</p>
<p>To enable this, we have to create an account on the OAuth provider. I was able to do it for Google and Facebook. For Apple, the OAuth requires me to create an Apple App, and that is locked behind a paywall. I need to pay 79 pounds a year to access that, so I won't be using that.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670941717984/YdBc-hAi6.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-problem-with-twitter-oauth">Problem with Twitter OAuth</h3>
<p>When I try to do it on Twitter, the funniest/most annoying issue happened. On the sign-up page, there is a button that asks do I want to receive marketing emails from them, I said no. Then they said they have sent a confirmation email to my email account, I didn't receive it, so I clicked have it resent. Guess what they said: "There was a problem resending the confirmation email, User has disabled email notification from Twitter...". I checked my Twitter setting and email notifications are on.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670941353702/F4vyiSMtc.png" alt class="image--center mx-auto" /></p>
<p>I then registered from the Twitter developer account on another Twitter account, this time ticking the marketing email field, this time I was able to get the confirmation screen and do it. So I think this was the issue. The funniest thing is that I was unable to fix this issue, I cannot find the setting to re-enable marketing email on the Twitter developer page.</p>
<h2 id="heading-email-sending">Email sending</h2>
<p>For email sending, originally I thought I can send emails from my VPS. So I searched for <em>How to send emails through VPS</em>. Every result uses Gmail or another managed email service, apparently, we can't simply use a VPS to send an email, I needed something called an SMTP server.</p>
<p>So I went on and search for <em>How to host an SMTP server on a VPS</em>. I found <a target="_blank" href="https://www.socketlabs.com/blog/setup-smtp-server/">an article</a> that says hosting a mail server is like building a jet, the article isn't trustworthy since it is from an email-sending company and they have a monetary incentive to say this. I wanted to find some open-source solutions that I can install in 1 step and be done, but there aren't any. Therefore I concluded the article was right and I gave up making my own SMTP server.</p>
<p>I had to find an email-sending service. I looked at different ones, SendInBlue has the largest free-email sending quota, 300 emails per day, which means 9000 emails per month. Of course, this is assuming 300 people sign up per day, which is highly unlikely, the number usually fluctuates. Still, this is the largest quota, so I went with them.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>In this blog post, I talked about the process of creating a Supabase Svelte project. Including challenges I faced when self-hosting Supabase, setting up Svelte, enabling third-party authentication and email sending. I also wanted to talk about routing, SSL, domain name and storing information inside the web browser, but I realized this article already got very long. I've decided I'll write about them in part 2 of this blog post. So stay tuned.</p>
]]></content:encoded></item><item><title><![CDATA[My Plan for building a prayer tracker website]]></title><description><![CDATA[Background
As Christians, we should pray. it is a way for us to communicate with God and develop a deeper relationship with Him. In the Parable of the Unjust Judge, Jesus taught us not to lose heart when praying. But sometimes, we forget to pray, for...]]></description><link>https://blog.cpbprojects.me/my-plan-for-building-a-prayer-tracker-website</link><guid isPermaLink="true">https://blog.cpbprojects.me/my-plan-for-building-a-prayer-tracker-website</guid><category><![CDATA[supabase]]></category><category><![CDATA[Svelte]]></category><category><![CDATA[Web Development]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[planning]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Tue, 13 Dec 2022 20:30:45 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1670939796885/r2B4DHdpQ.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-background">Background</h2>
<p>As Christians, we should pray. it is a way for us to communicate with God and develop a deeper relationship with Him. In the Parable of the Unjust Judge, Jesus taught us not to lose heart when praying. But sometimes, we forget to pray, forget what to pray about, and forget how God answered our prayers. Imagine if our friend forgot to talk to us, or even worse, forgot what we did for them.</p>
<p>That's why we should keep track of what we pray about. It can help us see how God is answering our prayers, as well as help us stay accountable and faithful in our prayer life.</p>
<p>Using a paper journal or a note-taking app to do so is fine, but searching through paper journals is hard, and note-taking apps may not be as convenient as they are not tailor-made for this purpose.</p>
<h2 id="heading-why-i-want-to-make-it">Why I want to make it</h2>
<p>Firstly, I see the need for it. While there are some great options out there, such as <a target="_blank" href="https://www.prayermate.net/app">PrayerMate</a> or the <a target="_blank" href="https://www.youversion.com/the-bible-app/">YouVersion Bible app</a>, they are limited to smartphones only, you cannot access them with a computer. Moreover, they are very complex with many functionalities. It would be nice to have a simpler one. Because</p>
<blockquote>
<p>“Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.”</p>
<p>― Antoine de Saint-Exupéry, Airman's Odyssey</p>
</blockquote>
<p>Secondly, I am interested in trying out different technologies, and this sounds like a cool challenge. I can see this actually being helpful for others. I'm sure I'd learn a lot from doing this.</p>
<p>That's why I've decided to develop it and make it available to the public.</p>
<h2 id="heading-features-of-the-prayer-tracker-website">Features of the prayer tracker website</h2>
<p>I want to build a website where users can:</p>
<ol>
<li><p>Store prayer items, including the ability to add, edit, and delete items.</p>
</li>
<li><p>Give users a daily prayer list based on the frequency set for each prayer</p>
</li>
<li><p>Record how has God answered their prayer, or when the prayer was not granted</p>
</li>
<li><p>Encrypt the prayer items optionally</p>
</li>
</ol>
<p>Since it is a website, we can assume the prayer items will be synced, and there will be account control. I choose a website over an app because a website can be accessed by any device connected to the internet, whereas apps must be downloaded and installed on a specific device.</p>
<h2 id="heading-how-do-i-plan-to-build-it">How do I plan to build it</h2>
<p>I've been following <a target="_blank" href="https://www.youtube.com/@Fireship">Fireship</a> on Youtube, so I've got some technologies on my radar that I want to try out. <a target="_blank" href="https://supabase.com/">Supabase</a> is one of them, it claims to be an open-source Firebase alternative Backend as a Service (BaaS). I'm happy to see that they have a self-host option, which means I wouldn't be charged a large sum overnight, as I am hosting it myself, I can control how many resources to give it.</p>
<p>For the front end, I'm planning to use <a target="_blank" href="https://svelte.dev/">Svelte</a>, simply because it is the javascript framework used in the Supabase to-do list example. I'm assuming it is a good framework for the official example to be built on that. I've also looked up the Fireship video on Svelte, and it looks promising.</p>
<p>For hosting the back end and front end, I'm going to use a VPS provider that limits my CPU and connection speed, but doesn't incur any additional cost nor have a data transmission limit, in my case, <a target="_blank" href="https://billing.oxide.host/aff.php?aff=166">Oxide Host</a>. It is much less stressful knowing that I won't be charged extra even if something went wrong.</p>
<h2 id="heading-what-are-the-other-building-blocks">What are the other building blocks</h2>
<p>I've decided to use Supabase, Svelte and a VPS as the foundation of my project. Now I have to choose the various building blocks that I'm going to use to build my website.</p>
<p>Before that, I want to talk about the two approaches to building something. The first approach is to start doing first and plan as you go, this leads to quicker development if nothing goes wrong, but it is also possible that you'll encounter a problem that sets you back a lot. The second approach is to plan everything before starting, this reduces the risk of finding problems mid-project and having to redo everything, but this slows down the progress, and there is only so much we can plan for, we may still find unexpected problems. I will use the latter approach and test each individual building block before using them to build the website.</p>
<p>At the time of writing this, I've already tested out all the technologies. The process is very difficult, I've been stuck at problems after problems. This process took me several days, and over 24 hours of figuring things out, so I'm going to spare the details and just explain what I'll be using. Later on, I'll write another blog post to talk about how I tested them.</p>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Functionality</td><td>Technology used</td></tr>
</thead>
<tbody>
<tr>
<td>Third-party authentication</td><td>Google and Facebook sign-in, as they are the most popular ones. I refrained from using Apple sign-in because it is locked behind a paywall, you need to pay to join their developer program to use it.</td></tr>
<tr>
<td>Routing</td><td>Nginx, since it is very popular.</td></tr>
<tr>
<td>Https, SSL</td><td>Certbot, since it is free and easy to set up.</td></tr>
<tr>
<td>Domain name</td><td>NameCheap, since it is cheaper than others and offers a 1-year free domain for students.</td></tr>
<tr>
<td>Storing encryption key locally</td><td>Local Storage in the browser, because unlike cookies, they do not get sent back to the server every request.</td></tr>
<tr>
<td>Email Authentication</td><td>MailInBlue, a relatively large email limit in the free tier</td></tr>
</tbody>
</table>
</div><h2 id="heading-conclusion">Conclusion</h2>
<p>This is my plan to build this project. I will make a blog talking about <a target="_blank" href="https://blog.cpbprojects.me/what-i-learned-from-self-hosting-a-supabase-svelte-project-part-1">my experience with testing the above building blocks</a>, as well as a complete tutorial on <a target="_blank" href="https://blog.cpbprojects.me/how-to-self-host-supabase-a-complete-guide">how to set up a Supabase Svelte project</a>.</p>
]]></content:encoded></item><item><title><![CDATA[Python script to find broken links in word document]]></title><description><![CDATA[I wrote a Python script to find broken links in a word document, the GitHub link is here. If you want to use it you can go to the GitHub page and the instructions are there. Below I will explain how it works and how I came up with this solution.
This...]]></description><link>https://blog.cpbprojects.me/python-script-to-find-broken-links-in-word-document</link><guid isPermaLink="true">https://blog.cpbprojects.me/python-script-to-find-broken-links-in-word-document</guid><category><![CDATA[Python]]></category><category><![CDATA[automation]]></category><category><![CDATA[Microsoft Word]]></category><category><![CDATA[Regex]]></category><category><![CDATA[http]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Sun, 11 Dec 2022 22:22:22 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797191710/0SEUkSMRK.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I wrote a Python script to find broken links in a word document, the <a target="_blank" href="https://github.com/chit-uob/findBrokenLink">GitHub link is here</a>. If you want to use it you can go to the GitHub page and the instructions are there. Below I will explain how it works and how I came up with this solution.</p>
<p>This program was inspired by helping a friend. The friend has to click on links inside documents one by one to check if they still work. Hearing that, I thought to myself, that sounds like a task that can be automated. Therefore I asked for some sample word documents and started testing the concepts.</p>
<p>In face of this big task, I decided to break the task into a few steps, figuring out a way to do each, and then piece them together.</p>
<h2 id="heading-step-1-finding-links-from-texts">Step 1: Finding links from texts</h2>
<p>If I want to extract some patterns from text, the first thing that pops into my mind is Regular Expression. They are a way to find patterns in texts and are often deemed difficult and feared by developers. Therefore I did what any sane developer would do: search for this problem online and copied the Regex from stack overflow.</p>
<p>The regex is <code>(https?:\/\/\S+)</code>, which is pretty easy to understand</p>
<div class="hn-table">
<table>
<thead>
<tr>
<td>regex</td><td>meaning</td></tr>
</thead>
<tbody>
<tr>
<td>http</td><td>http</td></tr>
<tr>
<td>s?</td><td>s for zero to one times</td></tr>
<tr>
<td>://</td><td>://, / stands for an escaped /</td></tr>
<tr>
<td>\S</td><td>any non-whitespace character</td></tr>
<tr>
<td>+</td><td>the previous token, \S, from zero to infinite times</td></tr>
</tbody>
</table>
</div><h3 id="heading-problem-1-the-succeeding-full-stop-is-also-matched">Problem 1: the succeeding full-stop is also matched</h3>
<p>Let's say I have a sentence that ends in a link, such as http://chit.hashnode.com. If I parse this text in the above regex, the last full-stop (.) will also be included, because it is a non-whitespace character.</p>
<p>the solution is to use another regex, <code>http[s]?:\/\/[^\s]+[^. ]</code>, here we have a <code>[^. ]</code>, which means: Match a single character not period (.) and not whitespace ( ), so it solves the problem of it catching the trailing period.</p>
<h3 id="heading-problem-2-having-a-newline-character-at-the-end">Problem 2: having a newline character at the end</h3>
<p>The previous regex solves the problem of having a period at the end of the line. But what if there is a newline character after it?</p>
<h4 id="heading-knowledge-dump-what-is-a-newline-character">Knowledge dump: What is a newline character</h4>
<p>The newline character is the invisible character that tells the text editor to go to the next line. For example, let's say <code>\n</code> is the newline character, <code>I am a line.\nI am another line</code> will become</p>
<pre><code class="lang-plaintext">I am a line.
I am another line
</code></pre>
<p>Let's use the website <a target="_blank" href="https://regex101.com/">regex101</a> to test it, when we have <code>A sentence ending in https://google.com.</code>, the <code>https://google.com</code> will be correctly captured.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797217032/v-ZLRF1Yo.png" alt class="image--center mx-auto" /></p>
<p>But with a newline character after it, it will be caught as well, since the newline character was not excluded in <code>[^. ]</code>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797231531/hqGh9X9Q9.png" alt class="image--center mx-auto" /></p>
<p>As I am writing this blog post, I figured the solution was to exclude the newline character as well, resulting in the regex <code>http[s]?:\/\/[^\s]+[^. ]</code>, but I was not this clear-minded when I was programming, the solution I came up with is to replace all the newline characters with the space bar. Then do the regex matching, which also worked, because there are no longer newline characters.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797235951/8_Aw5iBPR.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-problem-3-different-types-of-links">Problem 3: Different types of links</h3>
<p>The third problem is with links with no <code>http</code>, our regex only matches links starting with <code>http</code>, but some links start with <code>www</code>, or maybe even no <code>www</code>.</p>
<p>I am sure more clever regexs can accommodate this, but I didn't want to spend too much time at this stage, so I just used a python library to do the work for me. That's the beauty/problem with Python, there are so many libraries that you can just use one, and not care how it is implemented.</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> urlextract <span class="hljs-keyword">import</span> URLExtract
urlextracter = URLExtract()
urls = urlextracter.gen_urls(s)
</code></pre>
<h2 id="heading-step-2-check-the-links">Step 2 Check the links</h2>
<p>Now that we have all the URLs from a text extracted, we want to check them.</p>
<h3 id="heading-problem-4-links-without-a-protocol-specified">Problem 4: links without a protocol specified</h3>
<p>But remember how some links don't start with HTTP? I need to add an HTTP in front of them, or else the python library request will have a hard time knowing what protocol is required, so I used the following code to achieve that</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797247078/Js5jLjEUC.png" alt class="image--center mx-auto" /></p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> re
<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">formaturl</span>(<span class="hljs-params">url</span>):</span>
    <span class="hljs-keyword">if</span> <span class="hljs-keyword">not</span> re.match(<span class="hljs-string">'(?:http|ftp|https)://'</span>, url):
        <span class="hljs-keyword">return</span> <span class="hljs-string">'http://{}'</span>.format(url)
    <span class="hljs-keyword">return</span> url

urls = [formaturl(url) <span class="hljs-keyword">for</span> url <span class="hljs-keyword">in</span> urls]
</code></pre>
<p>here I try to match the regex <code>(?:http|ftp|https)://</code>, it tests if the URL starts with http/ftp/https, if not, we append <code>http://</code> in front of it</p>
<p>then we use list comprehension to do it on the entire url list.</p>
<h3 id="heading-actually-checking-the-link">Actually checking the link</h3>
<p>Now we get to the step of actually checking the link, we do that using the <code>requests</code> library. We wrap the requests.get() inside a try-catch block so that if other issues happen, the program will not crash, it will simply return false. and if the status_code of the response isn't 200, then we return false too, else we return true.</p>
<p>200 means all good, so if the website returns standard content, it will return all good.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">check_link</span>(<span class="hljs-params">url</span>):</span>
    print(<span class="hljs-string">f"checking <span class="hljs-subst">{url}</span>"</span>)

    <span class="hljs-comment"># Try and see if url have inherit problem</span>
    <span class="hljs-keyword">try</span>:
        response = requests.get(url)
    <span class="hljs-keyword">except</span>:
        <span class="hljs-keyword">return</span> <span class="hljs-literal">False</span>

    <span class="hljs-comment"># See if not 200</span>
    <span class="hljs-keyword">if</span> <span class="hljs-keyword">not</span> response.status_code == <span class="hljs-number">200</span>:
        <span class="hljs-keyword">return</span> <span class="hljs-literal">False</span>

    <span class="hljs-keyword">return</span> <span class="hljs-literal">True</span>
</code></pre>
<h2 id="heading-step-3-reading-word-document-text">Step 3: Reading word document text</h2>
<p>Now I have a function that extracts URLs, and another to check if the url's website is working. I have to extra the text from a word document.</p>
<p>To do that, I can simply use the save-as function inside word and call it a day. That would save the document in a text format and we can read the text file with the program and get the result.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797253134/2LIkAqdom.png" alt class="image--center mx-auto" /></p>
<p>But that would mean the user has to do more, so I was thinking, is it possible to read the word document directly with Python?</p>
<p>The answer is yes because actually, word documents are just zip files. To know if I'm telling you the truth, install <a target="_blank" href="https://www.7-zip.org/">7zip</a> and unzip a word document, and you will see the following result. A word document is just a zip file containing a lot of XML files.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1670797261175/dwdqv89WD.png" alt class="image--center mx-auto" /></p>
<p>A way to solve this problem is to read the word document as a zip file and find the links inside the XML files. But since I'm using Python, I decided to find if there are any libraries that can do this for me. I found <code>docx</code>, <code>docx2txt</code>, <code>docx2python</code>. After testing them all, the one I decided to use is <code>docx2python</code>, because it extracts the main text, headers, footers, and even footnotes. The syntax is as below:</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> docx2python <span class="hljs-keyword">import</span> docx2python

<span class="hljs-comment"># extract docx content</span>
<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">get_text_by_docx2python</span>(<span class="hljs-params">path</span>):</span>
    text = docx2python(path).text
    <span class="hljs-keyword">return</span> text
</code></pre>
<h2 id="heading-step-4-piecing-them-all-together">Step 4 Piecing them all together</h2>
<p>Now that all parts of the puzzle are here, it is time to implement the main logic, we first extract the text by <code>use_docx2python.get_text_by_docx2python</code>, then <code>urlextracter.gen_urls</code> to get the URLs, after adding <code>http</code> to them, we use <code>check_links.check_link</code> to check the link, and print a helpful message.</p>
<h3 id="heading-problem-some-links-show-a-different-result-when-we-request-it-from-the-program-than-when-we-actually-click-the-site">Problem: some links show a different result when we request it from the program than when we actually click the site</h3>
<p>This is because some sites do not like robots, but since Python requests have the default user agent as <code>python-requests/2.25</code>, the website sees this and forbids the program from getting the result. What we need to do to fix it is to use another user agent, I copied the user agent on my web browser, <code>Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/104.0.0.0 Safari/537.36</code>, then I used <code>response = requests.get(url, headers=HEADERS, stream=True)</code>, now the website responds correctly.</p>
<h2 id="heading-step-5-doing-it-asynchronously">Step 5 Doing it asynchronously</h2>
<p>Now waiting for the program to do all the things is fine, but it is also too slow. The slowest part is waiting for the websites to be fetched. Since we have to wait for each website to give us the result. This would be faster if we use concurrency, which means doing things at the same time.</p>
<p>Think of it as, burger shop A gives you a burger in 5 minutes, and Boba shop B gives you a boba in 3 minutes. You can get a burger and then a boba in 8 minutes. If you order both at the same time and then collect them when they are ready, you'll only need 5 minutes.</p>
<p>Concurrency code is a bit more complex so I'm not going to explain it here, you can visit <a target="_blank" href="https://realpython.com/python-concurrency/">RealPython</a> on this topic, they have a great tutorial on this topic.</p>
<h2 id="heading-step-6-warp-up">Step 6 Warp up</h2>
<p>At last, I added a code to show a file dialogue to further simplify this for users.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> tkinter <span class="hljs-keyword">as</span> tk
<span class="hljs-keyword">from</span> tkinter <span class="hljs-keyword">import</span> filedialog
file_path = filedialog.askopenfilename()
</code></pre>
<h2 id="heading-conclusion">Conclusion</h2>
<p>That's how I did it and I hope you learned something from it.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [Conclusion] What I learned from Voi scooter data this summer]]></title><description><![CDATA[This blog post concludes what I learned from the data collected in the Where's my Voi scooter series. I started this study at the beginning of the summer aiming to find ways to locate a Voi scooter by collecting data on Voi scooters. I tracked the sc...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-conclusion-what-i-learned-from-voi-scooter-data-this-summer</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-conclusion-what-i-learned-from-voi-scooter-data-this-summer</guid><category><![CDATA[Voi]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Python]]></category><category><![CDATA[Data Science]]></category><category><![CDATA[data analysis]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Tue, 20 Sep 2022 09:48:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1663667231908/U1uWHWlH1.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This blog post concludes what I learned from the data collected in the <a target="_blank" href="https://chit.hashnode.dev/series/where-is-my-voi-scooter">Where's my Voi scooter series</a>. I started this study at the beginning of the summer aiming to find ways to locate a Voi scooter by collecting data on Voi scooters. I tracked the scooter location and the battery level of scooters in my city at one-minute intervals using the Application Program Interface (<a target="_blank" href="https://aws.amazon.com/what-is/api/">API</a>) on their mobile app. Below is a summary of what I have learned from the study.</p>
<h2 id="heading-mean-battery-level-of-scooters">Mean battery level of scooters</h2>
<h3 id="heading-graph-for-the-entire-duration">Graph for the entire duration</h3>
<p>This is the mean battery level of all scooters around the city during the entire duration of the project. The graphs may look small, you can right-click the image and click <code>open image in new tab</code> to enlarge them.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663335774138/2M_uyA5tQ.png" alt="1.png" /></p>
<p>In general, the mean battery level fluctuated between 55% and 60%.</p>
<p>We can observe that before 7 July, there were flat lines every day. This was because back then scooters went offline at 10 pm, and online at 6 am. During that time my program could not collect data and it assumed the battery level didn't change. After 7 July, you can see that the transitions in the graph were smooth again with no jumps.</p>
<p>A scooter "goes offline" when the API no longer returns information about the scooter. The API that I used is the API that tells users where the idle scooters are. So when someone is using the scooter, when the battery of the scooter is being charged, or when the scooter is sent back to the warehouse, the scooter is considered "offline".</p>
<p>We can also see errors in the data collection. There were two sloped straight lines during June, as back then my program would crash if the API server didn’t return information. This was fixed afterwards by adding a try-catch and retrying after failures. The second error was on 1 July, the graph broke into half. This is because I calculated the mean battery level each month, and the program guesses the battery levels of offline scooters by forward filling, ie. if the scooter went offline with 70% battery, the program assumes it would stay at 70% while being offline. When the new month started, the scooters with unknown battery levels were not included in the calculation, so there was a jump in the value. The reason I used forward filling despite this disadvantage is given in <a target="_blank" href="https://chit.hashnode.dev/wheres-my-voi-scooter-7-analysis-of-scooter-battery-with-graphs">the 7th blog in this series</a>.</p>
<p>We can also see a sharp decline in battery levels around 7 July. This might be because scooters no longer went offline at night and the scooter service area was expanded, so the battery drained even during the night, and with a wider area, it was more difficult to reach them and replace the battery of them.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663337429768/PW0XSt1lV.png" alt="2.png" /></p>
<h3 id="heading-graph-for-weeks">Graph for weeks</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663337919371/ETDIxOF8M.png" alt="3.png" /></p>
<p>These graphs were weeks starting from 10 July, they all started on Sundays. There was a general trend that the battery level was higher in the second half of the week.</p>
<h3 id="heading-graph-for-days">Graph for days</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663338279692/e23KK63Ql.png" alt="4.png" /></p>
<p>These graphs were some days, they all started at around 0 am. There was a general trend that the battery level would increase in the first half of the day, usually, at 10 am, then start decreasing. This was strange, an increase in battery level should mean someone was swapping out the batteries, or new scooters with a full battery was being introduced. But it was hard to imagine people swapping batteries in the middle of the night.</p>
<h2 id="heading-available-scooter-count">Available scooter count</h2>
<h3 id="heading-graph-for-the-entire-duration">Graph for the entire duration</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663345088354/b9EPikZP5.png" alt="5.png" /></p>
<p>Same as the mean battery level, the days before 7 July had the count drop to zero at night because they were offline. And the number of available scooters also sharply declined from around 1150 scooters to under 500 scooters on 11 July. Presumably, it was because of the change in operation too.</p>
<p>We can also observe that more scooters were put into operation from 26 July, I believed it was due to the Commonwealth Games 2022 on 28 Jul 2022 – 8 Aug 2022. When the Commonwealth Games started, the number of available scooters, along with the mean scooter battery level dropped. It should be because of the increased ridership. This can be confirmed on <a target="_blank" href="https://www.intelligenttransport.com/transport-news/138525/voi-ridership-2022-commonwealth-games/">the news</a> where it said Voi sees record e-scooter ridership during 2022 Commonwealth Games at over 66,000 journeys.</p>
<p>Then the count returned to normal until 16 Aug but the count increased again on 19 Aug when the number of scooters exceeded 1750 scooters.</p>
<h3 id="heading-graph-for-weeks">Graph for weeks</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663346756796/H4cBJ5Nd4.png" alt="6.png" /></p>
<p>These are graphs of weeks starting from Sunday. There seemed to be a slight trend where the number of scooters in the second half of the week was higher.</p>
<h3 id="heading-graph-for-days">Graph for days</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663347012233/PLwDWH4JL.png" alt="7.png" /></p>
<p>These are graphs of days starting at 0 am, usually, the number of scooters increases to about the middle of the day, then decreases.</p>
<h2 id="heading-ridership">Ridership</h2>
<h3 id="heading-total-graph">Total graph</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663534469329/gxJbKZYZN.png" alt="8.png" /></p>
<p>This graph corresponded positively to the available scooter count graph, when more scooters became available, more rides happened too. In July, there are fewer rides, but in July and August, the number of rides increased.</p>
<p>At the beginning of the graph, there were two places where the count hit close to zero, that was because of the error of my data collection program. My program crashed during that two times so I wasn't able to recognise rides during that two periods.</p>
<h3 id="heading-commonwealth-games">Commonwealth games</h3>
<p>With this data, we can answer the question: did ridership increase during the Commonwealth Games 2022? The answer is it did. From 28 Jul 2022 to 8 Aug 2022, the ridership peaked at around 6400 rides per day, but the increase didn't stop there, it seemed even after the Commonwealth Games and the increase in scooter count and coverage, the increased ridership was maintained.</p>
<p>The sum of the rides during the Commonwealth Games calculated in my program is 65951, which was super close to 66000 said in <a target="_blank" href="https://www.intelligenttransport.com/transport-news/138525/voi-ridership-2022-commonwealth-games/">this article</a>. There was a discrepancy in the number due to the logic of my program. A ride was counted in my program when a scooter became offline and then online again within 45 minutes. But it did this at a one-minute interval. So if a rider parked a scooter and another rider rented it within a minute, my program would count it as one ride when there were supposed to be two. Another problem was that when a scooter temporarily went offline for a battery swap, my program still considered it a ride since the scooter went offline and then online.</p>
<p>I limited the maximum number of minutes in a ride because scooters might be taken to the warehouse, fixed, and re-released, so I could not count all instances of scooters being offline and then online again as a ride, so I chose 45 minutes as it was the maximum riding duration for a Voi pass.</p>
<h3 id="heading-per-week">Per week</h3>
<p>Now I want to answer the question, does the ridership correspond with weekdays?</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663536657898/TsWtslP_-.png" alt="9.png" /></p>
<p>These weeks were the same as the ones with available scooter count. They generally corresponded with each other, days with more available scooters generally result in more ridership. There was no clear pattern in which weekdays had greater ridership.</p>
<h3 id="heading-per-day">Per day</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663537403311/8Jt15q5tN.png" alt="10.png" /></p>
<p>The same days were chosen as the ones in the available scooter count graphs. These graphs were generally the opposite of those graphs, the times with the least available scooter were the times with most rides, which makes sense because people were using scooters so there are fewer of them available.</p>
<p>We can also see on most days the ride count was minimum in the morning around 5 am, then increased to peak at 8 am, then fell back again, then increased again from 12 pm to around 8 pm, then slowly decreased. I think the first peak was when people get to work, and the second peak was for people leaving work.</p>
<h2 id="heading-heatmap">Heatmap</h2>
<h3 id="heading-day">day</h3>
<p>I figured it made little sense to measure this over weeks or the entire period, so I decided to only measure this over a day.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663586408033/OMPBdRRmF.gif" alt="11.gif" /></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663586564548/SGwHBl6dt.gif" alt="12.gif" /></p>
<p>The scooters generally became more concentrated on certain spots on the map from 0 to 12 pm, then became more spread out after that.</p>
<h2 id="heading-battery-consumption">Battery consumption</h2>
<h3 id="heading-during-ride">During ride</h3>
<p>The way I calculated this was the same as ridership, so this analysis suffered from the same problems. However, I think these results would still be very close to the true value, as the calculation only fails in a few edge cases. The average level of battery drop was 0.461% per minute when unlocked. The battery consumption varied in each ride since scooters could be idle while unlocked, or be used in battery-consuming activities such as going up slopes.</p>
<p>I also filtered the battery consumption based on which month the data was collected, and found a surprising result. The battery consumption was decreasing. This is strange since I was taking on average all rides, so there shouldn't be much difference. Some possible explanation for this phenomenon was a change in rider behaviour, a change in battery capacity, or Voi figuring out a way to use less battery on its scooters.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663611599390/-ZPG2pLei.png" alt="13.png" /></p>
<p>I also filtered the results based on the starting battery level, and there are some interesting observations too. The results for the first half make sense, it is the same for our phones, the battery generally drains slower at a higher battery level. However, the battery drain decreased when it is below 40%, which is strange, usually, the battery drains the quickest when it is closest to zero.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663612238244/tSV9pgRt8.png" alt="14.png" /></p>
<h3 id="heading-when-idle">When idle</h3>
<p>I also wanted to know what is the battery draining speed when the scooter is sitting there idle. The average battery drained, while the scooter was idle, was 0.006830% per minute.</p>
<p>Starting at different battery levels, the draining speed was different too, the less battery there was, the faster the battery drains, this was what I had expected.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663614258451/mKSKEvBea.png" alt="15.png" /></p>
<p>In my experience, lithium batteries, used in smartphones, usually drain faster when the battery level is lower. I wasn't able to pin down why, but I saw some explanations on <a target="_blank" href="https://www.reddit.com/r/askscience/comments/buttbd/does_each_percent_of_your_phone_battery_last_the/">the</a> <a target="_blank" href="https://www.quora.com/Why-does-my-phone-battery-drop-from-20-to-1-in-about-1-minute-but-then-will-rest-on-1-for-a-good-amount-of-time">internet</a> what I would like to share, one has to do with the lithium discharge voltage curve, in a lower stage of charge the voltage decreases and therefore the current has to be increased to maintain the power output, therefore the battery drains quicker. But I wasn't able to find solid proof of this, so if you know more, please let me know in the comments down below.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1663613962183/mozSUFjcX.png" alt="16.png" /></p>
<p>This brings us to another question, why during rides, the battery drained on 0-20% wasn’t the quickest, but during idle, the battery drained on 0-20% was the quickest? It could be an error in my data processing. It could also be that Voi's display battery level was not the actual battery level, it displayed a lower level when the battery was near zero so that it could travel further during the supposed near zero battery, to prevent the scooter from running out of battery mid-ride, but this is just a theory. </p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>This project started out with the goal of locating a scooter in the city, which I am now able to do with the data I collected. But I found greater interest in studying the broader trend of scooters around my city. So below are my findings:</p>
<p>I can conclude that the scooter business is doing well. The ridership increased during the Commonwealth Games 2022 and the usage didn't drop back down after the games. The batteries are being replaced in time so their mean battery level and available scooter count are stable. The scooter's battery drain when idle is about 65 times slower than when on a ride, so leaving scooters on the street isn't costing them a lot.</p>
<p>Other than what I learned from the data, I also learned a lot about data analysis in the project, like how to collect and organize data, how to analyse them, and some handy Python modules to make it easier.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [10] Analysis of data in July 2022]]></title><description><![CDATA[In this blog, I will analyse the data of Voi scooters in my area in July using the Python program I built previously. All the messy code will not be shown here, but you can see the beautiful plots and graphs generated from the data. If you are intere...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-10-analysis-of-data-in-july-2022</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-10-analysis-of-data-in-july-2022</guid><category><![CDATA[Voi]]></category><category><![CDATA[2Articles1Week]]></category><category><![CDATA[Python]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Data Science]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Wed, 07 Sep 2022 11:03:11 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1662548521650/noBibX9Hw.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog, I will analyse the data of Voi scooters in my area in July using the Python program I built previously. All the messy code will not be shown here, but you can see the beautiful plots and graphs generated from the data. If you are interested in the code behind this, you can check out my <a target="_blank" href="https://chit.hashnode.dev/series/where-is-my-voi-scooter">Where's my Voi scooter series</a>.</p>
<h2 id="heading-downloading-the-data">Downloading the data</h2>
<p>I have been using a VPS (virtual private server) to collect the data regularly and automatically. Therefore I have to first download the data from it using a tool called <a target="_blank" href="https://termius.com/">Termius</a>.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118562471/9vRTXYyBo.png" alt="1.png" /></p>
<p>I selected all the data items from the start of July and copied them to my local machine. There are about 2GB of data.</p>
<h3 id="heading-recap-of-what-data-i-collect">Recap of what data I collect</h3>
<p>I collect the data of all Voi scooters around my city every minute by using the API provided by the Voi mobile app. The data includes each available scooter's id, battery level and location. I modify it into a JSON list. Then I save the result to a JSON file every hour. Therefore the data I have looks something like this:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118568024/vahv3V2yp.png" alt="2.png" /></p>
<h2 id="heading-importing-the-data-and-plotting-the-graph">Importing the data and plotting the graph</h2>
<p>All data files are in the same folder, so I generate a list of all file paths inside the folder and sort them in ascending order. Then I sanitise the data by trying to parse it using Python's try-catch. If the data file is correct, Python would not catch any error, so if Python caught an error, I know I have to fix the data file. Once that's done, I read all the files in ascending order and stored the time stamp and vehicle count in two separate numpy arrays. I then plot it using matplotlib. The graph looks something like this:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118572386/8iSGGlzNn.png" alt="3.png" /></p>
<p>We can observe:</p>
<ol>
<li>Before 7/7, the scooter count always goes to zero at the end of the day. Because scooters are disabled at night, after that, Voi decided that scooters would be available 24/7, therefore the scooter counts never dropped to zero again.</li>
<li>Near 12/7, the scooter count dropped to very low, that is just after the change, so maybe Voi did some changes in its workforce so that fewer workers are replacing batteries on scooters, or they are doing some checking on scooters so a lot is brought back for checking. Either way, the scooter count increased and returned to normal.</li>
<li>After 30/7, the number of available scooters seems to plummet, I suspect that is because of the Commonwealth games.</li>
</ol>
<h2 id="heading-finding-the-set-of-all-available-scooters">Finding the set of all available scooters</h2>
<p>Then I run through all files, running the set union function from Python, to find the set of all scooters available throughout July. I do that by the list of all available scooters at each time stamp to a set, and take the union of the set of all scooters registered. So in the end, the union of all sets is the total amount of available scooters. </p>
<p>After running that code, I found out there were 2544 scooters of unique id that were available at a point in July.</p>
<h2 id="heading-get-all-dates-list">Get all dates list</h2>
<p>What I want is a gigantic pandas table with scooter ids as columns, and time stamps as rows. I already know the ids, so I have to find the time stamps, I do so by iterating through all the data items and adding the time stamps to a list.</p>
<h2 id="heading-fill-all-the-tables">Fill all the tables</h2>
<p>I run the code to fill all the tables, it does this by iterating through all the data files, in each individual data item, and adding them to the corresponding pandas table. This code is extremely slow because we have 44499 data of individual time stamps. And in each data, there are 1000 scooters, and for each scooter, we are storing their latitude, longitude and battery level, so we are adding 44499 <em> 1000 </em> 3 items to tables, which is over one hundred million. The code wasn't able to finish in 2 hours, therefore I had to stop it, and continue the code from where I left off the other day.</p>
<h2 id="heading-plotting-individual-scooters">Plotting individual scooters</h2>
<p>Some scooters are more active, like</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118580013/Ob_b7kEo3.png" alt="4.png" /></p>
<p>Some are decommissioned for some time, maybe for fixing, or just recycling scooters. </p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118588230/WDU1Krs_6.png" alt="5.png" /></p>
<p>But keep in mind that we only know the id of the scooter, the same scooter could have its id switched, we don't know what the id means.</p>
<h2 id="heading-plotting-the-mean-battery-count">Plotting the mean battery count</h2>
<p>now I wish to plot the mean scooter battery count. I do so by first dropping columns that were all null values. then I use the method of forwarding filling to fill in null values. Then at last I plot the graph using the plot function and the mean function with axis = 1. there are null values in my code reasons: when a scooter is rented, it is no longer available, so the value becomes null. when the scooter is called back for fixing or other purposes, it is not available. Before scooters are available all day, at night, all scooters are not available.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1662118594549/9K5JhsETO.png" alt="6.png" /></p>
<h2 id="heading-plotting-location-heatmap">Plotting location heatmap</h2>
<p>Then I generate the heatmaps using seaborn and matplotlib, I generate one for each hour, then piece them together as a gif file. The negative of having it as a gif file is that a gif file generally has a bigger file size as it compresses images individually, while mp4 is capable of storing the change between images, hence if the images in the video are similar, a lot of storage space can be saved. I convert the gif to mp4 using the answer suggested by <a target="_blank" href="https://stackoverflow.com/a/40726572/16929051">this StackOverflow answer</a> by using moviepy. and the result is as below.</p>
<iframe width="728" height="410" src="https://www.youtube.com/embed/EljqGsqNzd4"></iframe>

<h2 id="heading-summary-and-whats-next">Summary and what's next</h2>
<p>There are a lot of interesting observations that can be made using these data. But for the length of this blog, I am just going to show the data collected and fewer observations. Again if you are interested in how the data was collected, you can check out <a target="_blank" href="https://chit.hashnode.dev/series/where-is-my-voi-scooter">Where's my Voi scooter series</a>. </p>
<p>I think there are many conclusions to be drawn, such as where will scooters be concentrated at what time, and around what time most scooters are available.</p>
]]></content:encoded></item><item><title><![CDATA[Python program running at regular time interval tutorial using time and datetime module]]></title><description><![CDATA[This tutorial will teach you how to set up a Python program to run at regular time intervals using the time and datetime module. This is useful if you want to check something constantly, so you have a python program running in the background doing th...]]></description><link>https://blog.cpbprojects.me/python-program-running-at-regular-time-interval-tutorial-using-time-and-datetime-module</link><guid isPermaLink="true">https://blog.cpbprojects.me/python-program-running-at-regular-time-interval-tutorial-using-time-and-datetime-module</guid><category><![CDATA[Python]]></category><category><![CDATA[time]]></category><category><![CDATA[Tutorial]]></category><category><![CDATA[2Articles1Week]]></category><category><![CDATA[Programming Blogs]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Wed, 24 Aug 2022 10:57:28 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1661338292959/aD44Gm9U1.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This tutorial will teach you how to set up a Python program to run at regular time intervals using the time and datetime module. This is useful if you want to check something constantly, so you have a python program running in the background doing the checking for you at regular intervals.</p>
<h2 id="heading-solution">Solution</h2>
<p>I will put the solution on top. If you are interested in how I found it, you can read about it below.</p>
<h3 id="heading-simple-solution-for-if-what-you-do-is-lightweight">Simple solution for if what you do is lightweight</h3>
<p>Lightweight means that it doesn't use much CPU power and time, like a simple print function.</p>
<p>Assuming you want to run the code every 10 seconds.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> time
<span class="hljs-keyword">from</span> datetime <span class="hljs-keyword">import</span> datetime

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">do_things</span>():</span>
    print(datetime.now())

<span class="hljs-keyword">while</span> <span class="hljs-literal">True</span>:
    do_things()
    time.sleep(<span class="hljs-number">10</span>)
</code></pre>
<p>you should get something like this</p>
<pre><code><span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">10</span>:<span class="hljs-number">59</span>:<span class="hljs-number">35</span>.<span class="hljs-number">247516</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">10</span>:<span class="hljs-number">59</span>:<span class="hljs-number">45</span>.<span class="hljs-number">259224</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">10</span>:<span class="hljs-number">59</span>:<span class="hljs-number">55</span>.<span class="hljs-number">272715</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">00</span>:<span class="hljs-number">05</span>.<span class="hljs-number">284097</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">00</span>:<span class="hljs-number">15</span>.<span class="hljs-number">290823</span>
</code></pre><p>There you have it, you can replace the sleep time if you want a different interval, and you can replace the content of the do_things() function with what you want the program to do.</p>
<h4 id="heading-problem-with-the-simple-solution">Problem with the simple solution</h4>
<p>But if you are like me, what you wish to do every interval takes time, whether it is waiting for I/O from another program, or it is doing something very CPU intensive. Then if we still do what we did, we will get inaccuracies in wait time like the below.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> time
<span class="hljs-keyword">from</span> datetime <span class="hljs-keyword">import</span> datetime

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">do_things</span>():</span>
    number = <span class="hljs-number">50</span>_000_000
    sum(i*i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> range(number))
    print(datetime.now())

<span class="hljs-keyword">while</span> <span class="hljs-literal">True</span>:
    do_things()
    time.sleep(<span class="hljs-number">10</span>)
</code></pre>
<p>Here I used <code>sum(i*i for i in range(number))</code> on a huge number to force the CPU to do a lot of work, don't mind the underscore in 50_000_000, it is the same as 50000000. The result of this program is:</p>
<pre><code><span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">07</span>:<span class="hljs-number">54</span>.<span class="hljs-number">377208</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">08</span>:<span class="hljs-number">07</span>.<span class="hljs-number">599538</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">08</span>:<span class="hljs-number">20</span>.<span class="hljs-number">838108</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">08</span>:<span class="hljs-number">34</span>.<span class="hljs-number">043164</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">08</span>:<span class="hljs-number">47</span>.<span class="hljs-number">255688</span>
</code></pre><p>As you can see, we want the program to execute every 10 seconds, but it took the program about 14 seconds each loop, this is because some extra time is spent computing. This is why you need the advanced solution.</p>
<h3 id="heading-advanced-solution">Advanced solution</h3>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> time
<span class="hljs-keyword">from</span> datetime <span class="hljs-keyword">import</span> datetime, timedelta

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">do_things</span>():</span>
    number = <span class="hljs-number">50</span>_000_000
    sum(i*i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> range(number))
    print(datetime.now())

<span class="hljs-keyword">while</span> <span class="hljs-literal">True</span>:
    previous_datetime = datetime.now()
    do_things()
    time_took_running_do_things = datetime.now() - previous_datetime
    remaining_time_in_secs = (timedelta(seconds=<span class="hljs-number">10</span>) - time_took_running_do_things).total_seconds()
    time.sleep(remaining_time_in_secs)
</code></pre>
<p>and the result is:</p>
<pre><code><span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">14</span>:<span class="hljs-number">40</span>.<span class="hljs-number">821099</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">14</span>:<span class="hljs-number">50</span>.<span class="hljs-number">797054</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">15</span>:<span class="hljs-number">00</span>.<span class="hljs-number">806950</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">15</span>:<span class="hljs-number">10</span>.<span class="hljs-number">834795</span>
<span class="hljs-attribute">2022</span>-<span class="hljs-number">08</span>-<span class="hljs-number">24</span> <span class="hljs-number">11</span>:<span class="hljs-number">15</span>:<span class="hljs-number">20</span>.<span class="hljs-number">845102</span>
</code></pre><h4 id="heading-explanation">Explanation</h4>
<p>Here you can see that the codes run every 10 seconds even though we are doing the heavy lifting. This is because we compensated for the time taken to run do_things() by waiting for less afterwards. If it took the program 3 seconds to run do_things(), we only wait 7 seconds afterwards, so the program still do things every 10 seconds.</p>
<h4 id="heading-limitation">Limitation</h4>
<p>This solution is limited by, the action you want to do must to take longer than the interval you want to loop. If you want to do an operation that takes 10 seconds every 1 second, this solution cannot help you. You might want to look into <a target="_blank" href="https://realpython.com/python-concurrency/#how-to-speed-up-a-cpu-bound-program">Concurrency from this RealPython article</a>.</p>
<h2 id="heading-how-i-got-there">How I got there</h2>
<p>Now I will talk about how I figured this out.</p>
<p>I am working on<a target="_blank" href="https://github.com/chit-uob/usageTracker"> a program that sees if I'm playing computer games at regular intervals to remind myself how long I'm playing</a>, but the problem is that it is very inaccurate. I would play a game for 45 minutes, then the program will tell me that I only played for 30 minutes.</p>
<p>I first thought this is a problem with time.sleep(), maybe because the CPU is doing a lot of work while gaming, the wait is wrong.</p>
<p>But this is proven wrong by when I'm not playing games, the program still runs at incorrect intervals.</p>
<p>Then I commented out the things the program does within each interval, only leaving the print time statement, and then the problem was fixed. It turns out the problem is what I do inside the interval is costing above 2 seconds!</p>
<p>I then immediately went on to try to over-engineer the problem. I thought to myself, maybe I want the loop and the do_things() function in separate threads via threading or asyncio, or maybe even separate CPU using multiprocessing. I watched the entire tutorial on <a target="_blank" href="https://realpython.com/python-concurrency/#how-to-speed-up-a-cpu-bound-program">Speed Up Your Python Program With Concurrency by Jim Anderson on RealPython</a> until I drew this graph.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1661337810919/qmXSDdtnU.png" alt="graph.png" /></p>
<p>I then realize, what if I just do them all linearly, then compute how much time I need to wait? It is then that it dawned on me, that the problem was never this complicated. So I fixed the problem and decided to write this article to help others like me. So there you have it.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [9] Masking scooter heatmap on the map using Python PIL module]]></title><description><![CDATA[In this blog, I plan to incorporate the map into the heatmap made in the previous blog post, because, with only the heatmap, you can see which coordinates have a lot of scooters, but you have to reference the map with the coordinates, which is inconv...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-9-masking-scooter-heatmap-on-the-map-using-python-pil-module</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-9-masking-scooter-heatmap-on-the-map-using-python-pil-module</guid><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Voi]]></category><category><![CDATA[Python]]></category><category><![CDATA[image processing]]></category><category><![CDATA[2Articles1Week]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Wed, 03 Aug 2022 22:00:41 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563934511/0GZxnmXkI.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog, I plan to incorporate the map into the heatmap made in the <a target="_blank" href="https://chit.hashnode.dev/wheres-my-voi-scooter-8-creating-heatmap-for-distribution-of-scooters">previous blog post</a>, because, with only the heatmap, you can see which coordinates have a lot of scooters, but you have to reference the map with the coordinates, which is inconvenient. So I want to put the information on the Heatmap directly on the map so that the information is shown clearly.</p>
<h2 id="heading-heatmap">Heatmap</h2>
<h3 id="heading-overlapping-image">Overlapping image</h3>
<p>To overlap the image of the map and the heatmap, I need to find out how to have a half-transparent image on top of another image.</p>
<p>I first have to make an image transparent, I do that by using the <code>putalpha(128)</code> to make an image half transparent. It works by putting an alpha, or transparency level of 128 on that image. An alpha level of 0 means completely transparent, and level 255 means completely opaque. So 128 is in the middle, hence half-transparent.</p>
<p>Then I tried using <code>image.paste(another_image)</code> to put the half-transparent image on top of the other. but what happened is the image complete overwrites the other. what I needed was <code>alpha_composite</code>. For testing, I alpha composited the half-transparent map on top of the sat.</p>
<p>I created this <code>map_img</code> by getting the four coordinates I used for the range of the heatmap. and Taking a screenshot with exactly the four corners.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563154225/M_BzlDdL2.png" alt="1.png" /></p>
<p>sat_map
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563160588/uNQMG9P1O.png" alt="2.png" /></p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> PIL <span class="hljs-keyword">import</span> Image
map_img = Image.open(<span class="hljs-string">'img/map.png'</span>)
map_img.putalpha(<span class="hljs-number">128</span>)
sat_map = Image.open(<span class="hljs-string">'img/sat_map.png'</span>)
sat_map.alpha_composite(map_img)
sat_map.save(<span class="hljs-string">'img/result.png'</span>)
</code></pre>
<p>result.png
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563168516/ecPnxvMXK.png" alt="3.png" /></p>
<h3 id="heading-masking">Masking</h3>
<p>Now I want to mask the heatmap image on the map image so that we can see the areas with a lot of scooters on the map. I wanted to use alpha composite like the last example, but I figured that using masking would be better. Image masking is putting image A on image B using image C as a guideline. I plan to put a dark image over the map, using the heatmap as a guideline. Places with more scooters will be darker so that we can tell where the scooters are on the map.</p>
<p>The function to do the masking is <code>Image.composite()</code>, it requires all images to have the same size. So I plan to use the size width 600 height 600 for all. I first make a background image of all white, resize the map image to the size of the heatmap using the <code>thumbnail</code> function, and paste it into the background. then I read the heatmap image and fix its size and convert it to <code>L</code>, which means greyscale for some reason. then I make a red image of all red, and I composite all of them together.</p>
<pre><code class="lang-python">bg_img = Image.new(<span class="hljs-string">"RGBA"</span>, (<span class="hljs-number">600</span>, <span class="hljs-number">600</span>), <span class="hljs-number">0</span>)
map_img = Image.open(<span class="hljs-string">'img/map.png'</span>)
map_img.thumbnail((<span class="hljs-number">474</span>, <span class="hljs-number">546</span>), Image.ANTIALIAS)
bg_img.paste(map_img, (<span class="hljs-number">36</span>, <span class="hljs-number">21</span>))
heatmap_img = Image.open(<span class="hljs-string">'img/heatmap.png'</span>).resize((<span class="hljs-number">600</span>, <span class="hljs-number">600</span>)).convert(<span class="hljs-string">"L"</span>)
bk_img = Image.new(<span class="hljs-string">"RGB"</span>, (<span class="hljs-number">600</span>, <span class="hljs-number">600</span>), <span class="hljs-number">255</span>).convert(<span class="hljs-string">"RGBA"</span>)
final_img = Image.composite(bk_img, bg_img, heatmap_img)
final_img.save(<span class="hljs-string">'img/result2.png'</span>)
</code></pre>
<p>heatmap_img
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563175814/BCW-fOdjO.png" alt="4.png" /></p>
<p>The result is, for the places with more scooters, a red overlay will be on them, so we can tell where the scooters are focused.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659563873772/Awb-PCnn9.png" alt="5.png" /></p>
<h2 id="heading-whats-next">What's next</h2>
<p>In the next blog, I plan to analyse the Voi scooter data from July, since my data collecting program has been running for over a month already. I'm sure there will be interesting data collected.</p>
]]></content:encoded></item><item><title><![CDATA[Project Exposure: [1] Creating Social Media Presence and Cross-posting on Dev and Medium]]></title><description><![CDATA[In this blog, I continue my journey of increasing my exposure by designing my brand and making social media presence for my blog on Twitter, Instagram, Facebook, and Reddit. As well as cross-posting on other blogging platforms.
My brand
To make socia...]]></description><link>https://blog.cpbprojects.me/project-exposure-1-creating-social-media-presence-and-cross-posting-on-dev-and-medium</link><guid isPermaLink="true">https://blog.cpbprojects.me/project-exposure-1-creating-social-media-presence-and-cross-posting-on-dev-and-medium</guid><category><![CDATA[social media]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[2Articles1Week]]></category><category><![CDATA[Blogging]]></category><category><![CDATA[Developer Blogging]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Thu, 28 Jul 2022 10:50:48 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005006620/EkShSWy5o.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog, I continue my journey of increasing my exposure by designing my brand and making social media presence for my blog on Twitter, Instagram, Facebook, and Reddit. As well as cross-posting on other blogging platforms.</p>
<h2 id="heading-my-brand">My brand</h2>
<p>To make social media accounts, I need a brand. I am just going to call it "Chit's Programming Blog". Then I also need an email for all the social media, but I don't want to use my personal email or work email because I don't want to get my emails entangled. So I made a new Gmail account since it is the most trusted. Popular social media may blacklist other email providers because they are less secured and may be used to create spam bots etc.</p>
<h3 id="heading-profile-picture">Profile picture</h3>
<p>I need a profile picture for social media. I don't want to use my photos, I'm also concerned about copyright and trademarks, so I cannot just snap an image online and use it. Therefore I need to design my own profile picture.</p>
<p>There are a handful of tools I can use. There are tools where I have to make everything from scratch, like Microsoft Paint, and Photoshop. There are also tools that give stock images and fonts for you, like Canva. There are also tools that claim to use AI to generate your logo if you tell it about what you want, like <a target="_blank" href="https://www.tailorbrands.com/logo-maker">Tailor Brands</a>, <a target="_blank" href="https://designs.ai/en/logomaker">Designs AI</a>, and <a target="_blank" href="https://www.adobe.com/express/create/logo">Adobe Express</a>.</p>
<p>I used Adobe Express since it is the bigger brand, and it says free to use forever on their website.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005041637/HD0rSvgfw.png" alt="1.png" /></p>
<h2 id="heading-social-media-presence">Social Media Presence</h2>
<p>Then I created my social media page on different platforms.</p>
<h3 id="heading-twitter">Twitter</h3>
<p>Once I created my basic Twitter account, I get to upgrade it into a <a target="_blank" href="https://business.twitter.com/en/basics/get-your-business-started-with-twitter.html">Twitter Business Account</a>, where I have to tell Twitter what company I am. Since there is no choice for content creator or blogger, I chose Information Technology Company. It also asks me if I am a shop or a creator, which I chose creator.</p>
<p>Then my Twitter account is finished. It gave me a handle "@BlogChit", which I am not so happy with, then I found out on <a target="_blank" href="https://help.twitter.com/en/managing-your-account/change-twitter-handle#:~:text=Navigate%20to%20Settings%20and%20privacy%20and%20tap%20Account.&amp;text=Tap%20Username%20and%20update%20the,prompted%20to%20choose%20another%20one.&amp;text=Tap%20Done.">their website</a> that I can change the handle. I want the handle to be memorable and easy to type out. Looking at popular Twitter users, most of them use the same as their display name, so I think I will do that too.</p>
<p>But I discovered my name is too long, so at last, I decided to go with "ChitProgramming"</p>
<h3 id="heading-instagram">Instagram</h3>
<p>At first, I thought I must use my real name on Instagram. But it turns out I can use my business name too, so I used it for searchability.</p>
<p>When I thought that I'm done making my Instagram account, I found a button to "Switch to professional account". I found <a target="_blank" href="https://www.socialmediatoday.com/news/6-reasons-why-you-need-to-switch-to-an-instagram-business-profile/519704/">this article</a> talking about why should I switch to a professional account, and the point Instagram Business Profiles Can Share Links in Instagram Stories sold me, it means that I'll be able to share my blog post with a single click. So I clicked the button, selected "Creator" to describe me better than "Business", choose "Personal Blog" as my category, and then I have a professional Instagram account.</p>
<h3 id="heading-facebook">Facebook</h3>
<p>This is basically the same with Instagram since it is owned by the same company, Meta. But one annoying thing is that it makes posts that I don't want. After I created the profile and added images, it made a post about my birthday, how I updated my profile picture and updated my cover photo.</p>
<p>Facebook seems to force you to use your real name to make an account, so I was confused why there are so many business pages on Facebook, it turns out that peoples use "profile", while businesses use "page", and "profile" to create "pages".</p>
<p>So I created indicating that my category is "Information Technology Company", added my blog's information, then it was done.</p>
<h3 id="heading-reddit">Reddit</h3>
<p>Reddit is the easiest of all, I just have to select a username, and enter my blog's info. I don't have to set up a professional account or anything.</p>
<h2 id="heading-cross-posting">Cross-posting</h2>
<p>I thought I had to make a program to automatically post my blog on other websites, then make a canonical link back to the original one. But turns out, that both Dev and Medium have features that help you publish articles from other sources.</p>
<h3 id="heading-devto">Dev.To</h3>
<p>Inside the Settings -&gt; Extensions page, there is an option where I can publish to the Dev community from RSS.  </p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005049430/IOgTTCnaO.png" alt="2.png" /></p>
<p>I get my RSS on the upper right corner of my blog, then I paste it on Dev.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005053338/fRJg31KMZ.png" alt="3.png" /></p>
<p>Then all my posts from Hashnode end up there, that's so cool and useful.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005056515/xdKYyJMTd.png" alt="4.png" /></p>
<p>I have to add tags manually, I have to check what tags are trending on Dev. I check this on their <a target="_blank" href="https://dev.to/tags">tags page</a></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005059077/xh4g1EIXZ.png" alt="5.png" /></p>
<h3 id="heading-medium">Medium</h3>
<p>As for Medium, I cross-post by going to Settings -&gt; Stories -&gt; Import a story.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005064116/VpC4W-aej.png" alt="6.png" /></p>
<p>Then I input the link to my article and click import.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005066940/N3qZxeZtY.png" alt="7.png" /></p>
<p>This one gave me some extra spacing between lines, which I have to remove.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1659005070710/r7PjA4uQh.png" alt="8.png" /></p>
<p>Medium doesn't share what tags are trending, it only recommended some <a target="_blank" href="https://medium.com/me/following/suggestions">topics</a>, so I have to find out in <a target="_blank" href="https://humbaa.com/60-most-popular-tags-on-medium/">a third-party website</a>.</p>
<h2 id="heading-whats-next">What's next</h2>
<p>In this article, I created social media accounts and cross-posting accounts. Now I can share my blog posts on social media and engage in discussions. And I will cross-post my previous blog posts at a regular interval, to avoid being spammy. As for SEO, since it takes time for Google to index my blog posts, I think I'll have to wait, and post an update once I can find myself on Google. </p>
]]></content:encoded></item><item><title><![CDATA[Project Exposure: [0] How do I increase my blog's view count]]></title><description><![CDATA[I've been writing this blog for about 2 months now and averages around 4 views per day, which is fair, considering how much competition is on Hashnode.  In light of this, I aim to increase my blog view in this project - Project Exposure.

Possible re...]]></description><link>https://blog.cpbprojects.me/project-exposure-0-how-do-i-increase-my-blogs-view-count</link><guid isPermaLink="true">https://blog.cpbprojects.me/project-exposure-0-how-do-i-increase-my-blogs-view-count</guid><category><![CDATA[SEO]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Hashnode]]></category><category><![CDATA[Developer Blogging]]></category><category><![CDATA[Blogging]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Mon, 25 Jul 2022 20:22:43 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1658780275044/aypl8-o4D.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I've been writing this blog for about 2 months now and averages around 4 views per day, which is fair, considering how much competition is on <a target="_blank" href="https://hashnode.com/">Hashnode</a>.  In light of this, I aim to increase my blog view in this project - Project Exposure.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779448923/iF1BafF0V.png" alt="1.png" /></p>
<h2 id="heading-possible-reasons-for-low-viewer-count">Possible reasons for low viewer count</h2>
<p>I start by investigating the shortcomings of my blog, to see how I can improve.</p>
<h3 id="heading-time-on-the-platform">Time on the platform</h3>
<p>I've only written several blog posts. There is a small chance in each of my blog posts to be appreciated and attract followers. So the longer I am on this platform, the higher chance this happens.</p>
<h3 id="heading-my-writing-skills">My writing skills</h3>
<p>I haven't written a blog before, and seldom write essays since I study Computer Science. Therefore my writing may wordy, unclear, and unpleasant to read. To fix this, I need to write more, as well as read good blog posts to learn.</p>
<h3 id="heading-my-writing-topics">My writing topics</h3>
<p>Looking at the trending blog posts on Hashnode, it is clear the trending ones teach you something or tell you something you don't know. But the ones I write document a programming journey, so it is not what most people look for.</p>
<p>However, I do not intend to change my article topics, as my blog is for documenting my programming journey and my target audience is people interested or wanting to learn from my experience. </p>
<p>I also notice a lot of popular blog posts have a number in them, like <a target="_blank" href="https://danyal.hashnode.dev/6-habits-i-have-picked-up-from-working-in-tech-for-3-years"><strong>6</strong> Habits I have picked up from working in tech for <strong>3</strong> years</a>, <a target="_blank" href="https://jamesqquick.hashnode.dev/15-common-beginner-javascript-mistakes"><strong>15</strong> Common Beginner JavaScript Mistakes</a>.  I thought about using AI, feeding in article topics and views, to see if there is any correlation with it, but I'll save this idea for later.</p>
<h3 id="heading-people-lack-a-chance-to-see-my-articles">People lack a chance to see my articles</h3>
<p>I watched a video from Youtuber <a target="_blank" href="https://www.youtube.com/c/HowMoneyWorks">How Money Works</a>, called <a target="_blank" href="https://www.youtube.com/watch?v=tNtnlzmvAw0">Why Finance "Gurus" Want You To Hate Them - How Money Works</a>. In that video, the author made an interesting point. The reason why fake gurus - people who give bad financial advice, act so unbearable, is because "any publicity is good publicity", working in conjunction with the sales funnel. </p>
<h4 id="heading-the-sales-funnel">The sales funnel</h4>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779457924/emjCaUtVT.png" alt="2.png" /></p>
<p><a target="_blank" href="https://www.crazyegg.com/blog/sales-funnel/#:~:text=The%20sales%20funnel%20is%20each%20step%20that%20someone%20has%20to%20take%20in%20order%20to%20become%20your%20customer.">The sales funnel is each step that someone has to take in order to become your customer</a>, as visualised above. Actually, it is more like a sieve or a filter, since a funnel keeps everything poured in, but some customers will be lost in the process. For the sake of recognisability, I'll keep calling it a sales funnel.</p>
<p>Let's say in each step of my sales funnel, I keep 10% of the customers. If I want more people to enter the last step, follow me, I can:
a) do better in each stage to keep more customers, such as making clickbait titles and writing a more interesting first paragraph of my article to increase the percentage of people staying in each stage.
b) put more people in the initial stage of the sales funnel, so that more people will enter the final stage.</p>
<p>That's why I came up with the idea of making a program which automatically cross-posts my blogs and also creates social media presence to promote my blog to increase my exposure and put more people inside my sales funnel.</p>
<h3 id="heading-popular-blogging-sites">Popular blogging sites</h3>
<p>To do that I will have to know where should I cross-post my blog, below are websites that I found online.</p>
<ul>
<li>Hashnode, where I'm on</li>
<li>HackerNoon and FreeCodeCamp, have dedicated editors to review blog posts, so I may only submit high-quality blog posts</li>
<li>Dev.To, simular to Hashnode</li>
<li>Medium, a website for all kinds of articles</li>
<li>Google Blogger / WordPress / Wix..., these sites while giving me a lot of freedom, there is no way for people to find me there, as there is no <code>explore</code> button, so I will not consider these, as at this stage, I want to be discovered</li>
<li>LinkedIn, a place with professionals and experts, so I need to be careful what I post there</li>
<li>Tumblr, a mini blogging website with built-in social media capabilities</li>
</ul>
<h3 id="heading-social-media-where-i-can-create-a-presence">Social media where I can create a presence</h3>
<ul>
<li>Twitter</li>
<li>Reddit</li>
<li>Instagram</li>
<li>Facebook</li>
<li>Snapchat</li>
<li>Youtube and TikTok, if I ever consider making videos</li>
<li>Pinterest, seems to be image-based but I don't think I have any images to show</li>
</ul>
<h3 id="heading-search-engine-optimization">Search Engine Optimization</h3>
<p>Ideally, my blog posts should be searchable from Google, so there is an additional way for people to visit my blog. But currently, my blogs cannot be searched on Google.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779463515/GNYcl5ILi.png" alt="3.png" /></p>
<p>The first reason is that I have an apostrophe in my blog title and that "chit" is usually used with "chit chat", so Google thinks it was a typo. If I force Google to search "Chit's", my blog appears on top.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779467081/P7KMvdqAS.png" alt="4.png" /></p>
<p>The other reason is that I recently changed the name from "Chit's tech blog" to "Chit's programming blog" because I realized that a tech blog is a blog that talks about new technology, which is not what I cover.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779471046/UcN04Hj19.png" alt="5.png" /></p>
<p>When I search for the title of my first article, I cannot find the result. It is because the past me thought it would be a good idea to make a custom SEO title. If I have time, I want to edit all my past blog posts' SEO titles and descriptions, for Google to better index my blog posts.</p>
<p>Also, notice that my article published on DevTo was indexed instead. This is undesirable because my articles on different platforms are fighting for views, what I should do instead is to use a canonical link, to tell Google, that the article you found on DevTo, comes from Hashnode, so Google knows where to send people.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779474452/pPdWSSMvy.png" alt="6.png" /></p>
<p>When I search for newer articles, they do not appear in the result, I assume it is because it takes time for Google to index my blog posts.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779481210/myyCrn0_q.png" alt="7.png" /></p>
<p>I also designed an image for social media sharing, so it looks prettier when I share my blog.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779484314/avHwvH69E.png" alt="8.png" /></p>
<p>I have to say, <a target="_blank" href="https://www.canva.com/">Canva</a> is really easy to use, and has some nice templates.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658779488200/_ZqCRMlTb.png" alt="9.png" /></p>
<p>Now All I have to do is wait for Google to index my new and improved blog.</p>
<h2 id="heading-action-plan">Action plan</h2>
<p>In the next blog, I want to create a social media presence on the above social media platforms, as well as investigate how can I cross-post to the other blogging platforms.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [8] Creating Heatmap for distribution of scooters]]></title><description><![CDATA[In this blog, I aim to investigate the distribution of scooters.
Finding out the latitude and longitude of my city
Scooter data uses latitude and longitude to store where they are. Latitude and longitude are a way to represent where something is in t...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-8-creating-heatmap-for-distribution-of-scooters</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-8-creating-heatmap-for-distribution-of-scooters</guid><category><![CDATA[Voi]]></category><category><![CDATA[Python]]></category><category><![CDATA[Data Science]]></category><category><![CDATA[data analysis]]></category><category><![CDATA[Programming Blogs]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Thu, 21 Jul 2022 13:13:06 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1658409084389/TX4wEmzRs.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog, I aim to investigate the distribution of scooters.</p>
<h2 id="heading-finding-out-the-latitude-and-longitude-of-my-city">Finding out the latitude and longitude of my city</h2>
<p>Scooter data uses latitude and longitude to store where they are. Latitude and longitude are a way to represent where something is in the world by how far it is from the equator and the prime meridian.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408666963/3ptcugnZh.png" alt="1.png" /></p>
<h3 id="heading-heatmap">Heatmap</h3>
<p>I aim to use a heat map to plot the concentration of the scooters. Heat maps show which part of the map is hot. It is used for three-dimensional data where the x and y axis are for the first two dimensions, and the colour (heat) is for the third. In my case, the x and y axis will be used for the latitude and longitude, and the colour will represent how many scooters are in that area.</p>
<h3 id="heading-changing-data-representation">Changing data representation</h3>
<p>Currently, my scooter location data is split into two dataframe, one for latitude and one for longitude. I want to get all the scooter locations from a timestamp, so I will get the data from the same row from both dataframes.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408673343/Uz7RvwA14.png" alt="2.png" /></p>
<p>Then I will fill the data into a dataframe where the latitude and the longitude are the two axes, and the value in the middle means how many scooters are in that coordinate range. So I want to find the 4 corners of the city where there are scooters, so I know how to set the range for the latitude and longitude.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408676896/vukAtL9z6.png" alt="3.png" /></p>
<h3 id="heading-fixing-floating-point-precision-using-string-formatting">Fixing floating point precision using string formatting</h3>
<p>Then I turn the four corners into a range, I use the <code>np.arange()</code> function, however, due to the precision problem of floating point numbers, the numbers are not accurate.</p>
<pre><code class="lang-python">list(np.arange(<span class="hljs-number">52.59</span>, <span class="hljs-number">52.40</span>, <span class="hljs-number">-0.005</span>))
</code></pre>
<p>the output is, [52.59,
 52.585,
 52.58,
 52.574999999999996,
 52.56999999999999,...</p>
<p> And the dataframe looks like this</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408683351/iZnoau2GA.png" alt="4.png" /></p>
<p>Therefore I used string formatting and list comprehension to turn this list of floating point numbers into a string, which doesn't have the problem of precision.</p>
<pre><code class="lang-python">lat_index = np.arange(<span class="hljs-number">52.59</span>, <span class="hljs-number">52.40</span>, <span class="hljs-number">-0.005</span>)
str_lat_index = [<span class="hljs-string">"%.3f"</span> % i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> lat_index]
</code></pre>
<h3 id="heading-creating-the-dataframe">Creating the dataframe</h3>
<p>Now I can create the dataframe from the str index and column, I used <code>fillna(0)</code> because at first of creating the dataframe, all values are null values, and since I am counting the number of scooters, everything should be 0 at the beginning.</p>
<pre><code class="lang-python">scooter_area_df = pd.DataFrame(index=str_lat_index, columns=str_lng_column, dtype=np.int64).fillna(<span class="hljs-number">0</span>)
</code></pre>
<h3 id="heading-filling-the-dataframe">Filling the dataframe</h3>
<p>Then I extracted the specific row from the dataframe, for the starting index, I took the value at 2022-06-25 at 5 am.</p>
<pre><code class="lang-python">df_starting_index = <span class="hljs-number">26852</span>

latitude_df_row = latitude_df.iloc[df_starting_index]
longitude_df_row = longitude_df.iloc[df_starting_index]
</code></pre>
<p>And I fill in the <code>scooter_area_df</code> by going through the latitude dataframe and longitude dataframe, fetching the value from both of them and incrementing the particular cell in the scooter_area_df. Notice that I used the lazy approach, which is to round down the data to the nearest 2 decimal point, this is not accurate because what I should have down is to round up/down depending which one is closer.</p>
<pre><code class="lang-python"><span class="hljs-keyword">for</span> lat, lng <span class="hljs-keyword">in</span> zip(latitude_df_row, longitude_df_row):
    <span class="hljs-keyword">try</span>:
        <span class="hljs-keyword">if</span> <span class="hljs-keyword">not</span> pd.isnull(lat):
            str_lat = <span class="hljs-string">"%.2f"</span> % lat
            str_lng = <span class="hljs-string">"%.2f"</span> % lng
            scooter_area_df.loc[str_lat, str_lng] += <span class="hljs-number">1</span>
    <span class="hljs-keyword">except</span> KeyboardInterrupt:
        <span class="hljs-keyword">break</span>
    <span class="hljs-keyword">except</span>:
        traceback.print_exc()
</code></pre>
<h3 id="heading-plotting-the-graph">Plotting the graph</h3>
<p>I then plot the graph using the seaborn library and add the appropriate labels to it.</p>
<p><a target="_blank" href="https://stackoverflow.com/a/44484758/16929051">fun fact</a>: We usually import seaborn as sns because Samuel Norman "Sam" Seaborn is a fictional character portrayed by Rob Lowe on the television serial drama The West Wing.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> seaborn <span class="hljs-keyword">as</span> sns
fig, ax = plt.subplots(figsize=(<span class="hljs-number">12</span>, <span class="hljs-number">10</span>))
fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
ax = sns.heatmap(scooter_area_df, square=<span class="hljs-literal">True</span>)
plt.xlabel(<span class="hljs-string">"Longitude"</span>)
plt.ylabel(<span class="hljs-string">"Latitude"</span>)
plt.title(<span class="hljs-string">"The concentration of scooters"</span>)
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408692825/TbnUwcVKr.png" alt="5.png" /></p>
<h2 id="heading-problems-with-the-current-plot">Problems with the current plot</h2>
<p>This result is not desirable because:</p>
<ol>
<li>The squares are too big, I wish we could be more precise</li>
<li>It may not be geographically accurate, as 1 latitude is usually not as long as one longitude, ideally the heatmap will look the same as if it is plotted on an actual map.</li>
</ol>
<h3 id="heading-problem-1-being-more-precise">Problem 1: being more precise</h3>
<p>To do this, How much we decrease the square size matters a lot, because if we decrease it too much. If there is a bunch of scooters next to each other, they might be separated into different areas, and we cannot see how concentrated they are.</p>
<p>The following is an example of what happened if we decreased the step, which is the square size, by 10 times, you can hardly see the dots.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408698142/AneBndsUX.png" alt="6.png" /></p>
<p>So I need to find an appropriate amount to decrease the step.</p>
<h3 id="heading-problem-2-the-actual-ratio-of-the-map">Problem 2: the actual ratio of the map</h3>
<p>To do this, I plan to measure the relationship between the distance between the two points, and the differences in latitude and longitude.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408702562/hDTG7rIPO.png" alt="7.png" /></p>
<p>The horizontal distance is 13.53 km, and the longitude difference is 0.2</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408708593/-o-ZC2gBC.png" alt="8.png" /></p>
<p>The vertical distance is 15.57 km, and the latitude difference is 0.14</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408712043/c0Lao6hwmc.png" alt="9.png" /></p>
<p>So let's say if I want each square to be 100 meters, then the vertical step will be 0.14/155.7, and the horizontal step will be 0.2/135.3.</p>
<h3 id="heading-find-the-closest-value-in-a-sorted-array">Find the closest value in a sorted array</h3>
<p>Previously, I took the lazy route and just used string formatting to do the rounding, but since now the intervals are not multiple of 10, I need to find a new way. I aim to find which value from the vertical range or horizontal range is the closest to the value, so a binary search should be useful.</p>
<p>I copied the code from this <a target="_blank" href="https://stackoverflow.com/a/12141511/16929051">stackoverflow answer</a>, which uses the bisect library for a quick way to find the nearest result, the code looks like this:</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> bisect <span class="hljs-keyword">import</span> bisect_left

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">take_closest</span>(<span class="hljs-params">myList, myNumber</span>):</span>
    <span class="hljs-string">"""
    Assumes myList is sorted. Returns closest value to myNumber.
    If two numbers are equally close, return the smallest number.
    """</span>
    pos = bisect_left(myList, myNumber)
    <span class="hljs-keyword">if</span> pos == <span class="hljs-number">0</span>:
        <span class="hljs-keyword">return</span> myList[<span class="hljs-number">0</span>]
    <span class="hljs-keyword">if</span> pos == len(myList):
        <span class="hljs-keyword">return</span> myList[<span class="hljs-number">-1</span>]
    before = myList[pos - <span class="hljs-number">1</span>]
    after = myList[pos]
    <span class="hljs-keyword">if</span> after - myNumber &lt; myNumber - before:
        <span class="hljs-keyword">return</span> after
    <span class="hljs-keyword">else</span>:
        <span class="hljs-keyword">return</span> before
</code></pre>
<p>I also changed the code to fit the code that updates the value inside the scooter_area_df, notice that I used lat_index_reversed, since take_cloest only works on an ascendingly sorted array, I had to make a reversed latitude array to make it work, take_cloest returns a value instead of an index, so I don't have to modify the returned value.</p>
<pre><code class="lang-python">str_lat = <span class="hljs-string">"%.3f"</span> % take_closest(lat_index_reversed, lat)
str_lng = <span class="hljs-string">"%.3f"</span> % take_closest(lng_column, lng)
scooter_area_df.loc[str_lat, str_lng] += <span class="hljs-number">1</span>
</code></pre>
<p>now the heatmap looks like this, which when looking at the Voi app, is very similar to how my city looks</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408720868/MNDyVK8NM.png" alt="10.png" /></p>
<h3 id="heading-turning-it-all-into-a-function">Turning it all into a function</h3>
<p>I now made <code>plot_concentration_of_scooters</code> a function, in which I first get the two dataframes outside of the function so it will only be called once, then declare where latitude and longitude start and end.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> seaborn <span class="hljs-keyword">as</span> sns

longitude_df = store[<span class="hljs-string">'longitude_df'</span>]
latitude_df = store[<span class="hljs-string">'latitude_df'</span>]

lat_start_index = <span class="hljs-number">52.54</span>
lat_end_index = <span class="hljs-number">52.40</span>

lng_start_index = <span class="hljs-number">-1.99</span>
lng_end_index = <span class="hljs-number">-1.79</span>
</code></pre>
<p>Then I calculate the vertical step and horizontal step, create the ranges, make the scooter area dataframe, fill it, and plot the result.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_concentration_of_scooters</span>(<span class="hljs-params">df_index, square_side_in_meters</span>):</span>
    lat_step = (lat_end_index - lat_start_index) / (<span class="hljs-number">15.57</span> * <span class="hljs-number">1000</span> / square_side_in_meters)
    lng_step = (lng_end_index - lng_start_index) / (<span class="hljs-number">13.53</span> * <span class="hljs-number">1000</span> / square_side_in_meters)

    lat_index = np.arange(lat_start_index, lat_end_index, lat_step)
    lat_index_reversed = lat_index[::<span class="hljs-number">-1</span>]
    str_lat_index = [<span class="hljs-string">"%.3f"</span> % i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> lat_index]

    lng_column = np.arange(lng_start_index, lng_end_index, lng_step)
    str_lng_column = [<span class="hljs-string">"%.3f"</span> % i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> lng_column]

    scooter_area_df = pd.DataFrame(index=str_lat_index, columns=str_lng_column, dtype=np.int64).fillna(<span class="hljs-number">0</span>)

    latitude_df_row = latitude_df.iloc[df_index]
    longitude_df_row = longitude_df.iloc[df_index]

    <span class="hljs-keyword">for</span> lat, lng <span class="hljs-keyword">in</span> zip(latitude_df_row, longitude_df_row):
        <span class="hljs-keyword">if</span> <span class="hljs-keyword">not</span> pd.isnull(lat):
            str_lat = <span class="hljs-string">"%.3f"</span> % take_closest(lat_index_reversed, lat)
            str_lng = <span class="hljs-string">"%.3f"</span> % take_closest(lng_column, lng)
            scooter_area_df.loc[str_lat, str_lng] += <span class="hljs-number">1</span>

    fig, ax = plt.subplots(figsize=(<span class="hljs-number">12</span>, <span class="hljs-number">10</span>))
    fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
    ax = sns.heatmap(scooter_area_df, square=<span class="hljs-literal">True</span>)
    plt.xlabel(<span class="hljs-string">"Longitude"</span>)
    plt.ylabel(<span class="hljs-string">"Latitude"</span>)
    plt.title(<span class="hljs-string">f"The concentration of scooters on <span class="hljs-subst">{longitude_df_row.name}</span>"</span>)
</code></pre>
<p>here is what happens when I use different values</p>
<p><code>plot_concentration_of_scooters(38339, 500)</code></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408727314/g5p0jkJ76.png" alt="11.png" /></p>
<h3 id="heading-making-multiple-results-into-a-gif">Making multiple results into a gif</h3>
<p>With a still frame, I can only tell the distribution in a static manner, what if I want to show the change over time? I want to make a gif to show the change, like in a day.</p>
<p>I found <a target="_blank" href="https://towardsdatascience.com/basics-of-gifs-with-pythons-matplotlib-54dd544b6f30">this article</a> that is really helpful, it talks about how to make multiple plots into a graph.</p>
<p>What I came up with is to have a temporary folder, then I plot each graph, after that saving them with <code>plt.savefig()</code>, and then using the <code>imageio</code> library to turn them into a gif file, after that deleting the temporary plot images.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> imageio

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">make_scooter_concentration_gif</span>(<span class="hljs-params">df_start_index, df_end_index, step, square_size_in_meters</span>):</span>
    temp_folder = Path(<span class="hljs-string">'temp/scooter_concentration_gif/'</span>)
    df_index_range = range(df_start_index, df_end_index, step)
    <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
        plot_concentration_of_scooters(df_index, square_size_in_meters)
        plt.savefig(temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>))
        plt.close()

    <span class="hljs-keyword">with</span> imageio.get_writer(<span class="hljs-string">'graphs/mygif.gif'</span>, mode=<span class="hljs-string">'I'</span>, duration=<span class="hljs-number">1</span>) <span class="hljs-keyword">as</span> writer:
        <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
            image = imageio.imread(temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>))
            writer.append_data(image)

    <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
        temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>).unlink()
</code></pre>
<p>the result of <code>make_scooter_concentration_gif(28288, 29304, 100, 200)</code> is</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408734391/oZBN0KlQj.gif" alt="12.gif" /></p>
<h3 id="heading-improving-the-function">Improving the function</h3>
<p>This is good, but one that I have against it is that it loops too quick and it is hard to tell when is the graph ending, I can either change the loop variable for the gif to make it not loop, or I can add a blank frame at the end so the users know if a new cycle begins, I prefer the latter solution.</p>
<p>Also, in each image, the scale for the anchor the colourmap is different, sometimes the max is 20, sometimes 10, so it is hard to tell for a single pixel, whether it is gaining scooters or not, so I modified the code the have a <code>vmax</code> argument, to fix the maximum for the scale.</p>
<p>the result looks like this <code>make_scooter_concentration_gif(28288, 29304, 100, 200)</code></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408739558/ywDCbBGsN.gif" alt="13.gif" /></p>
<p>at another day <code>make_scooter_concentration_gif(36903, 37920, 60, 200, 15)</code></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1658408744446/3iL21R-D1.gif" alt="14.gif" /></p>
<h3 id="heading-final-version-of-the-code-for-reference">Final version of the code for reference</h3>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_concentration_of_scooters</span>(<span class="hljs-params">df_index, square_side_in_meters, vmax=None</span>):</span>
    lat_step = (lat_end_index - lat_start_index) / (<span class="hljs-number">15.57</span> * <span class="hljs-number">1000</span> / square_side_in_meters)
    lng_step = (lng_end_index - lng_start_index) / (<span class="hljs-number">13.53</span> * <span class="hljs-number">1000</span> / square_side_in_meters)

    lat_index = np.arange(lat_start_index, lat_end_index, lat_step)
    lat_index_reversed = lat_index[::<span class="hljs-number">-1</span>]
    str_lat_index = [<span class="hljs-string">"%.3f"</span> % i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> lat_index]

    lng_column = np.arange(lng_start_index, lng_end_index, lng_step)
    str_lng_column = [<span class="hljs-string">"%.3f"</span> % i <span class="hljs-keyword">for</span> i <span class="hljs-keyword">in</span> lng_column]

    scooter_area_df = pd.DataFrame(index=str_lat_index, columns=str_lng_column, dtype=np.int64).fillna(<span class="hljs-number">0</span>)

    latitude_df_row = latitude_df.iloc[df_index]
    longitude_df_row = longitude_df.iloc[df_index]

    <span class="hljs-keyword">for</span> lat, lng <span class="hljs-keyword">in</span> zip(latitude_df_row, longitude_df_row):
        <span class="hljs-keyword">if</span> <span class="hljs-keyword">not</span> pd.isnull(lat):
            str_lat = <span class="hljs-string">"%.3f"</span> % take_closest(lat_index_reversed, lat)
            str_lng = <span class="hljs-string">"%.3f"</span> % take_closest(lng_column, lng)
            scooter_area_df.loc[str_lat, str_lng] += <span class="hljs-number">1</span>

    fig, ax = plt.subplots(figsize=(<span class="hljs-number">12</span>, <span class="hljs-number">10</span>))
    fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
    ax = sns.heatmap(scooter_area_df, square=<span class="hljs-literal">True</span>, xticklabels=<span class="hljs-number">5</span>, yticklabels=<span class="hljs-number">5</span>, vmax=vmax)
    plt.xlabel(<span class="hljs-string">"Longitude"</span>)
    plt.ylabel(<span class="hljs-string">"Latitude"</span>)
    plt.title(<span class="hljs-string">f"The concentration of scooters on <span class="hljs-subst">{longitude_df_row.name}</span>"</span>)
</code></pre>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> imageio

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">make_scooter_concentration_gif</span>(<span class="hljs-params">df_start_index, df_end_index, step, square_size_in_meters, vmax=None</span>):</span>
    temp_folder = Path(<span class="hljs-string">'temp/scooter_concentration_gif/'</span>)
    df_index_range = range(df_start_index, df_end_index, step)
    <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
        plot_concentration_of_scooters(df_index, square_size_in_meters, vmax=vmax)
        plt.savefig(temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>))
        plt.close()

    <span class="hljs-keyword">with</span> imageio.get_writer(<span class="hljs-string">'graphs/mygif.gif'</span>, mode=<span class="hljs-string">'I'</span>, duration=<span class="hljs-number">1</span>) <span class="hljs-keyword">as</span> writer:
        <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
            image = imageio.imread(temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>))
            writer.append_data(image)
        writer.append_data(imageio.imread(<span class="hljs-string">'temp/scooter_concentration_gif/blank.png'</span>))

    <span class="hljs-keyword">for</span> df_index <span class="hljs-keyword">in</span> df_index_range:
        temp_folder.joinpath(<span class="hljs-string">f'<span class="hljs-subst">{df_index}</span>.png'</span>).unlink()
</code></pre>
<h2 id="heading-whats-next">What's next</h2>
<p>For the next blog, I aim to incorporate the map into the plots, like making the map as a background image in the heatmap or plotting the scooters at dots on the map.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [7] Analysis of scooter battery with graphs]]></title><description><![CDATA[I will be trying to learn from the data set I collected, starting with the battery of scooters.
Deal with missing data
Because of my data collection program crashing, there are some data missing, I now need to find them out, so that I don't mess up m...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-7-analysis-of-scooter-battery-with-graphs</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-7-analysis-of-scooter-battery-with-graphs</guid><category><![CDATA[Voi]]></category><category><![CDATA[Data Science]]></category><category><![CDATA[Python]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[data analysis]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Mon, 11 Jul 2022 19:28:24 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1657567482911/qUzLkT1ol.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I will be trying to learn from the data set I collected, starting with the battery of scooters.</p>
<h1 id="heading-deal-with-missing-data">Deal with missing data</h1>
<p>Because of my data collection program crashing, there are some data missing, I now need to find them out, so that I don't mess up my data analysis by making incorrect assumptions.</p>
<h2 id="heading-finding-missing-data">Finding missing data</h2>
<p>I am going to iterate through all the data I have and check the separation between it and the previous data item. The data is collected every 1 minute, so any time longer than that indicates there is something wrong. But since running the program also takes time, the separation should be 1 minute + the duration of each loop. There might also be cases where there is a slight error and the program skips a minute. So I will check for data separated by more than 3 minutes to be safe.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">find_missing_data</span>():</span>
    battery_level_df = store[<span class="hljs-string">'battery_level_df'</span>]
    previous_timestamp = battery_level_df.index[<span class="hljs-number">0</span>]
    <span class="hljs-keyword">for</span> current_timestamp <span class="hljs-keyword">in</span> battery_level_df.index:
        <span class="hljs-keyword">if</span> current_timestamp - previous_timestamp &gt; pd.to_timedelta(<span class="hljs-string">"3 minutes"</span>):
            print(<span class="hljs-string">f"Missing data from <span class="hljs-subst">{previous_timestamp}</span> to <span class="hljs-subst">{current_timestamp}</span>"</span>)
        previous_timestamp = current_timestamp

find_missing_data()
</code></pre>
<p>The program finds that there are missing data in the following positions:<br />Missing data from 2022-06-02 21:59:47.420510 to 2022-06-02 23:24:21.071122<br />Missing data from 2022-06-04 14:46:03.130303 to 2022-06-05 19:08:46.083742<br />Missing data from 2022-06-15 09:04:56.485744 to 2022-06-17 21:15:31.558383<br />Missing data from 2022-06-22 09:05:24.070133 to 2022-06-22 15:09:05.898395  </p>
<h2 id="heading-finding-day-cuts">Finding day cuts</h2>
<p>Scooters are available from 6 am to 10 pm, so knowing when the day starts and ends is important to our analysis. Since there is no direct way to tell, I need to find them out.</p>
<p>It works by iterating through the timestamp and vehicle count array, then whenever the vehicle count drops the zero, we know it is 10 pm and the day ended, and we set <code>zero_flag</code> to True, then when the vehicle count leaves zero while <code>zero_flag</code> is True, we know it is 6 am, this repeats until the last data item.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">find_day_cuts</span>():</span>
    timestamp_np_array = np.load(<span class="hljs-string">'cached_data/timestamp_np_array.npy'</span>, allow_pickle=<span class="hljs-literal">True</span>)
    vehicle_count_np_array = np.load(<span class="hljs-string">'cached_data/vehicle_count_np_array.npy'</span>, allow_pickle=<span class="hljs-literal">True</span>)
    zero_flag = <span class="hljs-literal">False</span>
    <span class="hljs-keyword">for</span> index, item <span class="hljs-keyword">in</span> enumerate(zip(timestamp_np_array, vehicle_count_np_array)):
        <span class="hljs-keyword">if</span> item[<span class="hljs-number">1</span>] == <span class="hljs-number">0</span> <span class="hljs-keyword">and</span> <span class="hljs-keyword">not</span> zero_flag:
            print(<span class="hljs-string">"become zero"</span>, index, item[<span class="hljs-number">0</span>])
            zero_flag = <span class="hljs-literal">True</span>
        <span class="hljs-keyword">if</span> zero_flag <span class="hljs-keyword">and</span> item[<span class="hljs-number">1</span>] != <span class="hljs-number">0</span>:
            print(<span class="hljs-string">"leave zero"</span>, index, item[<span class="hljs-number">0</span>])
            zero_flag = <span class="hljs-literal">False</span>

find_day_cuts()
</code></pre>
<p>The program outputs each day like this:\
become zero 594 2022-06-02 23:24:21.071122\
leave zero 930 2022-06-03 05:00:53.124146</p>
<h1 id="heading-analysis-of-battery">Analysis of battery</h1>
<h2 id="heading-individual-scooter">Individual scooter</h2>
<p>I am going to analyse the battery of a scooter. I first read the dataframe from the HDFS storage, then I generate the series from the dataframe, and change the first and last value to zero so that it will not be skipped.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_scooter_battery</span>(<span class="hljs-params">starting_index, steps, scooter_id</span>):</span>
    <span class="hljs-keyword">global</span> battery_level_df
    series = battery_level_df.iloc[starting_index: starting_index + steps, battery_level_df.columns.get_loc(scooter_id)]
    series[<span class="hljs-number">0</span>] = <span class="hljs-number">0</span>
    series[<span class="hljs-number">-1</span>] = <span class="hljs-number">0</span>
    fig, ax = plt.subplots()
    fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
    series.plot(title=<span class="hljs-string">f"Battery of scooter <span class="hljs-subst">{scooter_id}</span> from index <span class="hljs-subst">{starting_index}</span> to <span class="hljs-subst">{starting_index + steps}</span>"</span>, xlabel=<span class="hljs-string">"timestamp"</span>, ylabel=<span class="hljs-string">"battery level"</span>, ax=ax)

plot_scooter_battery(<span class="hljs-number">930</span>,<span class="hljs-number">1400</span>,<span class="hljs-string">'npv9'</span>)
</code></pre>
<p>To illustrate my point, the graph without skipping zero looks like this</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566629058/A6h35w_mv.png" alt="img1.png" /></p>
<p>By setting the last item in the series as zero, the graph is forced to be extended, so we can see the whole thing.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566649155/cGEZjAU-r.png" alt="img2.png" /></p>
<h3 id="heading-critical-mistake">Critical mistake</h3>
<p>Then I realized that by changing the starting and ending value of the series, I actually accidentally edited the dataframe that I am not supposed to edit. This is because by calling the <code>iloc</code> function, the returned series is actually a reference, instead of a copy, so any edit done to that reference will reflect on the dataframe as well.</p>
<p>I was afraid that I have to re-run the function which took an hour to run to make the dataframe again. Luckily, I didn't save my edit to the HDFS storage, so I could load it again.</p>
<p>So I changed the function to make a copy of the series by calling <code>.copy()</code>, and I only edit the series if it is NaN, the updated function is:</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_scooter_battery</span>(<span class="hljs-params">starting_index, steps, scooter_id</span>):</span>
    <span class="hljs-keyword">global</span> battery_level_df
    series = battery_level_df.iloc[starting_index: starting_index + steps, battery_level_df.columns.get_loc(scooter_id)].copy()
    series[<span class="hljs-number">0</span>] = <span class="hljs-number">0</span> <span class="hljs-keyword">if</span> pd.isnull(series[<span class="hljs-number">0</span>]) <span class="hljs-keyword">else</span> series[<span class="hljs-number">0</span>]
    series[<span class="hljs-number">-1</span>] = <span class="hljs-number">0</span> <span class="hljs-keyword">if</span> pd.isnull(series[<span class="hljs-number">-1</span>]) <span class="hljs-keyword">else</span> series[<span class="hljs-number">-1</span>]
    fig, ax = plt.subplots()
    fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
    series.plot(title=<span class="hljs-string">f"Battery of scooter <span class="hljs-subst">{scooter_id}</span> from index <span class="hljs-subst">{starting_index}</span> to <span class="hljs-subst">{starting_index + steps}</span>"</span>, xlabel=<span class="hljs-string">"timestamp"</span>, ylabel=<span class="hljs-string">"battery level"</span>, ax=ax)
</code></pre>
<h3 id="heading-other-graphs">Other graphs</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566656115/Bmh_TvDZF.png" alt="img3.png" /></p>
<p>what happens when that scooter is not available on that entire day</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566660717/0ZCAoR-Fu.png" alt="img4.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566662184/n08RWj5MB.png" alt="img5.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566663798/oJFyXwTKh.png" alt="img6.png" /></p>
<h2 id="heading-mean-battery-level-of-all-scooters">Mean battery level of all scooters</h2>
<h3 id="heading-difficulty">Difficulty</h3>
<p>This plot sounds simple at first, I just have to use the <code>mean()</code> function from pandas and everything is done. The challenge is with NaN values, there are scooters that are not available for the entire day, and there are also scooters that are being used. </p>
<p>To illustrate my point, I drew this simulation, there are 3 scooters, scooters A, B and C, the line is its battery level, in between the lines is the time when the scooter is being used, and the value is NaN. At p1, no scooter is being used, so the mean is the mean of all scooters. At p2, scooters A and B are being used, so they are not being used to calculate the mean, so the means is the remaining scooter, scooter C. At p3, all scooters are back, so the mean goes back up.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566669917/96PSFhPC3.png" alt="img7.png" /></p>
<p>This isn't the real mean battery of all scooters. This is what the graph looks like without any editing.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566679250/tK8fi-fsN.png" alt="img8.png" /></p>
<h3 id="heading-plan-of-action">Plan of action</h3>
<p>Ideally, I want to fill in the blanks with a line going from the previous point to the next point, as if the battery consumption is linear. However, that would be difficult. When I am iterating through the elements, I'll have to find the point when the data stops for a particular scooter, go forward until the data reappear, and fill in the data linearly. That is fine for a scooter, but will be very difficult with so many scooters. </p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566683454/87RgwmetH.png" alt="img9.png" /></p>
<p>Moreover, with so many scooters, a forward fill, where I fill NaN values with the previous value will work very similarly because of the sample size.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566688468/jYBhpBEKR.png" alt="img10.png" /></p>
<p>So what I did is to first copy the specific area of the dataframe I want to analyse and copy it, so I won't change the original dataframe. Then I use the <code>dropna()</code> function to drop columns where all the values are NaN, meaning that particular scooter wasn't available that entire day, and we don't need it. Then I looped through all the elements in the copied dataframe to fill in the NaN values.</p>
<pre><code class="lang-python">starting_index = <span class="hljs-number">930</span>
steps = <span class="hljs-number">1945</span><span class="hljs-number">-930</span>

battery_level_df = store[<span class="hljs-string">'battery_level_df'</span>]
selected_area_df = battery_level_df.iloc[starting_index: starting_index + steps].copy()
selected_area_df.dropna(axis=<span class="hljs-string">'columns'</span>, how=<span class="hljs-string">'all'</span>, inplace=<span class="hljs-literal">True</span>)
selected_area_df_without_filling_nan = selected_area_df.copy()

<span class="hljs-keyword">for</span> row_num, (row_index, row_content) <span class="hljs-keyword">in</span> enumerate(selected_area_df.iterrows()):
    <span class="hljs-keyword">for</span> item_num, (item_index, item_content) <span class="hljs-keyword">in</span> enumerate(row_content.iteritems()):
        <span class="hljs-keyword">if</span> pd.isnull(item_content):
            selected_area_df.iloc[row_num, item_num] = selected_area_df.iloc[row_num - <span class="hljs-number">1</span>, item_num]
</code></pre>
<p>Then I plot it using the following code:</p>
<pre><code class="lang-python">plt.figure(facecolor=<span class="hljs-string">'white'</span>, figsize=(<span class="hljs-number">15</span>, <span class="hljs-number">8</span>))
plt.plot(selected_area_df.mean(axis=<span class="hljs-number">1</span>))
plt.ylim([<span class="hljs-number">55</span>,<span class="hljs-number">65</span>])
plt.title(<span class="hljs-string">"Mean battery level over time using my own way to fill na"</span>)
plt.xlabel(<span class="hljs-string">"Timestamp"</span>)
plt.ylabel(<span class="hljs-string">"Battery Level"</span>)
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566696617/_l4dqizFW.png" alt="img11.png" /></p>
<p>But then I realized that I can use the inbuilt <code>fillna</code> function with the <code>ffill</code> method, which forward fills the value. I ran the function and plotted the graph, this function ran faster than my own.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566701681/Okil8wp6M.png" alt="img12.png" /></p>
<h3 id="heading-different-results-from-my-implementation-and-the-inbuilt-method">Different results from my implementation and the inbuilt method</h3>
<p>But very confusingly, although these two plot looks similar, there are some very minor differences. These two functions are supposed to work the same way, therefore I think I have to investigate what caused the differences.</p>
<p>I suspect it is because in my implementation, I did <code>selected_area_df.iloc[row_num, item_num] = selected_area_df.iloc[row_num - 1, item_num]</code>, if the first element is NaN, the value from the end will be used to fill it because of how negative indices works. I investigated and found there are some scooters that were initially unavailable, but then were activated in the middle of the day. That is why the result is different, I forgot about the edge case. </p>
<h3 id="heading-assumption-about-the-data">Assumption about the data</h3>
<p>This brings me to the problem, what should I assume about the scooters that aren't available at the start? There are 3 possibilities:</p>
<ol>
<li>It was broken and had to be fixed</li>
<li>It was out of battery and need a battery replacement</li>
<li>It was being used, but this is not possible, because it is just after the scooters are available, so no one has time to unlock scooters yet</li>
</ol>
<p>I think most of them will be out of battery instead of broken, so I plan to fill the lowerest value of battery to them, I know for a fact that scooters don't get deactivated when they hit zero, they usually are deactivated before that to prevent accidents where a scooter run out of battery mid-ride happen, so the problem is when, when does the scooter get diactivated? I looked through the dataframe, and usually, the scooter gets deactivated when they reach 8 per cent battery, so I will assume scooters that are not available at the start of the day are at 8 per cent battery.</p>
<p>To compare, I kept the same y limit to this graph, as you can see, the average battery count decreased a lot.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566707937/u0Q_bKemZ.png" alt="img13.png" /></p>
<h3 id="heading-my-assumptions-were-incorrect">My assumptions were incorrect</h3>
<p>This is actually still inaccurate, as at the end, let's say a scooter is unlocked with 15 per cent battery left, when it is locked if the battery level is lower than 8 per cent, Voi don't want you to unlock it again, so the scooter will not be available again, in this case, we incorrectly assume the battery is still at 15 per cent because of forward filling, when in fact it should be 8 or less per cent.</p>
<p>I intend to fix it by filling all NaN values at the end with 8 per cent, and backward fill only those values up until they were last unlocked, I believe this will give us a more realistic view of the battery condition.</p>
<p>I am thinking of a shortcut, which is to forward fills the first half and backwards fill the second half, this would be fast, and create mostly what we want, except there will be a jump in the middle</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566713327/iVyVy-39O.png" alt="img14.png" /></p>
<p>as you can see, this is very problematic.</p>
<h3 id="heading-my-assumptions-were-still-incorrect">My assumptions were still incorrect</h3>
<p>I found out that some scooters are disabled while having 50% battery or above, which means they are not disabled because of a lack of battery, but for other reasons, that is why assuming scooters end with 8 per cent battery is so wrong.</p>
<p>My solution is to use the previous day to find the last seen battery count. I first compute the <code>one_day_before_index</code>, which is either one day before the starting index, or 0 when we don't have the data for one day before. Then I first get the columns of data used for that specific day, then I use forward fill on the data plus one day before, then cutting the previous day, so that the data is in the desired range.</p>
<p>final code</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_mean_battery</span>(<span class="hljs-params">starting_index, ending_index</span>):</span>
    one_day_before_index = max(starting_index - <span class="hljs-number">1440</span>, <span class="hljs-number">0</span>)
    starting_diff = starting_index - one_day_before_index
    battery_level_df = store[<span class="hljs-string">'battery_level_df'</span>]
    columns = battery_level_df.iloc[starting_index: ending_index].dropna(axis=<span class="hljs-string">'columns'</span>, how=<span class="hljs-string">'all'</span>).columns
    selected_area_df = battery_level_df[columns].iloc[one_day_before_index: ending_index].fillna(method=<span class="hljs-string">'ffill'</span>)[starting_diff:]
    plt.figure(facecolor=<span class="hljs-string">'white'</span>, figsize=(<span class="hljs-number">15</span>, <span class="hljs-number">8</span>))
    plt.plot(selected_area_df.mean(axis=<span class="hljs-number">1</span>))
    plt.title(<span class="hljs-string">f"Mean battery level over time from <span class="hljs-subst">{selected_area_df.index[<span class="hljs-number">0</span>]}</span> to <span class="hljs-subst">{selected_area_df.index[<span class="hljs-number">-1</span>]}</span>"</span>)
    plt.xlabel(<span class="hljs-string">"Timestamp"</span>)
    plt.ylabel(<span class="hljs-string">"Battery Level"</span>)

plot_mean_battery(<span class="hljs-number">3539</span>, <span class="hljs-number">13166</span>)
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566719212/LmYXO-rc3.png" alt="img15.png" /></p>
<h3 id="heading-other-days">Other days</h3>
<p>Now I run this function on other days, and these are the result.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566723672/L_ybogzVx.png" alt="img16.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566724659/Y3gr-vJRo.png" alt="img17.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566725815/-IpCIMYOi.png" alt="img18.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566727178/ZUfm1e24-.png" alt="img19.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566728508/VcBvxZD3Q.png" alt="img20.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566729988/xxeBjtO6d.png" alt="img21.png" />
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566731009/3dtJtNo3x.png" alt="img22.png" /></p>
<h3 id="heading-entire-week">Entire week</h3>
<p>For the entire week, from 6/6 Monday to 12/6 Sunday.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657566746235/TfIv2XEni.png" alt="img23.png" /></p>
<h1 id="heading-whats-next">What's next</h1>
<p>I may try to explain the result of the average battery count, I may also try to start looking into the scooter locations.</p>
]]></content:encoded></item><item><title><![CDATA[Where's my Voi scooter: [6] Changing the data storage format]]></title><description><![CDATA[In this blog post I will first fix the data collection program, and then change how the data is stored, to facilitate further analysis.
Improving the data collection program
The data collection program crashed multiple times. Since data cannot be col...]]></description><link>https://blog.cpbprojects.me/wheres-my-voi-scooter-6-changing-the-data-storage-format</link><guid isPermaLink="true">https://blog.cpbprojects.me/wheres-my-voi-scooter-6-changing-the-data-storage-format</guid><category><![CDATA[Voi]]></category><category><![CDATA[Python]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Data Science]]></category><category><![CDATA[data structures]]></category><dc:creator><![CDATA[Chit]]></dc:creator><pubDate>Thu, 07 Jul 2022 14:01:13 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/unsplash/Wpnoqo2plFA/upload/v1657202339643/4zyHEDEve.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In this blog post I will first fix the data collection program, and then change how the data is stored, to facilitate further analysis.</p>
<h1 id="heading-improving-the-data-collection-program">Improving the data collection program</h1>
<p>The data collection program crashed multiple times. Since data cannot be collected while the program is crashed, I have a lot of missing data, which makes data analysis difficult. </p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657202165213/iLbpn7-XA.png" alt="img1.png" /></p>
<p>I need a try-except block to catch exceptions instead of letting the program crash. I should have done this long ago, but I was lazy. Now I need to pay the price.</p>
<p>In the main loop, I first ask for a response, then store the response, and wait 60 seconds before making another request, the two possible errors are:</p>
<h3 id="heading-failing-to-get-a-response-from-the-server">Failing to get a response from the server</h3>
<p>When I make a request to the API server, it has happened that it refuses to give me an access token, or response. In this case, I need to cancel future actions and make a request maybe 5 seconds later.</p>
<h3 id="heading-getting-an-invalid-response-from-the-server">Getting an invalid response from the server</h3>
<p>If the response is not what I expected, I will simply not parse the response. and wait for the next iteration.</p>
<h2 id="heading-the-fix-of-the-data-collection-program">The fix of the data collection program</h2>
<p>The following is the before and after of the main loop. I used two try-except blocks to wrap up the two functions that may crash. In the first block, if the program runs into an exception, I will print out what is the problem, and continue the loop after 5 seconds. In the second block, I simply print out the problem and do nothing. So now if something bad happens, the program will not crash, but instead, wait and try again later. In retrospect, this fix is so easy that I should have done this earlier, so then I don't have to deal with so much missing data.</p>
<p>Before:</p>
<pre><code class="lang-python"><span class="hljs-keyword">while</span> <span class="hljs-literal">True</span>:
    response = make_api_requests.get_scooter_locations()
    handle_response.store_response(response)
    time.sleep(<span class="hljs-number">60</span>)
</code></pre>
<p>After:</p>
<pre><code class="lang-python"><span class="hljs-keyword">while</span> <span class="hljs-literal">True</span>:
    <span class="hljs-keyword">try</span>:
        response = make_api_requests.get_scooter_locations()
    <span class="hljs-keyword">except</span>:
        print(<span class="hljs-string">"Failed to get scooter locations, retrying in 5 seconds"</span>)
        time.sleep(<span class="hljs-number">5</span>)
        <span class="hljs-keyword">continue</span>
    <span class="hljs-keyword">try</span>:
        handle_response.store_response(response)
    <span class="hljs-keyword">except</span>:
        print(<span class="hljs-string">"Failed to store response, will skip storing the response"</span>)
    time.sleep(<span class="hljs-number">60</span>)
</code></pre>
<h1 id="heading-moving-the-project-to-a-jupyter-notebook">Moving the project to a Jupyter notebook</h1>
<p>From my other <a target="_blank" href="https://chit.hashnode.dev/what-i-learned-from-the-10-hour-data-science-course-data-analysis-with-python-course-numpy-pandas-data-visualization">blog post</a>, I discovered about Jupyter notebook and decided to use it to complete the analysis. As a JetBrains user, I searched for tools from Jetbrains that support Jupyter notebooks. I found <a target="_blank" href="https://www.jetbrains.com/dataspell/">DataSpell</a>. It seems useful for me as it can:</p>
<ul>
<li>automatically starts a Jupiter server when I start running the code </li>
<li>allow me to view dataframes and Jupyter notebook variables</li>
<li>be used for free because of the <a target="_blank" href="https://education.github.com/pack">Github Student Developer plan</a>, so why not</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657202184157/2dSKxq5TY.png" alt="img2.png" /></p>
<h2 id="heading-importing-and-checking-data">Importing and checking data</h2>
<p>The first thing to do is to sanitise the data, I will have to iterate the data files. I copied the functions from my previous blog post.</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> pathlib <span class="hljs-keyword">import</span> Path

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">get_sorted_json_list</span>():</span>
    path_list = sorted(Path(<span class="hljs-string">'location_data/'</span>).glob(<span class="hljs-string">'*.json'</span>))
    <span class="hljs-keyword">return</span> path_list
</code></pre>
<p>Then I print all files that are not valid json, because if my data-gathering program ran correctly, all data should be valid json, so I want to double-check before doing any analysis.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> json

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">check_all_files</span>():</span>
    sorted_file_list = get_sorted_json_list()
    <span class="hljs-keyword">for</span> path <span class="hljs-keyword">in</span> sorted_file_list:
        <span class="hljs-keyword">with</span> open(path, <span class="hljs-string">'rb+'</span>) <span class="hljs-keyword">as</span> f:
            <span class="hljs-keyword">try</span>:
                json_data = json.load(f)
            <span class="hljs-keyword">except</span> ValueError:
                print(<span class="hljs-string">f"file <span class="hljs-subst">{path.name}</span> is not valid json"</span>)
    print(<span class="hljs-string">"all file ok"</span>)
</code></pre>
<p>Two files showed up as invalid json, so I fixed the file and tried again, this time all are valid</p>
<h2 id="heading-creating-numpy-array">Creating numpy array</h2>
<p>In the previous blog post in this series, I plotted the graph of vehicle counts, now I want to redo it with the new data I gathered. I first created the numpy array for timestamps and vehicle counts.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> numpy <span class="hljs-keyword">as</span> np
<span class="hljs-keyword">import</span> pandas <span class="hljs-keyword">as</span> pd

timestamp_np_array = np.empty(MAX_DATA_SIZE, dtype=pd.Timestamp)
vehicle_count_np_array = np.empty(MAX_DATA_SIZE, dtype=int)
counter = <span class="hljs-number">0</span>
</code></pre>
<p>Then I define the function that gets the vehicle count from a file, it opens the file and fills the two numpy arrays with data. I used the  <code>global</code> keyword here so that I don't have to worry about scopes.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> traceback

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">get_vehicle_count_list_from_file</span>(<span class="hljs-params">path</span>):</span>
    <span class="hljs-keyword">global</span> timestamp_np_array, vehicle_count_np_array, counter
    <span class="hljs-keyword">with</span> open(path, <span class="hljs-string">'r'</span>) <span class="hljs-keyword">as</span> f:
        <span class="hljs-keyword">try</span>:
            json_data = json.load(f)
            <span class="hljs-keyword">for</span> individual_data <span class="hljs-keyword">in</span> json_data:
                <span class="hljs-keyword">if</span> counter &lt; MAX_DATA_SIZE:
                    timestamp_np_array[counter] = pd.to_datetime(individual_data[<span class="hljs-string">'time_stamp'</span>])
                    vehicle_count_np_array[counter] = individual_data[<span class="hljs-string">'vehicle_count'</span>]
                    counter += <span class="hljs-number">1</span>

        <span class="hljs-keyword">except</span> ValueError:
            traceback.print_exc()
            print(<span class="hljs-string">f"file <span class="hljs-subst">{path.name}</span> is not valid json"</span>)
</code></pre>
<p>Then I execute the operation, truncate the two numpy arrays so they are only as long as the amount of data we have, and then save them so they can be loaded at another time, I am not going to work non-stop on this project, so I need to be able to load the data at another time.</p>
<pre><code class="lang-python">
<span class="hljs-keyword">for</span> path <span class="hljs-keyword">in</span> get_sorted_json_list():
    get_vehicle_count_list_from_file(path)

timestamp_np_array = timestamp_np_array[:counter]
vehicle_count_np_array = vehicle_count_np_array[:counter]

print(counter)

np.save(<span class="hljs-string">'cached_data/timestamp_np_array'</span>, timestamp_np_array)
np.save(<span class="hljs-string">'cached_data/vehicle_count_np_array'</span>, vehicle_count_np_array)
</code></pre>
<h1 id="heading-plotting-the-data">Plotting the data</h1>
<p>I now plot the graph. I stated how long, because I may not want to plot the entire graph, I also have a step variable so that the graph will not be too packed. There are 39356 data items to fit inside a graph, so not every point has to be considered</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> matplotlib.pyplot <span class="hljs-keyword">as</span> plt

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">plot_vehicle_count_graph</span>():</span>
    timestamp_np_array = np.load(<span class="hljs-string">'cached_data/timestamp_np_array.npy'</span>, allow_pickle=<span class="hljs-literal">True</span>)
    vehicle_count_np_array = np.load(<span class="hljs-string">'cached_data/vehicle_count_np_array.npy'</span>, allow_pickle=<span class="hljs-literal">True</span>)

    fig, ax = plt.subplots(figsize=(<span class="hljs-number">15</span>,<span class="hljs-number">8</span>), subplot_kw={<span class="hljs-string">"title"</span>: <span class="hljs-string">"Number of available scooters over time"</span>,
                                                       <span class="hljs-string">"xlabel"</span>: <span class="hljs-string">"timestamp"</span>,
                                                       <span class="hljs-string">"ylabel"</span>: <span class="hljs-string">"Number of scooters available"</span>})
    fig.patch.set_facecolor(<span class="hljs-string">'white'</span>)
    starting_time = <span class="hljs-number">0</span>
    how_long = <span class="hljs-number">39356</span>
    step = <span class="hljs-number">30</span>
    ax.plot(timestamp_np_array[starting_time:starting_time+how_long:step],
            vehicle_count_np_array[starting_time:starting_time+how_long:step])
    plt.show()

plot_vehicle_count_graph()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657202206190/7s-lDx5NZ.png" alt="img3.png" /></p>
<p>Note that I put everything in a function, this is so that when I come back to the project, and I need to run previous code blocks to import everything I need, I can do that without doing everything. I could've also separated all the imports and put them at the start of the Jupyter notebook.</p>
<h1 id="heading-placing-data-in-dataframe">Placing data in dataframe</h1>
<p>I decided that it is important to have the data in tabular form, because when I investigate battery consumption, having data in tabular form makes it easier to track individual scooters. As it is just going down a column, compared to digging through multiple files.</p>
<p>I precompute everything and put it inside a dataframe, instead of going through the files every time I do the analysis for the scooters.</p>
<p>In that case, I will need to know how many columns and rows there will be, as in the pandas' documentation when I create a dataframe, I need to specify the columns and index.</p>
<h2 id="heading-generating-all-vehicles-set">Generating all vehicles set</h2>
<p>We iterate through all the files, then we use list comprehension to accumulate all the ids, and we use the union operation to combine the set with the total set.</p>
<pre><code class="lang-python">vehicle_id_set = set()

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">all_vehicle_id_for_a_single_json_file</span>(<span class="hljs-params">path</span>):</span>
    <span class="hljs-keyword">global</span> vehicle_id_set
    <span class="hljs-keyword">try</span>:
        <span class="hljs-keyword">with</span> open(path, <span class="hljs-string">'r'</span>) <span class="hljs-keyword">as</span> f:
            json_data = json.load(f)
        <span class="hljs-keyword">for</span> time_data <span class="hljs-keyword">in</span> json_data:
            new_set = set([vehicle_data[<span class="hljs-number">0</span>] <span class="hljs-keyword">for</span> vehicle_data <span class="hljs-keyword">in</span> time_data[<span class="hljs-string">'vehicle_data'</span>]])
            vehicle_id_set = vehicle_id_set.union(new_set)
    <span class="hljs-keyword">except</span>:
        traceback.print_exc()
        print(<span class="hljs-string">f"In all_vehicle_id_for_a_single_json_file: file <span class="hljs-subst">{path.name}</span> is not valid json"</span>)

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">add_all_vehicle_id_to_set</span>():</span>
    <span class="hljs-keyword">global</span> vehicle_id_set

    <span class="hljs-comment"># Iterate through all files</span>
    sorted_json_list = get_sorted_json_list()
    <span class="hljs-keyword">for</span> json_path <span class="hljs-keyword">in</span> sorted_json_list:
        all_vehicle_id_for_a_single_json_file(json_path)

add_all_vehicle_id_to_set()
</code></pre>
<h2 id="heading-generating-all-timestamp-list">Generating all timestamp list</h2>
<p>Now I find out all the timestamps to use them as indexs.</p>
<pre><code class="lang-python">datetime_np_array = np.empty(shape=(<span class="hljs-number">39356</span>,), dtype=pd.Timestamp)
datetime_np_array_counter = <span class="hljs-number">0</span>

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">get_all_dates</span>():</span>
    <span class="hljs-keyword">global</span> datetime_np_array_counter
    sorted_json_list = get_sorted_json_list()
    <span class="hljs-keyword">for</span> json_path <span class="hljs-keyword">in</span> sorted_json_list:
        <span class="hljs-keyword">try</span>:
            <span class="hljs-keyword">with</span> open(json_path, <span class="hljs-string">'r'</span>) <span class="hljs-keyword">as</span> f:
                json_data = json.load(f)
            <span class="hljs-keyword">for</span> time_data <span class="hljs-keyword">in</span> json_data:
                datetime_np_array[datetime_np_array_counter] = pd.to_datetime(time_data[<span class="hljs-string">'time_stamp'</span>])
                datetime_np_array_counter += <span class="hljs-number">1</span>
        <span class="hljs-keyword">except</span>:
            traceback.print_exc()
            print(<span class="hljs-string">f"In get_all_dates: file <span class="hljs-subst">{json_path.name}</span> is not valid json"</span>)
</code></pre>
<h2 id="heading-actually-creating-the-dataframe">actually creating the dataframe</h2>
<p>Now I create a dataframe with all the vehicle_id as columns, and the number of collected data as rows, this will be a huge dataframe. This is rather straightforward, I just call their constructor with the index, columns and data type specified.</p>
<pre><code class="lang-python">battery_level_df = pd.DataFrame(index=datetime_np_array, columns=vehicle_id_set, dtype=np.int64)
longitude_df = pd.DataFrame(index=datetime_np_array, columns=vehicle_id_set, dtype=np.float64)
latitude_df = pd.DataFrame(index=datetime_np_array, columns=vehicle_id_set, dtype=np.float64)
</code></pre>
<h2 id="heading-filling-the-dataframe">filling the dataframe</h2>
<p>Now that I have this empty dataframe, I need to fill it, I will once again iterate through the json_files, and put the correct data in the correct place
when filling in the fields, I can either
1 .assume the order of data is correct and just fill them in one row by row without checking the index</p>
<ol>
<li>check the index every time, which is more expensive
I think I will check the index, just to be safe</li>
</ol>
<pre><code class="lang-python">%%time

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">fill_all_df</span>(<span class="hljs-params">battery_level_df, longitude_df, latitude_df</span>):</span>
    <span class="hljs-keyword">for</span> json_path <span class="hljs-keyword">in</span> get_sorted_json_list():
        print(<span class="hljs-string">f"doing <span class="hljs-subst">{json_path.name}</span>"</span>)
        <span class="hljs-keyword">with</span> open(json_path, <span class="hljs-string">'r'</span>) <span class="hljs-keyword">as</span> f:
            json_data = json.load(f)
        <span class="hljs-keyword">for</span> time_data <span class="hljs-keyword">in</span> json_data:
            current_row_datetime = pd.to_datetime(time_data[<span class="hljs-string">'time_stamp'</span>])
            <span class="hljs-keyword">for</span> vehicle_data <span class="hljs-keyword">in</span> time_data[<span class="hljs-string">'vehicle_data'</span>]:
                battery_level_df.loc[current_row_datetime, vehicle_data[<span class="hljs-number">0</span>]] = vehicle_data[<span class="hljs-number">1</span>]
                longitude_df.loc[current_row_datetime, vehicle_data[<span class="hljs-number">0</span>]] = vehicle_data[<span class="hljs-number">2</span>]
                latitude_df.loc[current_row_datetime, vehicle_data[<span class="hljs-number">0</span>]] = vehicle_data[<span class="hljs-number">3</span>]

fill_all_df(battery_level_df, longitude_df, latitude_df)
</code></pre>
<p>I used the <code>%%time</code> to measure how long this operation took, and the result was Wall time: 1h 35min 49s, this proofs that this operation uses a lot of computing resources.</p>
<h2 id="heading-storing-the-data">storing the data</h2>
<p>Now I wish to store the result of the computation, I want it to store quickly, and load quickly, according to this <a target="_blank" href="http://matthewrocklin.com/blog/work/2015/03/16/Fast-Serialization">article</a> from Matthew Rocklin, the cost of serializing text-based data and numeric data is compared, and HDFS store seems to perform well in numeric data, so I decided to use it.</p>
<pre><code class="lang-python">store = pd.HDFStore(<span class="hljs-string">"store.h5"</span>)
store[<span class="hljs-string">'battery_level_df'</span>] = battery_level_df
store[<span class="hljs-string">'longitude_df'</span>] = longitude_df
store[<span class="hljs-string">'latitude_df'</span>] = latitude_df
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1657202219288/WyvFQQk8R.png" alt="img4.png" /></p>
<h1 id="heading-whats-next">What's next</h1>
<p>I intended to talk about the results of my data analysis in this blog post, but this is already very long, so in the next one, I will use the dataframes to investigate why the number of scooters decreases throughout the day, and other interesting things.</p>
]]></content:encoded></item></channel></rss>