
Free Keyword Research Tool: The Free Stack That Works for Service Businesses
Which free keyword research tool should a Singapore clinic, firm, contractor or tutor use? Combine five free tools to find your first 30-50 keywords. See how.
From F&B to fintech, clinics to law firms, startups to enterprise. If your customers search on Google, we make sure they find you first, not your competitors.
One specialist team, focused only on the organic rankings that put you in front of ready-to-buy Singapore customers.
A clear, sequenced path from audit to rankings. You always know what we’re doing and why it matters for your leads.

Quick answer: Log file analysis SEO Singapore sites benefit from means reading your server’s own record of every request Googlebot made. It shows which URLs were actually crawled, how often, and what response each returned. Crawl tools simulate a search engine; logs show what the search engine truly did.
Every SEO tool you have ever used guesses. A crawler pretends to be a search engine and reports what a search engine would probably find. A rank tracker samples results and infers a position. Search Console summarises, aggregates and rounds. All of these are useful and all of them are interpretations. Your server logs are the only source in the entire discipline that is a record rather than an estimate. They are your web server writing down, line by line, every single request it received: who asked, for what, when, and what it sent back. That includes every visit from Googlebot. Nothing in SEO is closer to the truth than that file, and almost nobody outside enterprise teams ever opens it. This post explains what a log file is, the specific things it tells you that no other tool can, how to obtain yours from a typical Singapore hosting setup, and how to read one without being a developer. It is the most technical topic in this series and it is deliberately written for an owner rather than an engineer. The wider context sits within our technical SEO practice.
Every time anyone or anything requests a page from your website, your web server writes a line to a text file. That line typically contains the requesting IP address, the date and time, the exact URL requested, the HTTP status code returned, the number of bytes sent, and the user agent string, which is how the requester identifies itself.
A single line looks intimidating and means something simple. Read in plain English, it says: at this moment, this visitor asked for this address, and the server replied with this outcome. Multiply that by every request your site receives and you have a complete attendance record.
Two clarifications for owners who are used to analytics dashboards. Logs are not analytics. Google Analytics runs a piece of JavaScript in a visitor’s browser, which means it records humans with JavaScript enabled and misses bots almost entirely. Logs record everything that touched the server, including bots, blocked requests and failed responses. They are complementary sources answering different questions.
Logs are not Search Console’s crawl stats either. The Crawl Stats report in Search Console is a summary derived from Google’s side of the conversation, aggregated and sampled, with limited ability to drill into specific URLs. Your logs are unaggregated and unrounded. If a URL was requested at a specific second, the line is there.
The practical consequence is that logs answer questions no other tool can answer definitively, and we will spend the rest of this post on exactly which questions those are.
One, whether a page has ever been crawled at all. A crawl tool can tell you a URL exists and is linked. Only logs tell you whether Googlebot has actually requested it, and when. A page published three months ago with zero Googlebot requests in the log is not underperforming, it is undiscovered, and those are entirely different problems with different fixes.
Two, crawl frequency by section. Logs show that your blog is requested daily while your service pages are requested every eleven days and your case studies once a month. That distribution reflects how search engines weigh the sections of your site, and it is invisible everywhere else. When we see commercially critical pages in the low frequency tier, that is a finding worth acting on.
Three, crawl waste. This is the most common revelation. Logs regularly show a large share of crawler requests going to URLs with tracking parameters, filtered views, paginated archives, feed URLs, old redirects and non existent pages. A crawl tool will not tell you that search engines are spending their time there, because the crawl tool is not a search engine.
Four, errors search engines actually encountered. Your site may return a 200 status when you check it and a 500 at two in the morning when a backup job runs and the server is under load. Logs record what Googlebot received at the moment it asked, not what you receive when you check on a Tuesday afternoon.
We have used this specifically to settle arguments about whether a page was ever seen. When a set of pages has been live for months without ranking, the logs can show that crawlers have barely requested them, which turns a long content debate into a quick internal linking fix. Our insurance sector results show why that matters for content depth, with 10 financial guide articles and 6 product education pages built during the engagement, all of which need to be found before they can rank.
Five, the true impact of a change. Publish a new section, submit a sitemap, or restructure internal links, and logs show precisely when crawlers began requesting the new URLs and at what rate. Instead of waiting for rankings to move and guessing at causation, you can watch discovery happen within days. For large catalogue sites this is the difference between managing and hoping, and it is a standard part of how we work on e-commerce SEO projects above a certain size.
Access varies considerably by hosting arrangement, and this is usually the hardest step of the whole exercise rather than the analysis itself.
| Hosting type | Typical access route | Retention | Difficulty |
|---|---|---|---|
| Shared hosting (cPanel) | Raw Access Logs in the control panel | Often 24 hours to 7 days | Easy but short retention |
| Managed WordPress host | Support request or dashboard export | Usually 7 to 30 days | Easy, may need a ticket |
| VPS or dedicated server | Direct file access via SSH | As configured, often 30 days | Needs technical help |
| Cloudflare or CDN in front | Logs at the CDN layer, plan dependent | Varies by plan | May require a paid tier |
| Shopify | Not available to merchants | Not applicable | Use Search Console instead |
| Wix or Squarespace | Not available | Not applicable | Use Search Console instead |
Three Singapore specific notes. First, if a content delivery network sits in front of your site, requests served from the CDN cache may never reach your origin server, so origin logs will understate crawler activity. Get the logs from the CDN layer instead. Second, default log retention on shared hosting plans is frequently very short, and the fix is to ask your host to enable archiving before you need the data, not after. Third, personal data protection obligations apply here, since log files contain IP addresses. Store them securely, limit who has access, and delete them when the analysis is done.
Ask your host for at least thirty days of logs covering all hostnames, in raw or combined format. Fourteen days is workable for a first look. Under seven days tells you very little, because crawl patterns on a typical SME site operate on a weekly and monthly rhythm.
Before you draw a single conclusion, filter out the impostors. Any script can claim to be Googlebot in its user agent string, and a meaningful share of the traffic identifying itself that way in a typical log is something else entirely: scrapers, competitive intelligence tools, or outright malicious crawlers harvesting content and contact details.
The verification method is a reverse DNS lookup on the requesting IP address, which should resolve to a googlebot.com or google.com hostname, followed by a forward lookup on that hostname to confirm it returns the original IP. Google also publishes IP ranges for its crawlers that can be matched against directly, and most log analysis software does this automatically.
Skipping verification produces confidently wrong conclusions. We have reviewed logs where a substantial portion of apparent Googlebot activity failed verification entirely, and an analysis based on the unfiltered data would have suggested a crawl budget problem that did not exist. Verify first, then analyse.
While you are there, note the split between the smartphone and desktop Googlebot user agents. Since indexing is mobile first, the overwhelming majority of legitimate crawling should come from the smartphone agent. A heavy desktop skew is unusual and worth a question.
You do not need to learn command line tools, though they help. A spreadsheet handles a small site’s logs perfectly well, and dedicated log analysis software handles larger ones. Here is the sequence that produces useful answers fastest.
Filter to verified search engine crawlers only. Remove human traffic and unverified bots. What remains is your subject.
Count requests by status code. You are looking at the ratio. A healthy site shows the large majority as 200 responses, a modest share of 301 redirects, and very few 404s. A large share of 404s means search engines are repeatedly requesting URLs that do not exist, which is both wasteful and a signal that something still links to them. Any 5xx errors at all warrant investigation, because server errors during crawling directly suppress crawling afterwards.
Group requests by URL folder. Compare the share of crawl attention each section receives against the commercial value of that section. This single view produces more actionable findings than any other, because it makes crawl waste visible as a proportion.
Find URLs crawled most and least. The most crawled list often contains surprises: a feed URL, a parameterised page, a legacy directory. The least crawled list, when compared against your sitemap, reveals what is effectively undiscovered.
Compare logs against your sitemap. URLs in your sitemap that never appear in logs are discovery failures. URLs in your logs that are not in your sitemap are either legitimate pages missing from it or junk that should not be crawlable. Both lists are actionable.
Plot requests per day over time. A sudden drop usually follows a server problem, a robots.txt change or a redirect error. A sudden rise usually follows a structural change or a proliferation of new URLs.
Crawl budget means the number of URLs a search engine is willing to request from your site in a given period. It is determined by how much load your server can handle without slowing down and how much demand the search engine has for your content.
Here is where conventional wisdom gets it wrong, and it is worth being direct about. For a Singapore SME site with a few hundred pages, crawl budget is almost never your problem. Google will happily crawl a small, fast, well structured site far more thoroughly than it needs to. Spending money optimising crawl budget on a 150 page site is wasted effort, and any provider selling it as a priority for a site that size is selling the wrong thing.
Crawl budget becomes genuinely relevant above roughly ten thousand URLs, or on any site where parameter and filter combinations have multiplied the URL count well beyond the real page count, or where the server is slow enough that crawling is being throttled. Catalogue driven retailers, property portals, classifieds sites and large publishers hit this. A professional services firm with forty pages does not.
That said, log analysis remains useful below that threshold for entirely different reasons: finding undiscovered pages, spotting intermittent server errors, verifying that a migration was crawled cleanly, and confirming that bot traffic is what it claims to be. The value of logs is not limited to crawl budget, and conflating the two is why most small sites never look at them.
Findings map to a fairly small set of responses, which makes this phase quicker than it sounds.
If important pages are rarely or never crawled, the cause is usually discovery or importance rather than budget. Increase internal links pointing at them from pages that are crawled frequently, ensure they are in the sitemap, and reduce their click depth from the homepage.
If crawl attention is concentrated on parameter and filter URLs, handle parameters properly with canonical tags and server level rules. Do not block them in robots.txt, because that prevents the canonical from ever being read.
If 404 responses are a significant share, find what links to those URLs. Frequently it is an old sitemap, a stale internal link, or an external site linking to a page you moved without a redirect. Each of those has a different fix and all three are worth doing.
If 5xx errors appear in clusters at particular times, correlate with scheduled tasks, backups or traffic peaks. This is a hosting conversation rather than an SEO one, and it is usually resolved by a plan upgrade or a scheduling change.
If redirect chains appear in the logs, shorten them to a single hop and update the internal links that point at the old address so the redirect stops being requested at all.
Inventory driven businesses get the most out of this list, because their URL counts move constantly as stock turns over. We recommend a monthly log review for any site republishing inventory at that pace, and it is a rhythm informed by our car dealer SEO work.
It particularly suits inventory-led sites like the one behind our used car dealer results, where 14 make and model pages also show the dealer’s current inventory in each model.
For lead generation businesses where a small number of pages carry all the commercial weight, the section level crawl distribution view is usually the single most valuable output, and it is a lens we apply across finance and education sector work where deep pages often matter more than the homepage.
Faceted and filtered URLs are where crawl attention most often leaks away. In our ecommerce case study, a WooCommerce home and lifestyle store had only 34% of its product pages indexed. Phase 1 rewrote robots.txt to block the 14 faceted navigation parameter combinations generating duplicate content and submitted a clean XML sitemap covering all canonical product and category URLs. By the end of Month 2, product page indexation had moved from 34% to 79%. Log analysis is how you confirm whether the same leak is happening on your site.
Field notes: The finding we see most often is misallocation rather than shortage: crawlers spending their time on parameterised, sorted or paginated variants rather than the pages the business wants found, and deep pages with few internal links being requested rarely, if at all. Both problems were present in our ecommerce case study. Beyond the Phase 1 crawl repair, Phase 4 rebuilt the internal linking architecture to route PageRank from the site’s most-linked blog content to the 15 target category pages, added breadcrumb schema sitewide and moved home page link equity away from brand and policy pages toward top-revenue categories. By Month 9, product indexation was at 95% and two category pages had reached #1. We also routinely find traffic claiming to be Googlebot that fails reverse DNS verification, so always verify before drawing conclusions.
Log files convert SEO from a discipline of inference into one of observation. Every other tool in the stack is telling you what a search engine would probably do with your site. Your server is telling you what it actually did, at what time, and with what result. That distinction matters most in the two situations where teams waste the most money: when a page is not ranking and nobody has checked whether it has ever been crawled, and when a site is large enough that crawler attention is a genuinely scarce resource being spent in the wrong places. Get thirty days of logs, verify the bots, group the requests by folder and by status code, and compare the result against your sitemap. That is four steps and it will tell you more about how search engines regard your site than a month of rank tracking. If your host will not give you logs, that is worth knowing too, and it is a reasonable thing to ask before you renew. Should you want help reading the output, our audit and consulting work covers it, and a short conversation will tell you whether your site is large enough to justify the exercise.
We’ve seen a page that looked perfectly healthy in every crawler turn out to be visited by Googlebot once a month, which no amount of on-page work was ever going to fix, because the problem was never content. Our clients who adopt log analysis usually do it after a ranking mystery that ordinary tools could not explain, and most of those mysteries turn out to be a crawl pattern rather than a content gap.
It is a plain text file your web server writes automatically, adding one line for every request it receives. Each line records the requesting IP address, the timestamp, the URL requested, the HTTP status code returned, the bytes transferred and the user agent string identifying the requester. It captures everything, including search engine crawlers, which is what makes it valuable for SEO in a way analytics never can be.
Search Console gives you a summarised, sampled view from Google’s side, with limited ability to examine individual URLs or specific moments. Logs are the unaggregated record from your side, covering every crawler rather than just Google’s, with exact timestamps and exact responses. Search Console tells you the shape of crawling; logs tell you the detail of it.
Not routinely. A site with fifty to two hundred pages will rarely have a crawl budget problem, and Search Console covers most of what you need. Logs still earn their place in specific situations: after a migration, when pages are not being indexed and you need to know whether they have been crawled at all, or when you suspect intermittent server errors that do not appear during your own checks.
You generally cannot. These platforms do not expose server logs to merchants or site owners. That is not a reason to avoid the platform, but it does mean Search Console’s crawl stats and coverage reports are your ceiling for this kind of insight. If log level visibility is genuinely important to your business model, that is a factor to weigh when choosing a platform.
Thirty days is the comfortable target because it captures weekly and monthly crawl rhythms. Fourteen days gives you a usable first look. Anything under seven days is too short to distinguish a pattern from a coincidence. Many Singapore shared hosting plans retain only a few days by default, so ask your host to enable longer retention before you need it.
Crawl budget is how many URLs a search engine will request from your site in a given period, set by your server’s capacity and the engine’s interest in your content. It becomes a real constraint above roughly ten thousand URLs, or wherever filters and parameters have multiplied your URL count well past your real page count. Below that, it is rarely the limiting factor and should not be a spending priority.
Because any script can put Googlebot in its user agent string, and many do. Scrapers and competitive tools routinely impersonate search engine crawlers. If you analyse unverified data, you may conclude that crawling is heavy and budget constrained when much of the activity is not search engines at all. Verification is done through a reverse DNS lookup, and good log analysis software handles it automatically.
Log files contain IP addresses and can therefore engage personal data obligations. Treat them as you would any other dataset containing personal information: restrict access, store them securely rather than in a shared drive, do not send them to third parties without a clear basis, and delete them when the analysis is complete. If you engage an external provider, confirm how they will handle and dispose of the files.
Crawl misallocation. On larger sites, a substantial share of crawler requests goes to parameterised, filtered, sorted or paginated URLs rather than the pages that generate revenue. On smaller sites, the common finding is the opposite: pages that appear in the sitemap but receive no crawler requests at all, almost always because they sit too deep in the structure with too few internal links.
Sometimes, and it is one of the faster ways to rule causes in or out. If logs show server errors, a crawl volume collapse or a sudden shift in what is being requested around the date of the drop, you have a technical explanation to pursue. If crawling looks entirely normal through the period, you can confidently redirect your investigation towards content, competition or a search engine update instead of chasing technical ghosts.
Most Singapore businesses have never looked at a single line of their own server logs, and for smaller sites that is a perfectly reasonable decision. If your site runs to thousands of URLs, or if you have pages that simply will not index and nobody can say why, it stops being reasonable. We will take a look at a sample of your logs at no cost and tell you honestly whether a full analysis is worth commissioning or whether your budget belongs somewhere else entirely. Send us a note through the contact page with your platform and a rough page count, and we will let you know either way.
Natalie leads SEO strategy at Singapore SEO Agency, helping local and regional businesses build organic search programmes that drive qualified leads. She specialises in technical SEO and content-led authority building for Singapore SMEs.
Get a free SEO audit for your Singapore website — we'll show you exactly where you stand, what's holding you back, and what it would take to rank on page 1.
Get Your Free SEO Audit →
Which free keyword research tool should a Singapore clinic, firm, contractor or tutor use? Combine five free tools to find your first 30-50 keywords. See how.

SEO vs SEM is usually the wrong question. Learn what the terms really mean and how to use ads and Search Console data to decide which searches to earn or buy.

Google Search Console login problems usually start with verification and ownership. Learn how to get in, fix access errors and offboard agencies safely.

The click through rate formula is clicks divided by impressions. Learn what each platform counts, the averaging trap and how to set it up in Google Sheets.

A click through rate means nothing on its own. Learn to compare CTR by position, query type and SERP features, and see what low CTR is really telling you.

What a google analytics certification proves, what it misses, how to verify one and the practical questions to ask before you hire a marketer or freelancer.
Fast, no obligation. We reply within 24 hrs.
Singapore’s specialist SEO agency for SMEs. We rank your business on Google — and only Google. No distractions, just results.
© 2026 Singapore SEO Agency. All rights reserved.