The 42-Point SEO Audit Checklist — Published in Full, Including the Six Nobody Else Runs
This is the checklist behind our ₹12,500 audit, published rather than gated. If you have a developer and a free afternoon, you can run most of it yourself — and you should, before you pay anybody.
In short
Forty-two checks in six sections: crawl and access, indexation and duplication, structure and linking, rendering and speed, structured data, and AI crawler access. The last six are the ones almost no Chennai audit includes. Running it yourself takes a focused afternoon; we charge ₹12,500 to do it in 10 working days.
Reviewed 20 September 2026 · revised twice a year; the AI-crawler section is reviewed quarterly as user agents and tokens change
Why publish the thing we sell
Because the checklist is not the product. Anyone can read a list of 42 checks; what takes ten working days is running them properly on a real site, evidencing each finding with a URL or a screenshot, and — the part that actually matters — ordering the fixes by effort against impact so a developer can start on Monday. The list is the easy half and gating it would be a marketing decision, not an honest one.
There is a second reason, and we will state it because this site publishes its reasoning: a page like this earns citations. Procedural content carries a 2.8× schema multiplier with HowTo markup and how-to queries average 5.1 AI citations against 3.1 for commercial pages. Publishing the checklist is both the right thing and the profitable thing, which is a comfortable position to be in and worth being transparent about.
The 42 checks
Six sections. Run them in order — there is no point optimising a page that is not being crawled.
| Section | # | Check | Why it matters |
|---|---|---|---|
| A. Crawl and access | 1 | robots.txt does not block anything that earns money | A staging rule pushed live is the single most expensive one-line mistake in SEO |
| 2 | XML sitemap lists only canonical, indexable, 200-status URLs | A sitemap full of redirects and noindex pages teaches a crawler to distrust it | |
| 3 | No orphan pages — every important URL is linked from somewhere | If your own site does not link to it, nothing else will find it either | |
| 4 | Money pages sit within three clicks of the homepage | Crawl depth is a proxy for how important your own architecture says a page is | |
| 5 | No 5xx errors in the last 30 days of server logs | Intermittent server errors quietly reduce crawl rate for weeks afterwards | |
| 6 | Soft 404s identified — pages returning 200 with nothing on them | Common on empty category and filter pages; they dilute everything around them | |
| 7 | No redirect chains or loops | Every hop loses time and some signal, and loops lose the page entirely | |
| B. Indexation and duplication | 8 | Indexed page count roughly matches the number of pages you meant to publish | A 40,000-URL index on a 600-page site means parameters are being crawled |
| 9 | No noindex or nofollow on a page that is supposed to rank | Usually a leftover from a rebuild; costs months before anyone notices | |
| 10 | Canonicals are self-referencing and consistent | Conflicting canonical and sitemap signals are resolved by the engine, not by you | |
| 11 | Parameter and faceted URLs are controlled deliberately | The commonest source of index bloat in ecommerce | |
| 12 | No two pages target the same query | Cannibalisation splits signal. Two of the largest Chennai agencies do this to themselves | |
| 13 | Paginated sets handled with a stated approach | “Load more” without URLs makes page two invisible | |
| 14 | Internal search result pages are not indexable | They generate infinite low-value URLs from nothing | |
| C. Structure and linking | 15 | Exactly one H1 per page | Three of the twelve largest Chennai agency sites fail this on their own money pages |
| 16 | Heading hierarchy is logical, not decorative | Headings are the outline a retrieval system reads before the prose | |
| 17 | Internal links point to money pages with meaningful anchors | The cheapest ranking lever most sites have never pulled | |
| 18 | Anchor text varies and describes the destination | “Click here” tells an engine nothing about the page it points at | |
| 19 | Breadcrumbs present and marked up | Helps both the user and the BreadcrumbList that appears in results | |
| 20 | URL structure is consistent — case, trailing slashes, separators | Inconsistency creates duplicates that nobody intended | |
| 21 | Mobile and desktop expose the same content and links | Content hidden on mobile is content the mobile-first index may not weigh | |
| D. Rendering, speed, mobile | 22 | Core Web Vitals read from field data, not a lab score | Lab scores flatter. Field data is what users actually experienced |
| 23 | Key content exists before JavaScript hydration | If the answer only appears after hydration, assume it does not exist | |
| 24 | Images sized, compressed and in a modern format | Usually the largest single weight on an Indian SME site | |
| 25 | Layout shift sources identified — images, ads, late fonts | Cheap to fix once you know which element is moving | |
| 26 | Tap targets and viewport behave on a real phone | Test on a mid-range Android on mobile data, not on your desktop | |
| 27 | Server response time measured, not assumed | On WooCommerce sites hosting is usually the constraint, not the theme | |
| 28 | Third-party scripts audited for what they cost | Chat widgets and analytics stacks routinely outweigh the page they sit on | |
| E. Structured data and meta | 29 | Every schema type in use validates without errors | Invalid markup is worth nothing at all |
| 30 | One Organization or LocalBusiness entity, referenced by @id | A fresh copy on each page makes the entity harder to resolve, not easier | |
| 31 | FAQPage markup matches the questions visible on the page | Marking up questions a user cannot see is a guidelines problem | |
| 32 | Prices in Product, Service or Offer match the visible price | A wrong price in schema is worse than no schema — AI answers will quote it | |
| 33 | Titles are unique, descriptive and front-load the term | Still the highest-leverage 60 characters on any page | |
| 34 | Meta descriptions written for the click, not for a keyword | They do not rank you; they decide whether the ranking is worth anything | |
| 35 | Canonical, Open Graph and hreflang agree with each other | Half-done internationalisation is worse than none | |
| F. AI access and citability | 36 | GPTBot allowed in robots.txt | Checked by name — a blanket allow can still be overridden further down the file |
| 37 | ClaudeBot allowed | Same check, different token | |
| 38 | PerplexityBot allowed | Perplexity cites densely and is often the first engine a new site wins | |
| 39 | Google-Extended allowed | Separate from Googlebot; blocking it removes you from some AI surfaces while leaving search intact | |
| 40 | CDN, WAF and bot rules tested from outside the network | A robots.txt that allows and an edge rule that blocks is the commonest false pass | |
| 41 | llms.txt published and pointing somewhere useful | An hour of work, no guarantees, and no reason not to | |
| 42 | Answer-first structure and a 10-run baseline citation test | You cannot show a citation improvement without a starting rate |
How to prioritise what you find
A crawler will hand you a list of a thousand issues and no sense of proportion. The order that works is: anything blocking access first — robots.txt, noindex, crawler rules, server errors — because nothing else you do matters while a door is shut. Then anything splitting signal: cannibalised pairs, canonical conflicts, parameter bloat. Then anything missing: schema, internal links, answer-first structure. Then anything slow, which is usually the section everyone starts with and rarely the one that moves rankings.
Score each finding on effort and impact and do the low-effort, high-impact ones this week. In practice that is usually three or four fixes, and they are usually in the first two categories. The thousand-item list can then be triaged honestly: most of it is one rule applied a thousand times.
If you run this yourself and find something alarming
Send it to us and we will tell you whether it matters, free, with no obligation and no call required. We would rather answer a question from somebody who ran the checklist themselves than sell an audit to someone who did not need one. If it turns out you do need the full thing, it is ₹12,500, ten working days, and credited against a retainer started within 30 days.
Checklist questions
Can I really run this myself?
Most of it, in a focused afternoon, if you have Search Console access and a crawler. The parts that are hard without practice are prioritisation, log analysis and the AI-crawler tests from outside your own network. The list itself is not secret and never should have been — what takes ten working days is doing it properly on a real site.
What tools do I need?
Search Console, a crawler such as Screaming Frog, a structured data validator, and a way to see field Core Web Vitals rather than lab scores. For the AI-access checks you need to test robots.txt and your CDN rules from outside your network, because an edge rule can block what robots.txt allows.
Which checks matter most?
Anything that blocks access — robots.txt, stray noindex tags, crawler rules, server errors. Then anything splitting signal, like cannibalised page pairs and parameter bloat. Speed usually matters least of the four for ranking, though it matters plenty for users. Fix in that order and most sites see something move within six weeks.
Why are the AI-crawler checks separate?
Because GPTBot, ClaudeBot, PerplexityBot and Google-Extended each have their own robots.txt token and each has to be checked by name. We find at least one blocked on roughly half the sites we audit, including sites paying another agency ₹40,000 a month, and it is almost never a decision anybody made deliberately.
How is your paid audit different from this list?
Evidence and order. Every finding arrives with a URL or a screenshot, a severity, an effort estimate and a position in a fix queue, plus a video walkthrough and a written answer on whether the site is worth a retainer at all. Ten working days, ₹12,500, credited if you start a retainer within 30 days.
How often should a site be audited?
Annually for a stable site, and immediately after any migration, replatform or CMS upgrade. Migrations are where crawler access, canonicals and schema break silently — we have seen all four AI user agents blocked by a default rule the morning after a replatform and nobody notice for six weeks.
Can I give this checklist to my developer?
Please do. That is what it is for, and it is also what a good freelance brief looks like — a numbered list with a reason attached to each item. If your developer disagrees with an item, they are probably right about your specific setup and worth listening to.
Related pages
- ₹12,500 technical audit The same 42, run properly
- AEO readiness checklist Score your AI visibility 0–100
- llms.txt guide Check 41, in detail
- Schema markup Checks 29–32, implemented
- Freelancer vs agency Brief a freelancer with this list
- Ecommerce SEO Where index bloat actually lives
Send the URL. Get a tier, a price and a start date.
One working day, on WhatsApp, from a person who has already opened your site in Search Console-shaped eyes. No discovery call required before you get a number.