Technical SEO Checklist: 15 Checks for Crawling, Indexing and Speed
A practical technical SEO checklist: crawling, indexing, redirects, canonicals, sitemaps, Core Web Vitals, rendering and structured data, with how to check and fix each.
- Read time
- 17 min read
- Sections
- 24
- FAQs answered
- 15
- Topic
- Technical SEO
A technical SEO checklist covers the things that decide whether search engines can find, read and index your pages: crawling, indexing, status codes, duplicate pages, site structure, speed, mobile layout, rendering and structured data. Work through it in order, because a problem early in the list can make later improvements pointless.
This checklist is written to be used. Each item says what to check, how to check it, why it matters and what a fix looks like. It assumes you can open Google Search Console and, ideally, run a site crawler such as Screaming Frog or a similar tool. If you need a developer for some fixes, the "what to ask for" notes should help you brief them. For the bigger picture of how search works, see our SEO guide. If you would rather have a specialist do the audit, our technical SEO services and SEO audit cover it.
How to use this checklist
- Verify your site in Search Console (a domain property is best). Without it you are working blind.
- Crawl your site with a crawler, so you see it the way a bot does.
- Work top to bottom. Crawling and indexing issues come first, because they stop everything else.
- Rate each finding as critical, important or minor, and fix in that order.
- Record changes with dates, so you can link results to actions.
1. Crawling: can bots reach your pages?
Check robots.txt
Open yoursite.com/robots.txt. Look for Disallow rules that block important sections, a blanket Disallow: / left over from a staging site, and blocked CSS or JavaScript files that pages need to display properly. Google's documentation explains that robots.txt controls crawling, not indexing. We explain the difference and a common trap in why a blocked URL can still appear in Google.
Check that important pages are linked
Google discovers pages by following links, so a page with no internal links pointing to it may not be found. Your crawler's report on orphan pages, meaning pages in your sitemap that nothing links to, shows these. Link them from relevant pages. Our internal linking guide explains how.
Use crawlable links
Google's documentation says links should be <a> elements with an href attribute that resolves to a real address. Links built only with JavaScript click handlers, or with href="javascript:...", may not be followed. Anchor text should be descriptive. Avoid "click here".
Check server logs and crawl stats
In Search Console, the Crawl Stats report under Settings shows how often Googlebot requests your site, response codes and average response time. Frequent 5xx errors or slow responses can reduce how much Google crawls. Large sites should also review server logs to see which sections Googlebot spends time on.
2. Indexing: are the right pages in Google?
Review the Page indexing report
In Search Console, open Indexing and then Pages. The report lists indexed pages and the reasons others are not indexed. Common reasons and what they usually mean:
| Status | Usually means | Typical action |
|---|---|---|
| Crawled, currently not indexed | Google fetched the page but chose not to index it, often because of thin or duplicate content | Improve the page, merge it with a stronger one or remove it |
| Discovered, currently not indexed | Google knows the URL but has not crawled it yet | Strengthen internal links, check server speed |
| Duplicate without user-selected canonical | Several similar pages, no clear preferred one | Add a canonical tag pointing to the preferred page |
| Excluded by noindex tag | The page tells Google not to index it | Confirm that was intended. Remove the tag if not |
| Blocked by robots.txt | Google is not allowed to crawl it | Allow crawling if the page should be found |
| Page with redirect | The URL redirects elsewhere | Normal for old URLs. Update internal links to the final address |
| Not found (404) | The URL returns a missing page | Redirect if a replacement exists, otherwise leave as 404 or 410 |
Not every URL should be indexed. Filter pages, internal search results and thank-you pages often should not be. The goal is that every page you want to rank is indexed, and that few low-value pages are.
Inspect key URLs
Use the URL Inspection tool for your most important pages. It shows whether the page is indexed, which canonical Google chose, the rendered HTML and any problems. If you recently fixed something, use it to request indexing.
Check noindex and snippet controls
Search your crawl for noindex in meta tags or X-Robots-Tag headers, and for nosnippet or very low max-snippet values. These are sometimes added for staging and never removed. Google notes that restrictive preview controls also limit how your content can feature in its AI experiences.
3. Status codes and redirects
| Code | Meaning | Use it for |
|---|---|---|
| 200 | OK | Live pages |
| 301 | Moved permanently | Permanent URL changes and merged pages |
| 302 or 307 | Moved temporarily | Genuinely temporary moves |
| 404 | Not found | Pages that do not exist |
| 410 | Gone | Pages permanently removed with no replacement |
| 5xx | Server error | Never intended. Fix |
Checks to run:
- Redirect chains. A to B to C wastes crawling and slows visitors. Point each old URL directly at the final one.
- Redirect loops. Two URLs redirecting to each other break the page. Fix immediately.
- Soft 404s. Pages that say "not found" but return 200 confuse search engines. Return a true 404 or 410.
- Internal links to redirected URLs. Update them to the final address.
- Mixed versions. Make sure http redirects to https, and that www and non-www resolve to one preferred version.
When you merge or retire pages, plan redirects first. Use a 301 to the closest relevant page, not to the home page.
4. Duplicate content and canonical tags
When the same or very similar content is available at more than one address, search engines have to choose one, and may choose the wrong one. Common causes include www and non-www versions, trailing slashes, tracking parameters, filtered category pages, printer versions and pages that were copied for different regions.
- Add a self-referencing
rel="canonical"tag on each indexable page, pointing to its own preferred address. - Point duplicates to the preferred version with a canonical or, better, a 301 redirect where the duplicate does not need to exist.
- Do not canonical a page to something very different. Google treats the tag as a hint and may ignore it.
- Do not use robots.txt to handle duplicates, because blocked pages cannot be crawled to read the canonical.
- Consolidate competing pages. Two pages targeting the same topic usually do better as one stronger page. This is a content decision as much as a technical one.
Google's canonicalization troubleshooting guidance notes that it can take time for a changed canonical to be re-evaluated, so be patient after you fix things.
5. Site structure and URLs
- Keep important pages close to the home page. Aim for key pages within about three clicks. This is a rule of thumb, not a Google rule.
- Use a logical hierarchy. Categories, subcategories and pages that reflect how customers think.
- Use short, readable URLs. Use hyphens between words, lowercase letters and no unnecessary parameters.
- Use breadcrumbs where they help navigation. They also help Google understand the hierarchy.
- Do not change URLs without a reason. Every change risks lost rankings. If you must, redirect precisely.
- Control faceted navigation. On shops, filters can create thousands of near-duplicate URLs. Decide which combinations deserve indexing, and keep the rest from being crawled or indexed.
- Handle pagination sensibly. Each page in a series should be crawlable through links, with its own address.
6. XML sitemaps
A sitemap lists the URLs you want search engines to know about. It helps discovery, especially for new or large sites, but it does not guarantee indexing. Best practice:
- Include only canonical, indexable URLs that return 200.
- Exclude redirected, noindex, blocked and error URLs.
- Keep each sitemap file within the size limits, and use a sitemap index for large sites.
- Reference it in robots.txt and submit it in Search Console.
- Update the lastmod value only when a page really changed. Dishonest dates make the field untrustworthy.
- Keep it automatically generated, so it stays in step with the site.
7. Speed and Core Web Vitals
Google's Web Vitals guidance defines Core Web Vitals as three measures, each with a recommended threshold measured at the 75th percentile of page loads, separately for mobile and desktop.
| Metric | Measures | Good threshold |
|---|---|---|
| Largest Contentful Paint (LCP) | Loading | 2.5 seconds or less |
| Interaction to Next Paint (INP) | Responsiveness | 200 milliseconds or less |
| Cumulative Layout Shift (CLS) | Visual stability | 0.1 or less |
Source: web.dev, Web Vitals. Check the Core Web Vitals report in Search Console and PageSpeed Insights, which shows both real-user data where available and lab tests. Typical fixes:
- Slow LCP: compress and properly size hero images, serve modern formats, reduce server response time, remove render-blocking scripts, preload the main image or font.
- Poor INP: reduce heavy JavaScript, break up long tasks, delay non-essential third-party scripts.
- High CLS: set width and height for images and embeds, reserve space for ads and banners, avoid inserting content above existing content.
Speed matters for visitors first. Do not sacrifice content or usability for a score. Our page speed services handle the deeper work.
8. Mobile friendliness
- Use a responsive layout, so the same content appears on every device.
- Ensure the mobile version has the same important content and links as desktop, since Google indexes using the mobile version.
- Check that text is readable without zooming, tap targets are spaced and nothing is hidden behind intrusive pop-ups.
- Test on real phones, not only a browser emulator.
9. Rendering and JavaScript
Google can process JavaScript, but sites that rely heavily on it are harder to get right. Google's guidance says content in JavaScript is processed as long as it is not blocked. Checks:
- Use URL Inspection and view the rendered HTML to confirm that your main content and links appear.
- Make sure essential content does not depend on a user action, such as clicking or scrolling, to appear.
- Prefer server-side rendering or pre-rendering for important pages, which also helps other crawlers that may not run scripts.
- Use real links, not click handlers, for navigation.
- Make sure error pages return proper status codes, not a 200 with an error message.
10. HTTPS and security
- Serve every page over HTTPS with a valid certificate that renews on time.
- Redirect all http requests to https and fix mixed content, where secure pages load insecure resources.
- Check for hacked content or injected spam links with the Security issues report in Search Console.
- Keep software, plugins and themes updated.
11. Structured data
Structured data describes what a page is about in a machine-readable way, usually as JSON-LD. Use it where it matches the visible content and a Google feature supports it, such as Organization, Product, Article, LocalBusiness and BreadcrumbList. Validate it with Google's Rich Results Test and monitor enhancements in Search Console. Note that FAQ rich results no longer appear in Google Search, as we explain in our note on FAQ markup. Google also says structured data is not required for its generative AI features. Do not add markup for content that is not on the page. Our schema services can help.
12. International and multilingual sites
- Use separate URLs for each language or region, not content switched by cookies.
- Use hreflang annotations so Google serves the right version, and make them reciprocal.
- Translate properly and adapt currency, units and legal details. Do not just swap the currency symbol.
- Avoid redirecting visitors automatically by location in a way that stops crawlers from seeing all versions.
See our international SEO services if you operate in more than one market, as we do from Manchester and Mumbai.
13. Images and media
- Use descriptive file names and meaningful alt text for important images.
- Compress images and use modern formats.
- Lazy-load images below the fold, but not the main image at the top of the page.
- For video, follow Google's video SEO guidance: make the watch page indexable, use standard embedding elements and provide metadata.
14. Crawl budget, for large sites
Crawl budget matters mainly for very large or frequently changing sites. If your site has millions of URLs, or thousands of low-value URLs generated by filters, check that crawlers spend their time on pages that matter: block or noindex low-value parameter combinations, fix slow responses and errors, keep the sitemap clean and link your best pages prominently. Small sites rarely need to worry about it.
15. Ongoing monitoring
| Task | How often |
|---|---|
| Check Search Console for new errors and manual actions | Weekly |
| Review page indexing and Core Web Vitals reports | Monthly |
| Run a full crawl and compare with last month | Monthly or after big changes |
| Test key templates after every release | Every release |
| Review redirects, sitemaps and robots.txt | Quarterly |
| Review structured data and server logs | Quarterly |
A severity guide for your findings
| Severity | Examples | Fix within |
|---|---|---|
| Critical | Whole site noindexed, robots.txt blocking everything, key pages returning 5xx, broken redirects on top pages | Immediately |
| Important | Duplicate pages competing, redirect chains, slow LCP on key templates, missing canonicals, orphan pages | Within weeks |
| Minor | Missing alt text on decorative images, small title issues, harmless warnings | As time allows |
Common technical SEO mistakes
- Launching a redesign without a redirect map, so old URLs return 404.
- Leaving a staging site's noindex or robots block on after launch.
- Using robots.txt to hide pages from search results, rather than a noindex tag.
- Letting plugins create thousands of thin archive or tag pages.
- Fixing scores in a speed tool while visitors still find the site slow.
- Adding structured data that does not match the visible page.
- Ignoring Search Console warnings for months.
Briefing a developer
When you hand work to a developer, be specific. A good brief states the problem, the evidence, the desired outcome and how you will test it. For example: "These 214 product URLs return 200 but display a not-found message. They should return 410. Please change the response code. We will check with URL Inspection and a crawl." Attach the crawl export and Search Console screenshots. Ask for the change on a staging site first, and test before release.
A worked example: a site migration, and what to check
This is an invented example. A Glasgow retailer moves from an old platform to a new one, changing many URLs. Within a week, organic traffic falls by 40 percent. Here is how a methodical check recovers it.
Triage
- Is the site indexable? robots.txt allows crawling and there is no site-wide noindex. Good.
- Do old URLs redirect? A crawl of the old URL list shows that 600 of 900 old pages return 404. The redirect map was only partly applied.
- Do redirects go to the right places? Many category pages redirect to the home page, which dilutes relevance. Some product redirects loop.
- Are internal links updated? The new templates still link to old URLs, so every click goes through a redirect or a 404.
- Is the sitemap right? It still lists old URLs.
- Are canonicals right? Product pages point to a staging domain.
Fixes
- Apply permanent redirects from each old URL to the closest relevant new page, and test them in bulk.
- Update internal links and the sitemap to the new URLs.
- Correct canonicals to the live domain.
- Resubmit the sitemap, request indexing for key pages and monitor the Page indexing report.
Lessons
Before any migration, create a full redirect map, test on a staging site, crawl both versions, check robots.txt and canonicals on launch day and watch Search Console daily for two weeks. Recovery usually takes weeks, and Google says canonicalisation changes can take time to be re-evaluated.
A quarterly review template
| Area | Questions | Output |
|---|---|---|
| Crawlability | Any new blocks? Are key sections linked? Any server errors in Crawl Stats? | List of issues and owners |
| Indexing | Are key pages indexed? Any growth in excluded or error categories? | Indexed versus expected pages |
| Performance | Do key templates pass Core Web Vitals on mobile? | Pass rates by template |
| Duplication | New duplicate or competing pages? Canonicals correct? | Merge or fix list |
| Redirects and errors | New chains or loops? 404s from internal links? | Fixes shipped |
| Structured data | Valid and matching visible content? | Errors and warnings fixed |
| Security | Certificate dates, mixed content, issues in Search Console? | Actions |
| AI crawler access | Policy for Googlebot, OAI-SearchBot and PerplexityBot checked in robots.txt and firewall? | Confirmed policy |
A story of a small technical fix with a big effect
This is an invented composite. A Hull-based retailer of garden tools noticed that one of its best-selling categories, hand tools, had lost visibility over several months, while competitors' pages appeared more often. The team had published new product pages and a buying guide, and nothing seemed to help. Eventually, the marketing manager decided to look at the basics before creating more content.
She started with Search Console's Pages report. Under "Not indexed", she found a large group labelled "Crawled, currently not indexed", containing about 300 URLs from the hand tools category. She opened a few. They were filtered versions of the category page, for example "hand tools, brand X, price low to high, page 4". The site's filters created a new URL for every combination, and each had a canonical tag pointing to itself, not to the main category page. Google had crawled thousands of near-identical URLs and decided not to index many of them, and it appeared to be spending time on those instead of on the main product pages.
She then checked the internal links. The main category page linked to the filtered versions in a way that made them look important, and the sitemap listed them. She also noticed that the page's server response time was slow when filters were applied, because each combination required a database query.
She briefed the developer clearly: "Filtered URLs should not be in the sitemap. Filtered pages with no unique content should carry a canonical to the main category page. Links to filters should use a mechanism that does not create crawlable URLs for sort order and price filters. Please keep indexable filters, such as brand pages, which have real search demand. We will check with a crawl and Search Console." The developer made the changes over two weeks, including caching to speed up responses.
Over the next two months, the number of crawled-but-not-indexed URLs in that group fell sharply, and impressions and clicks for the main hand tools category and its brand pages increased. The manager noted that other seasonal factors might have helped, but she also saw that the product pages were being crawled more frequently. The fix did not involve any new content, any links or any clever tactics. It was a matter of looking at the data Google provides, understanding what it was telling her and fixing the underlying structure. The story is a reminder that when SEO performance drops, a technical audit should come before more content.
Where we can help
We are a digital marketing agency in Manchester, UK and Mumbai, India. Our technical SEO services begin with a crawl, a Search Console review and a prioritised list, and we agree what success looks like before we start. Contact us if you would like us to look at your site.
Your questions, answered in plain English
Technical SEO is the work that lets search engines crawl, index and understand your pages. It covers robots.txt, sitemaps, status codes, canonicals, site structure, speed, mobile layout, rendering and structured data.
Crawling, indexing, status codes and redirects, duplicate content and canonicals, site structure, sitemaps, Core Web Vitals, mobile friendliness, JavaScript rendering, HTTPS, structured data, international setup and ongoing monitoring.
According to Google's web.dev guidance, a good experience means LCP of 2.5 seconds or less, INP of 200 milliseconds or less and CLS of 0.1 or less, measured at the 75th percentile of page loads.
Crawling is when a bot fetches a page. Indexing is when the search engine analyses it and stores it in its index. A page can be crawled but not indexed, and a page that is not indexed cannot appear in results.
Not reliably. It stops crawling, but a blocked URL can still be indexed if it is linked elsewhere. To keep a page out of results, allow crawling and use a noindex directive.
Google fetched the page but decided not to index it, often because it is thin, duplicated or low in value. Improve it, merge it with a stronger page or remove it.
It is not required, but it helps discovery, especially on new or large sites. Include only canonical, indexable URLs that return 200, and do not expect it to guarantee indexing.
A canonical tag tells search engines which address is the preferred version when several URLs show the same or very similar content. Google treats it as a hint, not a command.
Use a 301 when a page has moved permanently or two pages have been merged. Redirect each old URL to the closest relevant new page, and avoid chains and loops.
Find the final destination and update each old URL to redirect straight to it, then update internal links to point to the final address. Re-crawl to confirm.
Google publishes Core Web Vitals as measures of user experience, and a slow, unstable page frustrates visitors. Improve speed for visitors first and treat scores as a guide, not a goal in themselves.
Yes, Google says it can process content in JavaScript as long as it is not blocked, but JavaScript sites are generally more complex to get right. Check rendered HTML in URL Inspection and prefer server-side rendering for important pages.
Check Search Console weekly, review indexing and Core Web Vitals monthly, crawl monthly or after big changes and review robots.txt, redirects and sitemaps quarterly.
Not for basic ranking. Use it where it matches visible content and supports a Google feature. Google also says structured data is not required for its generative AI features.
You can find and prioritise most issues yourself with Search Console and a crawler. Some fixes, such as redirects, rendering and speed work, usually need a developer, so brief them clearly with evidence.
Still curious? Send us your question and a strategist will get back to you.
Found this useful?
Want this done for you?
Start with an audit tied to revenue. We will tell you what is worth fixing and what is not.