Web

SEO for people who would rather be building

The short checklist I run before any page goes live, and why most of it is just being a good host.

7 min read

A Haunted House: how a web page gets found, told in one song. A house nobody visits is discovered through links, visited by a crawler, rendered, indexed, ranked, and finally shown to a visitor.

I like building things. I don't like wondering whether anyone will ever find them.

So I sat down and learned SEO properly. I expected a dark art. It turned out to be a short list, and most of it is just being a good host: tell people where the door is, let them in, and don't leave them waiting in the dark.

I also didn't want to write another list of tips. Google's own docs are right there, and they're good. So I drew a haunted house instead. That's the film above. This is the slow version.

Search is a pipeline

This is the one idea everything else hangs off. A page goes through six stages before anyone sees it in a search result, and every SEO technique you've ever heard of is a fix for one of them.

  1. 01 · Discover

    Google learns the URL exists, from links and sitemaps.

    Breaks onPages nothing links to. Navigation that only works through click handlers.

  2. 02 · Crawl

    Googlebot asks your server for the page.

    Breaks onA robots.txt that says no. Server errors. Redirect chains.

  3. 03 · Render

    It builds the page: HTML first, JavaScript later.

    Breaks onContent that only shows up after a client-side fetch or a click.

  4. 04 · Index

    It decides whether to keep the page, and which URL is the real one.

    Breaks onA stray noindex. Duplicates. A "not found" page that returns 200.

  5. 05 · Rank

    The page is matched against what someone typed.

    Breaks onContent that doesn't help. Nothing linking in. A slow, jumpy page.

  6. 06 · Display

    How the result looks, which decides whether it gets clicked.

    Breaks onA vague title. No description. A broken share card.

The useful part is the order. When something is wrong, find the earliest stage that's broken and fix that first. Polishing a title is pointless on a page that was never indexed.

If you write the code, you own almost all of stages one to four, and six. Stage five is mostly about whether the page is actually worth reading, which is a different kind of problem.

There are two versions of every page

The first is the response HTML: the exact string your server sends back. The second is the rendered DOM: what the page turns into after JavaScript has run.

You see the first with View Source or curl. You see the second in the Elements panel. They can be very different.

Anything in the response HTML is visible to every bot on the internet. Anything that only exists after JavaScript runs is visible only to bots that run JavaScript, and only if they wait long enough.

Google does run JavaScript, but as a second step. It reads the raw HTML first, takes the links, and renders the page later in a headless Chrome. That renderer doesn't click, doesn't scroll, and keeps no cookies. The bots behind link previews on WhatsApp, Slack and LinkedIn don't run JavaScript at all. On the best evidence I could find, neither do most AI crawlers.

Try this on something you've built. It takes a minute.

terminal
# What does the server really send?
curl -s https://blog.vedasdixit.com/seo | head -c 3000

# Status and headers only. Try a URL that shouldn't exist.
curl -sI https://blog.vedasdixit.com/nope

# The two files every crawler asks for first.
curl -s https://blog.vedasdixit.com/robots.txt
curl -s https://blog.vedasdixit.com/sitemap.xml

Then open DevTools, run Disable JavaScript from the command menu, and reload. What's left on screen is roughly what those bots get. On a plain React app, that's an empty <div id="root">. Which is the whole argument for rendering on the server.

The setup

The list is short, and I go through it in pipeline order. The code below is Next.js (the App Router), but the ideas are the same anywhere.

Google finds pages by following links from pages it already knows. So navigation has to be real <a href> links. A div with a click handler that calls the router is invisible to a crawler.

A sitemap is the map you hand over on top of that. I generate it from the list of posts, so there is nothing to keep in sync by hand.

app/sitemap.ts
export default function sitemap(): MetadataRoute.Sitemap {
  return [
    { url: absoluteUrl("/") },
    { url: absoluteUrl("/author") },
    ...getPosts().map((post) => ({
      url: absoluteUrl(`/${post.slug}`),
      lastModified: new Date(post.updated ?? post.date),
    })),
  ];
}

Two details. Google ignores priority and changeFrequency, so I left them out. And lastModified has to be true. Stamp today's date on everything and Google learns to ignore it.

A sitemap helps a page get found. It doesn't force anything, and it's no substitute for links. A page you can only reach through the sitemap is an orphan, and it tends to perform like one.

The sign in the yard

robots.txt tells crawlers what they're allowed to fetch. A good default says come in, and points at the map.

app/robots.ts
export default function robots(): MetadataRoute.Robots {
  return {
    rules: { userAgent: "*", allow: "/" },
    sitemap: absoluteUrl("/sitemap.xml"),
  };
}

This file is easy to get backwards, because it sits right next to another tool that sounds like it does the same job.

robots.txt Disallownoindex
ControlsCrawling: don't fetch this.Indexing: don't show this in results.
Removes a page from Google?No. A blocked URL can still be indexed from links, with no description.Yes, once it's crawled.
The catchGoogle can't see a noindex on a page it isn't allowed to fetch.The page has to stay crawlable for the tag to be read.

So to take a page out of search, you allow crawling and add noindex. Disallow is for saving the crawler's effort on junk. It's also public, so it isn't security.

The scary version of this mistake is launching with staging's Disallow: / still in place. Check with curl on launch day. Not from memory.

Words in the HTML

For a blog, the simplest thing is to keep every page static. Posts are turned into HTML at build time, so the very first response already has everything in it.

One misconception worth dropping: 'use client' doesn't mean "not in the HTML". Client components are rendered to HTML on the server too, then hydrated. The directive only means the component's JavaScript ships to the browser. The film at the top of this page is a client component, and its first frame is still in the response.

What actually goes missing is narrower:

  • Data fetched in useEffect, with no server render behind it.
  • next/dynamic with ssr: false.
  • Anything gated on window, mounted or localStorage.
  • Tab or accordion content that only renders after a click.
  • Everything past the first batch of an infinite scroll.

One address per page

The same page is often reachable at several URLs: with and without www, with a trailing slash, with ?utm_source= stuck on the end. Google picks one as the real one. A canonical tag is your vote.

Every page casts its own. It's one line in the page's metadata.

app/[slug]/page.tsx
export async function generateMetadata({ params }) {
  const { slug } = await params;
  const post = getPost(slug);
  return {
    title: post.title,
    description: post.dek,
    alternates: { canonical: `/${post.slug}` },
  };
}

The one place it must not go is the root layout. Put a canonical there and every page inherits it, which tells Google the whole site is the home page.

The other half of this stage is honest status codes. A made-up URL should return a real 404, not a friendly "not found" message with a 200. In Next.js, dynamicParams = false on the post route does this: any slug that isn't a real post is a 404 before anything renders.

How the result looks

The title is the strongest thing on the page that I control. I keep each one unique, with the topic first and the site name last.

The description doesn't affect ranking at all. It's ad copy for the click, and Google often swaps in its own snippet anyway. I still write one for every page.

Each post also carries a small block of JSON-LD. It states plain facts for machines: this is a blog post, here's the headline, the date, who wrote it.

app/[slug]/page.tsx
const jsonLd = {
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  headline: post.title,
  description: post.dek,
  datePublished: post.date,
  dateModified: post.updated ?? post.date,
  author: { "@type": "Person", name: "Vedas Dixit", url: "/author" },
};

It isn't a ranking boost. It makes the page eligible for a nicer-looking result, and that's all. The one rule is that it has to describe what's visibly on the page.

The weird part

I set the share image once, in the root layout. Later I gave posts their own openGraph block, for the title and the publish date.

That combination has a trap in it.

Next merges metadata shallowly. A page that sets openGraph replaces the layout's openGraph whole. It doesn't merge field by field. So the post keeps its new title and quietly loses the site's image.

Nothing warns you. The build is green, the page looks fine, and the link preview comes up blank.

What I ended up doing is keeping the shared fields in one place and spreading them into every page that sets its own.

lib/site.ts
export const OG_BASE = {
  siteName: SITE_NAME,
  locale: LOCALE,
  images: [OG_IMAGE],
};

// and in any page that sets its own openGraph
openGraph: {
  ...OG_BASE,
  type: "article",
  title: post.title,
  publishedTime: post.date,
},

Every page that touches openGraph has to remember to do this. It's one line, so I can live with it.

Things I'm ignoring

A lot of SEO advice is folklore. These are the things I decided not to spend time on.

  • The keywords meta tag. Google has ignored it for well over a decade.
  • Sitemap priority and changeFrequency. Also ignored.
  • FAQ schema. FAQ rich results stopped showing in May 2026. Rich result types get retired, so check what's still live before building for one.
  • llms.txt and other AI-only files. Google Search doesn't read them. Its AI answers are built on the normal index, so a page that's indexed and readable is already in the running.
  • A perfect Lighthouse score. Core Web Vitals are a real ranking signal and a small one. Google measures real users, not my laptop. And a fast page with a weak answer still loses to a slow page with the best one.

My test for any claim now: which stage of the pipeline does it affect, and does Google's documentation actually say so? If I can't answer either, I skip it.

The checklist

This is what I run before a page goes live. It's short on purpose.

That's the whole list. None of it is clever. It's mostly just being a good host.

If you want to go deeper