JavaScript SEO is the discipline of making sure search engines, and increasingly AI crawlers, can find, render and index the content a React or Next.js application builds in the browser. Handle it well and the same framework that gives users a fast, interactive experience also ships HTML that search engines can crawl and quote. Handle it badly, and Googlebot, ChatGPT and Perplexity all see a thin, near-empty page while your visitors see a fully working product.
How Google Actually Crawls, Renders and Indexes JavaScript
Google processes JavaScript pages in three separate steps: crawling, rendering and indexing. Google's own JavaScript SEO documentation describes Googlebot fetching a URL, checking robots.txt, then queuing the page separately for rendering once it returns a 200 status code. Rendering runs on an evergreen, continuously updated build of Chromium, the same engine behind Chrome, so most modern JavaScript executes correctly.
The catch is timing. Google states that a page can stay in the render queue for a few seconds, but it can take considerably longer, and only after rendering finishes does Google see the final DOM and extract the text, links and structured data it indexes. A page that returns useful content directly in the raw HTML skips that wait. A page that depends entirely on client-side JavaScript to paint its content is betting that rendering happens promptly and correctly, for every page, every time you ship a change.
This is also why Googlebot's renderer, the Web Rendering Service, behaves differently from a real visitor's browser in ways that matter for anything more complex than a static page. It does not retain state across page loads, so Local Storage, Session Storage and cookies are cleared between renders. It also caches aggressively and can ignore your cache headers, which explains a fair share of "it works for me but not for Google" reports.
CSR, SSR, SSG, ISR or Streaming: Which Rendering Strategy Fits Your Site?
None of these strategies is inherently bad for SEO. Google can index content produced by any of them, provided the content lands in the HTML that Googlebot actually renders. The decision that matters is which content must exist in that first render, and which can safely wait.
| Strategy | What It Means | SEO Implication | Best For | |---|---|---|---| | Client-side rendering (CSR) | Browser downloads a near-empty shell; JavaScript builds the page after load | Content only exists after rendering completes; crawlers that skip JS see nothing | Authenticated dashboards and internal tools with no organic-search role | | Server-side rendering (SSR) | Server renders full HTML on every request | Crawlers receive complete content on the first fetch, no render-queue wait | Pages with frequently changing or personalized data that still need to rank | | Static site generation (SSG) | HTML built at deploy time, served from a CDN | Fastest possible response; content is fully present before any crawl happens | Marketing pages, documentation, blog posts, comparison pages | | Incremental static regeneration (ISR) | Static HTML that revalidates on a schedule or on demand | Combines SSG's speed with content that will not go stale indefinitely | Large catalogs, pricing pages, anything fed by a CMS | | Streaming and Server Components | Server sends static content immediately, streams slower parts behind Suspense | Core content should still land in the initial response; reserve streaming for non-essential, below-the-fold data | Pages mixing fast static content with a slower personalized widget |
Next.js makes React Server Components the default in the App Router, which pushes new projects toward the safer end of this table automatically. A component stays a server component, and renders to plain HTML on the server or at build time, until you explicitly add use client. Reserve use client for genuinely interactive pieces, like a filter control or a cart, and keep the content itself, product descriptions, article bodies, pricing tables, in server components so it exists without any rendering gamble.
Illustrative scenario: assume a SaaS company's pricing page is currently a client component that fetches plan data from an internal API after the page loads, showing a loading spinner for roughly 400 milliseconds on a fast connection. Converting the page to a server component that fetches the same data during rendering closes that gap: the HTML Google, and every AI crawler, receives on first render already contains plan names and prices, instead of a spinner and an empty container. The change is typically a few hours of engineering work: move the fetch into the server component, delete the client-side loading state, and keep only the plan selector's click handling in a small client component.
The JavaScript SEO Mistakes That Quietly Cost Rankings
Most JavaScript SEO problems are invisible in a normal browser, because your browser renders JavaScript perfectly. They only surface when you look at the page the way a crawler does.
Soft 404s From Client-Side Routing
Single-page applications often handle a missing product or an invalid route entirely on the client, showing a "not found" message while the server still returns an HTTP 200 status code. Google's guidance is explicit that this pattern "can lead to error pages being indexed and possibly shown in search results." The fix in Next.js is a single function call, covered in the code examples below.
Hash-Based and Fragment Routing
Routing that lives entirely after a # in the URL, such as /#/products/42, stopped being crawlable when Google deprecated the old AJAX-crawling scheme in 2015. Use real paths and the History API instead. This is rarely a problem in a fresh Next.js App Router project, since file-based routing already produces real URLs, but it resurfaces when a team ports an older React single-page app into Next.js and keeps its legacy router.
Infinite Scroll Without Paginated URLs
Infinite scroll feels great for users and is a common indexing trap. Google's guidance on lazy-loaded and infinite-scroll content is specific: give each content chunk its own persistent URL, such as ?page=12, keep the content at that URL consistent every time it loads, link sequentially between chunks so crawlers can discover them, and update the visible URL with the History API as the user scrolls. Skip any one of those steps and everything past the first chunk effectively does not exist for search.
Links Without Real Href Attributes
Googlebot discovers new URLs by parsing href attributes on anchor elements in the rendered HTML. A "link" built as a div or button with an onClick handler and no real href will not pass link discovery or internal link equity, even though it works fine for a visitor's mouse click. Next.js's built-in Link component renders a genuine anchor tag by default, so this mostly stays safe unless a custom design system or a third-party component swaps it out.
Content Locked Behind Interaction or Permissions
Content that only appears after a click, a hover, a scroll past a threshold, or a browser permission prompt for the camera or microphone, is content Googlebot will not see. Google's own advice is blunt here: features that require user permission "don't make sense for Googlebot, or for all users," so give people, and crawlers, a way to reach the content without that gate.
Blocked JavaScript or CSS Resources
If robots.txt blocks the JavaScript bundles or stylesheets a page needs, Googlebot cannot render the page the way a visitor sees it, and may index a sparser version. Audit robots.txt every time you change your build pipeline or CDN configuration, not just when you launch.
| Symptom | Likely Cause | Fix |
|---|---|---|
| Deleted or invalid pages stay indexed | Client-side routing returns 200 instead of 404 | Call notFound() on the server, or return a real 404 status |
| Deep content never gets indexed | Hash routing or infinite scroll without unique URLs | Switch to real paths; paginate with unique, linked URLs |
| Internal pages have few or no inbound internal links | Navigation built with non-anchor click handlers | Use real anchor tags or next/link for anything that should be crawlable |
| Rendered page looks broken in URL Inspection | JS or CSS blocked in robots.txt | Allow crawling of the assets the page needs to render |
| Content missing from the indexed version only | Content gated behind scroll, hover or permission prompts | Render essential content by default; treat the gate as progressive enhancement |
Next.js App Router Patterns That Keep Every Page Indexable
The App Router's data and metadata APIs map closely onto the issues above, and Next.js's own API reference is worth bookmarking alongside Google's guidance. Four patterns close most of the gap between a page that renders fine in a browser and one that indexes cleanly.
Canonical URLs and Metadata With generateMetadata
Route parameters in the App Router are asynchronous, so params arrives as a Promise that you await before use. generateMetadata runs on the server, which means the resulting title, description and canonical tag are part of the HTML response rather than something a crawler has to wait for.
// app/blog/[slug]/page.tsx
import type { Metadata } from "next";
type Props = {
params: Promise<{ slug: string }>;
};
export async function generateMetadata({
params,
}: Props): Promise<Metadata> {
const { slug } = await params;
const post = await getPostBySlug(slug);
if (!post) {
return { title: "Post Not Found" };
}
return {
title: post.title,
description: post.description,
alternates: {
canonical: `/blog/${slug}`,
},
};
}
sitemap.ts and robots.ts
Both files are plain TypeScript, generated at build or request time from your real data instead of maintained by hand. For large sites, generateSitemaps splits the output across multiple files once you approach the 50,000-URL limit per sitemap.
// app/sitemap.ts
import type { MetadataRoute } from "next";
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const posts = await getAllPosts();
return posts.map((post) => ({
url: `https://example.com/blog/${post.slug}`,
lastModified: post.updatedAt,
changeFrequency: "monthly",
priority: 0.7,
}));
}
// app/robots.ts
import type { MetadataRoute } from "next";
export default function robots(): MetadataRoute.Robots {
return {
rules: { userAgent: "*", allow: "/" },
sitemap: "https://example.com/sitemap.xml",
};
}
Real 404s With notFound()
This is the direct fix for the soft-404 pattern Google warns about. Calling notFound() inside a server component renders the nearest not-found.tsx boundary and returns an actual 404 status, instead of a 200 with a "not found" message baked into the HTML.
// app/blog/[slug]/page.tsx
import { notFound } from "next/navigation";
export default async function BlogPostPage({ params }: Props) {
const { slug } = await params;
const post = await getPostBySlug(slug);
if (!post) {
notFound();
}
return <Article post={post} />;
}
Structured Data in Server Components
Because server components render on the server, a JSON-LD script tag built from real page data is present in the initial HTML, with no dependency on client-side execution.
// components/article-json-ld.tsx
export function ArticleJsonLd({
title,
description,
url,
datePublished,
}: {
title: string;
description: string;
url: string;
datePublished: string;
}) {
const json = {
"@context": "https://schema.org",
"@type": "Article",
headline: title,
description,
url,
datePublished,
};
return (
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(json) }}
/>
);
}
Keep this in a server component and pass it real data fetched during rendering. A version guarded behind use client defeats the purpose: the markup would exist for a user's browser but might not be present in time for the render a crawler sees.
Do AI Crawlers Read JavaScript the Way Google Does?
No, and the gap is bigger than most teams assume. Google can execute JavaScript, even if rendering is delayed. Most AI crawlers cannot execute it at all.
| Crawler | Operator | Executes JavaScript | What This Means |
|---|---|---|---|
| Googlebot | Google Search | Yes, via evergreen Chromium, though rendering can be delayed | Client and server-rendered content both get indexed eventually; server rendering is still faster and safer |
| Google-Extended | Google, for training its generative AI models | Not documented for this purpose; a separate crawler from Search | Does not affect Search ranking or inclusion; controlled independently in robots.txt |
| GPTBot, ChatGPT-User, OAI-SearchBot | OpenAI | No | Only the raw HTML your server sends is visible to these crawlers |
| ClaudeBot | Anthropic | No | Fetches JavaScript files in some requests but never executes them |
| PerplexityBot | Perplexity | No | Same constraint applies to citations shown inside Perplexity answers |
An analysis of crawler traffic published by Vercel, built with the technical SEO firm MERJ, found that none of the major AI crawlers it studied render JavaScript: GPTBot fetched JavaScript files in roughly 11.5% of requests and ClaudeBot in roughly 23.8%, but neither one executed what it fetched. Both crawlers parsed whatever text was already present in the raw HTML and moved on.
The practical takeaway is blunt. If your value proposition, pricing or product detail only exists after client-side JavaScript runs, GPTBot, ClaudeBot and PerplexityBot never see it, regardless of how well Googlebot eventually renders it. This matters for generative engine optimization specifically, since these crawlers feed the models that generate answers in ChatGPT and Perplexity. It is also worth knowing that Google's own documentation on AI features states plainly that no special markup or files are required to appear in AI Overviews or AI Mode: a page only needs to be indexed and eligible to show a normal snippet, which routes straight back to the rendering fundamentals covered above.
How to Test Whether Google Can Actually See Your Content
Do not guess. Testing rendering takes minutes and settles the argument.
- Fetch the raw HTML. Run a plain
curlrequest against the live URL and read the output before any JavaScript executes, to see exactly what every non-rendering crawler receives. - Compare it to the rendered DOM. Open Search Console's URL Inspection tool, run a live test, and view the rendered HTML and screenshot Google actually produced for that URL.
- Check the indexed version. Still in URL Inspection, confirm the page is indexed and review which canonical URL Google chose, since it can differ from the one you specified.
- Search for your critical text. Confirm the headline, price, or primary content block appears verbatim in the rendered HTML, not only in the visual screenshot.
- Repeat after every major release. A routing change, a new component library, or a shift to client-side data fetching can silently reintroduce a rendering gap that passed testing last quarter.
curl -s https://example.com/blog/your-post-slug | head -c 2000
If step 1 returns a near-empty div with an id and nothing else, your primary content depends entirely on client-side rendering, and every crawler that skips JavaScript execution is seeing that empty shell.
How Agentixly Approaches JavaScript SEO Audits
When Agentixly runs a JavaScript SEO engagement, the work stays inside your codebase instead of living in a slide deck nobody implements. A typical audit and fix moves through four phases:
- Render diff audit (days 1 to 5). We compare the raw HTML your server sends against the fully rendered DOM for your most important templates, category pages, product pages, articles, and flag every piece of content, link or metadata tag that only appears after client-side JavaScript runs.
- Rendering and metadata fixes (weeks 1 to 3). Engineers convert client components to server components where it is safe, wire up
generateMetadata, canonical tags and JSON-LD, and replace client-only routing with real URLs, all delivered as pull requests in your repository. - Crawler verification (from week 3). We test with Search Console's URL Inspection tool, compare raw HTML against what AI crawlers like GPTBot and ClaudeBot actually receive, and confirm indexing in production rather than in a staging environment.
- Regression monitoring (ongoing). We set up alerts for indexing drops and rendering regressions, so a routing refactor or a new component library six months from now does not quietly undo the fix.
Because Agentixly's web development team and SEO team work against the same codebase, fixes ship as code, not recommendations that wait in a backlog. That matters most on React and Next.js projects, where the gap between what a page looks like and what a crawler receives is invisible until rankings, or AI citations, quietly stall.
Next Steps: Ship Rendering You Can Verify
You do not need to guess whether Google, ChatGPT or Perplexity can see your content. Fetch the raw HTML, compare it against the rendered version, fix the specific gap with the App Router primitives covered here, and test again after every major release. Most JavaScript SEO problems trace back to one root cause: content that exists for a user's browser but not in the HTML a crawler actually receives.
For related technical ground, see our technical SEO checklist for developers, our Next.js performance optimization guide for the speed side of the same App Router APIs, and our structured data and schema markup guide for what to put inside the JSON-LD pattern shown above. If you are publishing at scale, our programmatic SEO guide covers the template and quality-gating side of the same rendering questions, and our website migration SEO checklist and llms.txt guide cover two more pieces of the same visibility puzzle.
If your React or Next.js site looks strong in the browser but underperforms in search and AI answers, the rendering layer is the first place to look. Agentixly's SEO team audits the gap between what your server sends and what crawlers actually index, then ships the fixes as pull requests instead of a slide deck. Contact us and we will answer within 24 hours.