Notable changes, newest first — new crawl modes and output formats, and the bugs worth
telling you about.
Pick exactly which WordPress content types to crawl
WordPress mode used to offer a fixed choice of posts and pages. Sites built by agencies keep most of their content somewhere else entirely.
Added Crawller now reads the site's own list of content types and shows every one it finds, by the name the site uses — Practice Areas, Case Studies, Team, Areas We Serve — each with its published count. Tick the ones you want.
Changed Detection results name content types properly too, instead of showing internal identifiers like mt_practice_areas.
Fixed Page-builder and SEO plumbing that reports a count but holds no readable content (block patterns, schema records, template libraries) is filtered out of the list.
WordPress crawls stopped losing content
Two bugs meant a WordPress crawl could quietly return a fraction of a site and still look like it had succeeded.
Fixed Custom post types are now crawled by default. A WordPress site's real content often lives in custom types — a law firm's practice areas, an agency's services, a team directory — and the old default of posts and pages silently left all of it out. On one test site that was the difference between 9 items and 36.
Fixed A failed request no longer deletes an entire content type. WordPress REST calls skipped the retry logic that ordinary page fetches use, so a single timeout or 502 dropped everything of that type and the crawl still reported success. Those calls now retry with backoff and honour Retry-After.
Added Incomplete crawls say so. Anything that could not be fetched is recorded per content type and flagged in the results panel, the Markdown header and the llms.txt header. A crawl that lost content can no longer pass for a clean one.
One page, one tool
The crawler moved to the home page. There is no longer a demo box on the front page and a separate advanced tool.
Changed The full crawler is now the home page, with no page limit imposed by the old hero demo. The previous /crawl address redirects, so existing links keep working.
Added Structured data across the site — the tool, its free pricing, a how-to, an FAQ and breadcrumbs — so search engines and AI crawlers can read what Crawller is without guessing.
Added A written FAQ covering rate limits, llms.txt, WordPress handling, orphan pages and robots.txt.
Fixed robots.txt pointed search engines at the wrong domain's sitemap, left over from the site template. It now points at crawller.dev, and explicitly welcomes GPTBot, ClaudeBot and PerplexityBot.
Fixed Corrected copy that was not true: the old feature list claimed no rate limits and no stored data. Crawls are limited to 12 per hour per IP, and completed crawls are saved so exports can be re-rendered without crawling again.
Single page mode, and a brake on runaway crawls
Added Single page mode. Extract one URL and nothing else — the fastest way to pull one article into a context window. It sits alongside Sitemap, WordPress and Follow links.
Changed Crawling no longer starts until you pick a method. Detection results and the advanced options open automatically, so the size of a crawl is visible before it runs.
Fixed The page limit no longer scales itself up to match the site. Detecting a 1,300-page site used to raise the cap to its maximum, which read as “No limit” on the slider — people hit Crawl and pulled thousands of pages by accident. Large sites now hold at a conservative default until you raise it.
Added A proper share image and complete Open Graph and Twitter card metadata. Links to crawller.dev had been previewing the unmodified site template.
The HTTP API and MCP server
The crawler became something other programs could call, not just a web page.
Added A public HTTP API: probe a site, start a crawl, poll its progress, and re-render a finished crawl in any output format.
Added An MCP server over both STDIO and HTTP, exposing inspect_site, crawl_website and crawl_page, so Claude and other MCP clients can crawl a site inside a conversation.
Added Site detection: before crawling, Crawller checks for a sitemap and an open WordPress REST API and recommends whichever reaches the most pages.
Added Export formats — Markdown, llms.txt, llms-full, plain text, clean JSON and raw JSON — re-renderable from a saved crawl without crawling again.
Added Request safeguards: URL validation that refuses private and reserved addresses, and per-IP rate limiting.
First build
Added The crawler engine: concurrent breadth-first crawling, sitemap discovery through robots.txt, sitemap indexes and gzipped sitemaps, WordPress REST enumeration, and content extraction through Trafilatura.
Added crawller.dev — the site, the warm Claude-inspired design system, and the first browser-based crawl box.