Free website crawler
Turn any website into clean, LLM‑ready data.
Paste a URL. Crawller works out whether the site publishes a sitemap or opens the
WordPress REST API, crawls it whichever way reaches the most pages, and hands it back as
Markdown, llms.txt, text or JSON.
Four ways to find every page
Most crawlers only follow links, so they stop at whatever the menu happens to point to. Crawller inspects the site first and tells you which route reaches the most pages.
Sitemap
Reads robots.txt and every common sitemap location, follows sitemap indexes and gzipped sitemaps. Reaches orphan pages nothing links to — usually the most complete crawl of a static site.
WordPress REST API
Enumerates published content straight from /wp-json. The most complete option on WordPress, and several times faster, because there is no HTML to parse.
Follow links
Walks internal links breadth-first from the start page, as deep as you tell it to go. Works on any site, but only reaches pages something links to.
Single page
Extracts the one URL you paste and nothing else. The fastest way to pull a single article into an AI context window.
Built for context windows, not for browsers
Extraction runs through Trafilatura, so what comes back is the article — not the nav, the cookie banner and four hundred lines of tag soup. Export the same crawl in any format.
Markdown
Clean, readable markdown per page, with the navigation, footers and cookie banners stripped out.
llms.txt
The emerging standard for handing a whole site to a language model as one file.
Plain text
Body copy only, for search indexes, embeddings and classic NLP pipelines.
JSON
Every page with its title, description, metadata, word count, schema.org blocks and link graph.
How to crawl a website
01
Paste a URL
Crawller checks the site before it crawls anything — whether it publishes a sitemap, whether the WordPress REST API is open, and how many pages each route would reach.
02
Pick how it crawls
Take the recommended method or choose your own, then set the page limit, depth and path filters. Nothing starts until you choose.
03
Export the content
Read it in the browser, download it as Markdown, llms.txt, text or JSON, or call the same crawler from Claude through the MCP server.
Crawl from inside Claude
Crawller is also an MCP server. Connect it once and your assistant can inspect a site, crawl it and read the markdown back without you leaving the conversation.
claude mcp add --transport http crawller https://crawller.dev/api/mcp.php Questions
Is Crawller free?
Yes. Crawller is free to use with no account and no API key. Crawls are rate limited to 12 per hour per IP address to keep the service available to everyone.
What is llms.txt and why would I want it?
llms.txt is a convention for publishing a site's content as a single, clean text file that a large language model can read in one pass. Crawller renders any crawl as llms.txt, so you can hand a whole site to an AI assistant without pasting page after page.
How does Crawller handle WordPress sites?
If a site exposes the WordPress REST API, Crawller reads published posts, pages and custom post types directly from /wp-json instead of parsing HTML. That reaches everything published, not just what the menus link to, and runs several times faster. When a page builder leaves the rendered content empty, Crawller falls back to fetching the live page.
Will it find pages that are not linked from the menu?
Usually, yes. Sitemap mode seeds the crawl from robots.txt and every common sitemap location, including sitemap indexes and gzipped sitemaps, so it reaches orphan pages that no internal link points to. A link-following crawl, by definition, cannot.
Can I use Crawller inside Claude or another AI assistant?
Yes. Crawller ships an MCP server over both STDIO and HTTP, exposing three tools: inspect_site, crawl_website and crawl_page. Once it is connected, you can ask your assistant to crawl a site and it gets the markdown back directly.
Does Crawller respect robots.txt?
It can, and there is a switch for it in the advanced options. It is off by default because the common case is crawling a site you own, where robots.txt rules written for search engines would hide your own content from you.
How many pages can one crawl cover?
Up to 1,000 pages per crawl through the web tool. The page limit starts at a conservative 150 so a large site never runs away with a crawl you did not intend; raise it with the slider when you want the rest.