Free tools · no sign-up
robots.txt tester
Paste a robots.txt, choose a crawler and a URL, and see whether that crawler may fetch it and which line decides.
Checked in your browser as you type. Nothing you paste is sent to domscout.
Blocked - line 2, "Disallow: /admin/", is the longest matching rule in the User-agent: * group (line 1).
Sitemaps declared
- https://example.com/sitemap.xml
How a crawler reads robots.txt
- A crawler obeys the group whose User-agent line names its product token, ignoring case. Only when no group names it does it fall back to User-agent: *. Groups that name the same crawler are combined.
- Within that group, the rule with the longest matching path decides. When an Allow and a Disallow match with paths of the same length, Allow wins.
- * matches any run of characters and a $ at the end anchors the end of the URL. Without $, a rule matches every URL that starts with it. Paths are case-sensitive.
- A robots.txt covers one protocol and host. https://example.com/robots.txt says nothing about https://www.example.com or http://example.com.
- robots.txt controls crawling, not indexing. A blocked URL can still appear in search results, without a description, when other pages link to it.
Check your head tags on every deploy
robots.txt is one way a deploy drops pages from search. A noindex that leaks from staging, or a canonical that points at a preview domain, is another. domscout reads the title, canonical, robots and JSON-LD a rendered page carries in one API call, so a build can fail before the change ships. See SEO regression checks or read the guide.
Questions
More free tools
Capture any public page as PNG, JPEG or WebP, at desktop or phone size.
Turn a web page into clean Markdown with a token estimate, ready to hand to an LLM.
See the title, description, og: tags and canonical URL a page publishes after its JavaScript runs.
Give your product a browser.
Get clean web content and visual proof into your workflow in minutes.