Allow / disallow rules
Parses every user-agent group and counts the Allow and Disallow rules so you can see what's blocked at a glance.
A free robots.txt checker, tester and validator in one: fetch any site's robots.txt, validate every rule and sitemap line, and test whether a specific URL is allowed or blocked for Googlebot, Bingbot, or any other user-agent. It replaces the robots.txt Tester Google removed from Search Console.
Start your 7-day trial — no credit card, free plan after.
Enter a site to fetch its robots.txt. Optionally add a path and user-agent to test whether that path is allowed or blocked.
Free check. No signup. Results are not published or indexed.
Enter a site to fetch its robots.txt. Optionally add a path and user-agent to test whether that path is allowed or blocked. Recurring checks are configured inside the NorthDuty app.
A single check parses the live robots.txt and tells you exactly how crawlers will treat your URLs.
Parses every user-agent group and counts the Allow and Disallow rules so you can see what's blocked at a glance.
Enter a path and user-agent to get an allowed-or-disallowed verdict using longest-match with Allow-wins-on-tie, the behaviour major crawlers use.
Lists any Sitemap: lines so you can confirm search engines are being pointed at the right sitemaps.
Tells you when no robots.txt exists — in which case everything is crawlable by default.
A single stray Disallow can deindex an entire section of a site, and it's easy to ship one by accident.
No signup — enter a site and optionally a path to test.
Provide a domain or URL. NorthDuty fetches robots.txt from that site's root.
The file is parsed into user-agent groups with their Allow and Disallow rules and any sitemap directives.
Add a path and user-agent to get an allowed-or-blocked verdict, with the matching rule shown.
NorthDuty's health checks cover SEO fundamentals including robots.txt, so an accidental block gets flagged on a schedule.
Google retired the legacy robots.txt Tester from Search Console in late 2023. Its replacement, the robots.txt report, shows which robots.txt files Google found for your site, when they were last crawled, and any fetch or parse errors. What it no longer does is let you paste a URL and see whether a rule blocks it.
That per-URL check is what this tester does. Enter your site, a path, and a user-agent, and it applies the same matching Google documents: the most specific user-agent group wins, then the longest matching rule, and Allow wins a tie. Use it before you ship a robots.txt change, then use URL Inspection in Search Console to confirm how Google treats a specific live page.
Every directive and pattern you are likely to need, and how Google treats it.
| Directive or pattern | What it does | Example |
|---|---|---|
| User-agent | Starts a group of rules for one crawler. * matches any crawler that has no group of its own. | User-agent: Googlebot |
| Disallow | Blocks crawling of URLs whose path starts with the value. An empty value blocks nothing. | Disallow: /wp-admin/ |
| Allow | Re-opens a path inside a blocked section. The longer (more specific) matching rule wins. | Allow: /wp-admin/admin-ajax.php |
| Sitemap | Points crawlers at a sitemap. Must be a full URL and can appear anywhere in the file. | Sitemap: https://example.com/sitemap.xml |
| * wildcard | Matches any sequence of characters inside a path. | Disallow: /*?add-to-cart= |
| $ end anchor | Matches only when the URL ends exactly there. | Disallow: /*.pdf$ |
| Crawl-delay | Ignored by Google. Some other crawlers, such as Bingbot, respect it. | Crawl-delay: 10 |
| Noindex | Not supported in robots.txt. Google stopped honouring it in 2019; use a meta robots tag or X-Robots-Tag header. | (use <meta name="robots" content="noindex">) |
Run any of these through the tester above with your own paths to see the verdict and the rule that decided it.
| Rules | URL tested | Result |
|---|---|---|
| Disallow: /wp-admin/ + Allow: /wp-admin/admin-ajax.php | /wp-admin/admin-ajax.php | Allowed: the Allow rule is longer, so it wins |
| Disallow: /search | /search-results/shoes/ | Blocked: rules are prefix matches, so /search also matches /search-results/ |
| Disallow: /*.pdf$ | /files/guide.pdf?v=2 | Allowed: $ requires the URL to end in .pdf, and this one ends in ?v=2 |
| Disallow: /Checkout/ | /checkout/ | Allowed: paths are case-sensitive |
| User-agent: * with Disallow: / and User-agent: Googlebot with Allow: / | / as Googlebot | Allowed: Googlebot obeys only its own group and ignores the * group |
| Disallow: (empty) | Any URL | Allowed: an empty Disallow blocks nothing |
Most robots.txt damage is accidental and ships with a deploy or a migration.
A staging site's User-agent: * / Disallow: / gets copied to the live site during a launch or migration, and search engines stop crawling everything. This is the first thing to check after any relaunch.
Blocking a URL stops crawling, not indexing. A blocked page can still appear in results if other pages link to it. To keep a page out of the index, allow crawling and add noindex.
If Google cannot fetch your theme's CSS and JS, it cannot render the page properly, which can hurt how the page is understood and ranked. Leave asset folders crawlable.
If robots.txt returns a 5xx server error, Google temporarily treats the whole site as disallowed. A 404 is treated as no rules at all. Monitor the file, not just the homepage.
robots.txt only works at the root of each host and protocol. https://www.example.com/robots.txt does not cover https://shop.example.com/. Each subdomain needs its own file.
Blocking cart and add-to-cart parameters is fine, but a broad rule like Disallow: /*? can also block filtered category pages and paginated product lists you want crawled.
If there is no robots.txt file on the server, WordPress serves a virtual one: it disallows /wp-admin/, allows /wp-admin/admin-ajax.php, and since WordPress 5.5 adds a Sitemap line pointing at /wp-sitemap.xml. SEO plugins such as Yoast and Rank Math let you edit it from the dashboard, and a physical robots.txt file in the site root overrides the virtual one.
For WooCommerce stores, the pages worth keeping crawlable are products, product categories, and your main content. Cart, checkout, and account pages rarely need crawling, but they are better handled with noindex than with robots.txt, so Google can still see the tag. After any plugin or theme change, re-test a product URL and a category URL above.
Use the tool preview for a quick answer, then move into recurring monitoring for your most important pages and journeys.
Feature
Learn how NorthDuty combines editable user journeys and website health checks in one website monitoring project.
Explore Website MonitoringFeature
Monitor uptime every 5 minutes by default with HTTP, SSL, DNS, blank-page detection, broken resources, JavaScript errors, and API call tracking.
Explore Uptime MonitoringArticle
A one-off robots.txt test only describes today. Here is what breaks robots.txt in production, how Google reacts to each failure, and how to monitor the file.
Read robots.txt MonitoringArticle
Use this 41-point website monitoring checklist to cover uptime, SSL, DNS, forms, checkout, journeys, alerts, launch checks, and incident response.
Read Website Monitoring Checklist: 41 ChecksArticle
Why a website breaks after deploy without going down: the 5 most common post-deploy failures uptime monitors miss, and how to catch them before your users do.
Read 5 Ways a Website Breaks After a DeployPricing
NorthDuty plans are sized by how many checkout, signup and login journeys you monitor: Free, $29 Starter, $79 Pro, $199 Business. 7-day trial, no card.
Compare pricing plansAnswers about this diagnostic preview and when to move into recurring monitoring.
All three, and the difference is mostly wording. It validates the file by fetching it and parsing every user-agent group, checks a live site rather than pasted text, and tests a specific path against those rules for the crawler you choose.
Yes. It's free and requires no signup. Enter a site to fetch and parse its robots.txt, and optionally test a specific path.
It picks the most specific matching user-agent group, then applies the longest matching rule, with Allow winning ties — the same approach major search-engine crawlers use. Wildcards (*) and end-anchors ($) are supported.
If no robots.txt is found, the tester reports that — and by the standard, everything on the site is crawlable by default.
No. The tester only reads the live file. To change crawling rules you edit robots.txt on your own server.
No. Google removed the legacy robots.txt Tester in late 2023 and replaced it with the robots.txt report, which shows fetch status and parse errors but does not test individual URLs. This tool fills that gap: enter a path and user-agent to get an allowed or blocked verdict with the matching rule.
Not reliably. robots.txt controls crawling, not indexing. A blocked URL can still be indexed, usually without a description, if other pages link to it. To remove a page from results, allow crawling and use a noindex meta tag or X-Robots-Tag header.
At the root of the host, for example https://example.com/robots.txt. It applies only to that exact protocol and host, so subdomains need their own file. Google reads up to 500 KiB of the file and ignores anything beyond that.
The paths in Allow and Disallow rules are case-sensitive, so Disallow: /Private/ does not block /private/. Directive names like User-agent and Disallow are not case-sensitive.
A robots.txt mistake can quietly deindex pages for weeks. NorthDuty monitors SEO fundamentals continuously, so an accidental block gets caught fast.
7 days with Pro features and limits, no credit card — then keep one daily journey on the free plan.