Skip to content
HomeBlogAllow all except /api/ and /debug/...
Search limits

Allow all except /api/ and /debug/

Search limits: live robots.txt Allow / with Disallow /api/ /debug/; what that means for crawl honesty; distinct from robots politeness and x-robots essays.

O

Oernoe Editorial Team

Writer

Published September 8, 202610 min read
The live robots.txt on www.oernoe.com is short on purpose. Mediapartners-Google is allowed across the site. Google-Display-Ads-Bot is allowed across the site. The ordinary user-agent is allowed across the site, with Disallow only for /api/ and /debug/. The file declares the host as https://www.oernoe.com and points at the public listing of URLs the company claims to publish. This essay is Search limits writing about that Allow-almost-everything design. It is not a remake of the robots-txt-as-politeness essay. It is not a remake of the X-Robots-Tag fix essay. Those hinges already exist. The hinge here is crawl honesty on the current file: what Allow / with only /api/ and /debug/ blocked means, and why that is more honest than a long Disallow list that pretends thin routes are invisible.

## What the live file says, line by line

A guest fetch of https://oernoe.com/robots.txt on 8 September 2026 returned the current rules. Ads crawlers are allowed so they can evaluate publisher inventory. The wildcard user-agent is allowed for ordinary crawling. Disallow lines name /api/ and /debug/ only. Host points at https://www.oernoe.com. The public listing of claimed URLs is declared for discovery of finished documents. There is no trailing-slash theater. There is no attempt to hide application shells behind mismatched path strings.

What we publish ends with the matching mechanical checks. robots.txt should allow Google’s ads crawlers. Application shells send a noindex header; they are not hidden behind a trailing-slash Disallow, because that stopped Google from reading the header. Corrections, updated 26 August 2026, tells the older story that forced the current shape. robots.txt used to disallow paths with a trailing slash for application shells. The live routes were the bare paths. Google could fetch the bare URLs, which rendered as near-empty shells. A reviewer who typed those addresses saw a site padded with empty pages. The fix is the opposite of blocking more paths. Shells now send X-Robots-Tag noindex, nofollow. robots.txt only disallows /api/ and /debug/, so a crawler can fetch the URL and honour the header.

Allow all except those two prefixes is therefore not a lazy default. It is the corrected design after a concealment attempt failed.

## What crawl honesty means here

Crawl honesty means the written robots rules match the live URL shapes, and the tools do the jobs they can actually do. A Disallow line is an instruction not to fetch. It is not an instruction to drop a URL after a fetch. A noindex header is a ranking instruction that requires a fetch to arrive. Confusing those tools is how a publisher stays indexed by accident. Corrections states the lesson in one sentence worth keeping: a Disallow line that stops Google from reading noindex is how you stay indexed by accident.

/api/ and /debug/ are the paths the company currently treats as not-for-fetch. They are machine surfaces, not publisher documents. Leaving them Disallowed is ordinary. Leaving almost everything else Allowed is also ordinary once shells are handled with headers instead of pretend invisibility. Ads crawlers need Allow so they can evaluate eligible inventory named in Privacy and Cookies. Ordinary crawlers need Allow so they can fetch finished guides, How-To, company pages, and journal documents that still belong in the public offer. Shells that still resolve for old links or app entry need to be fetchable so their noindex headers can be read.

How Oernoe Search works lists robots respect next to crawl rate limits as ordinary crawler behavior on the open web. Continuous crawling means the work runs around the clock. Respecting robots means honouring robots.txt files and rate limits set by website owners. Privacy-first crawl claims in that guide say the company does not store identifying data during crawling as advertising harvest. Ranking refuses search-history advertising profiles. Funding is optional premium features and contextual ads on selected publisher pages. Queries are not sold or used to build advertising profiles. Search honouring other people’s robots files is neighbor behavior. The publisher publishing a short, accurate robots file on www is self-description. Both refuse the fantasy that a long Disallow list is the same thing as a clean index.

## Why a long Disallow list would be less honest

A long Disallow list looks strict. It can be theater. If the strings do not match live routes, crawlers fetch the real pages anyway. If the strings do match and block fetches of shells that needed noindex headers, crawlers never receive the ranking instruction. If the list tries to hide thin promotional URLs that still belong in a public Corrections story, the company is choosing concealment over the recovery rule in What we publish. Promotional stubs that stay up for old links are supposed to be noindexed and de-monetized, not secretly unfetched while still resolving for humans.

Allow / with narrow Disallow lines forces other honesty tools to do their jobs in public. What we publish separates pages that belong in the index from pages that stay up but leave the index. Corrections defines a correction as the old wording no longer being the page, or the URL being noindexed, or both. Privacy Policy section 4 keeps short or promotional posts, legal pages, error pages, and protected application pages out of AdSense. Cookie Policy treats application shells as non-inventory. The publisher can allow ads crawlers to fetch and still refuse to monetize hollow routes. That pair only works if robots.txt is not pretending those routes do not exist.

Finite index language also belongs here. How Search works explains that Search keeps an index of public pages it has fetched. Missing is ordinary. A path Disallowed under /api/ or /debug/ should be missing because it was never fetched under the rule. A shell that sends noindex should be missing from the public offer even if a person can still type the URL. Status on Tracker answers whether named services are up. It does not answer whether a foreign index has finished dropping a noindexed shell. Corrections does not promise same-day disappearance from every cache. It promises a public list that gets longer when the company is wrong.

## What this design is not

It is not a claim that every Allowed URL is high quality. Index quality still depends on What we publish, Editorial Standards, and Corrections. Indexable journal articles still need enough original text to be a document, currently at least seven hundred words, and still fail if the title is hype or the body is an unfinished launch note. Promotional exclusions still remove brochure from the offer. Overlapping privacy essays still had to be given distinct jobs. Thin company pages still had to be expanded. Allowing a crawler to fetch a URL is not a promise that the URL deserves to rank.

It is not a claim that Search ignores robots.txt on the open web. The politeness essay already took neighbor behavior. This essay takes the publisher’s own short file as a Search-limits fact: crawl honesty on www looks like Allow / plus two Disallow prefixes, not like a cloak. It is not a claim that X-Robots-Tag is unnecessary. The header essay already took that tool. This essay takes the robots side of the same corrected design: without Allow on shells, the header cannot be read.

It is not a claim that /api/ and /debug/ are the only sensitive surfaces on the network. Login lives at account.oernoe.com. Product work lives on product hosts. Privacy still admits ordinary technical and security logs. Ads still process data under Google’s policies on eligible www pages. The robots file is not a privacy policy. It is a fetch contract.

## Product truth still sits on the homepage

www.oernoe.com is the public publisher: journal, guides, How-To, company pages, and legal. Search lives at search.oernoe.com. Login lives at account.oernoe.com. The live Services list sends people to Search, Health, Chat, AI, Docs, Drive, and Tracker. Unreleased product routes are not sold as joinable homepage products. About says that if a page here disagrees with the homepage directory, the homepage is the one to trust until the page is fixed. Corrections already had to unwind a coming-soon capture page that collected addresses for software that did not exist; that route is now a status notice, noindexed, without ads. robots.txt is not how you make an unfinished route look finished. Allowing the fetch so a noindex header can work is how you stop unfinished routes from ranking as if they were essays.

Public writing is issued as Oernoe Editorial Team unless a person is named. The company is Anoepal. The founder is Angel Mejia Rodriguez. Product claims get checked against the homepage. Policy claims get checked against Privacy, Cookies, and Terms. If those sources disagree, the article loses. A robots file that disagreed with live routes already lost once. The current file is the corrected agreement.

## Ads crawlers Allowed is part of honesty, not a loophole

Allowing Mediapartners-Google and Google-Display-Ads-Bot across the site looks strange to readers who think robots.txt exists only to hide things. On this publisher it is the opposite signal. Eligible inventory lives on finished How-To, finished guides, and substantial journal articles that still pass originality and length checks. Ads crawlers need to fetch those documents to evaluate them. Blocking ads crawlers while asking to join a publisher network would be theater: ask for review, then refuse the fetch that review requires.

Privacy and Cookies already limit where units may load. Allowing an ads crawler to see Home or About does not place a unit there. Allowing an ads crawler to see a noindexed shell does not monetize the shell. The fetch contract and the inventory map are different tools. Crawl honesty keeps them both readable instead of stuffing inventory policy into Disallow lines. About’s corrections section names the old robots miss among public failures. The current Allow-all-except design is one repaired item on that list, not a claim that the list is finished.

## What a reviewer should verify

Fetch robots.txt and read it end to end. Confirm Mediapartners-Google and Google-Display-Ads-Bot are Allowed. Confirm the wildcard Allow. Confirm Disallow only for /api/ and /debug/. Confirm Host points at https://www.oernoe.com. Open Corrections and read the robots trailing-slash entry end to end. Open What we publish and read the closing mechanical checks. Open How Search works for robots respect and rate limits as Search-side behavior. Open Privacy section 4 if you need the inventory map that ads crawlers are Allowed to evaluate. Do not invent crawl volume numbers. Do not invent index sizes. Do not treat a short robots file as proof that every shell on the internet has already dropped from every foreign cache.

## What this essay refuses to invent

It does not invent a second robots philosophy beyond the live file. It does not invent Disallow lines that are not there. It does not treat Allow / as a quality certificate for every URL. It does not remake politeness theory or the header fix. It trusts the live sources: Allow all except /api/ and /debug/ is the corrected crawl-honesty design, ads crawlers are Allowed on purpose, shells rely on headers rather than pretend invisibility, and Search limits follow from rules that match live routes.

Allow all except /api/ and /debug/ is crawl honesty printed in public. A reviewer who wants concealment theater will be disappointed. A reviewer who wants a fetch contract that matches the site will find the current file shorter, and more useful, for that reason.
O

Oernoe Editorial Team

Writes for the Oernoe Journal. Questions about this article can go to the contact page.

Get in touch

Related Articles

Want to Learn More?

Explore our complete guides and knowledge base for more insights on privacy, technology, and best practices.

Browse Our Guides