Skip to content
HomeBlogRate limits on the crawler are a feature...
Search limits

Rate limits on the crawler are a feature

Search limits: How Search works rate limits + robots; politeness ≠ outage; distinct from finite-crawl essay.

O

Oernoe Editorial Team

Writer

Published September 8, 202610 min read
How Oernoe Search works does not describe crawl rate limits as a bug that slipped past marketing. It lists them next to robots.txt respect as ordinary crawler behavior: continuous crawling, respecting robots files and crawl rate limits set by website owners, privacy-first crawling that does not store identifying data during the fetch, and content analysis of pages and links. This essay is Search limits writing about that politeness as a feature. It is not a remake of the finite-crawl essay that taught crawl description is not a coverage promise, and it is not a remake of the finite-index essay that taught missing as ordinary. Those hinges already exist. The hinge here is narrower: a crawler that slows down or refuses a path on purpose is doing the job the guide advertised. Politeness is not an outage.

## What the guide actually says

How Oernoe Search works explains that Search keeps an index of public web pages it has fetched. Crawlers visit pages they already know about, follow public links, and come back when a page changes. Continuous crawling means the work runs around the clock to stay current with new and updated content. Respecting robots means the system honours robots.txt files and crawl rate limits set by website owners. Privacy-first crawling means the company does not store IP addresses, device information, or identifying data during crawling, does not extract email addresses for advertising lists, and does not track user identities or store behavioral signals as part of that fetch. Content analysis means page content, structure, and links are used to understand relevance.

The FAQ keeps the ambition honest. Search is a working engine at search.oernoe.com, not a claim that the company matches every other index in size. The difference it will stand behind is narrower: Search does not use personal advertising profiles to rank results, and the publisher site discloses Google ads when they appear on selected pages such as that guide. Funding comes from optional premium features and contextual Google ads on selected publisher pages. Search queries are not sold or used to build advertising profiles. None of that funding language buys infinite crawl speed. None of it converts a rate limit into a Status incident.

The Search homepage prints the quiet half of the same limit. Results come from websites and knowledge entries in the Oernoe index. Not every page on the web is indexed. That footer is about the stored set. Rate limits are about the process that feeds the set. You can accept a finite index and still misread a slow or refused crawl as the product being down. This page exists to block that misread.

## Robots and rate limits are two politeness tools

Robots.txt is the hard public rule. A publisher that disallows a path is telling crawlers not to fetch it. Search that honours the rule is behaving as documented. The absence that follows is not a secret demotion. It is compliance. Crawl rate limits are softer. They trade speed for host health. A site owner can ask crawlers not to hammer a server. A crawler that ignores that request in the name of “better coverage” is not being more useful. It is being a bad neighbor. How Search works refuses that theater by naming rate limits in the same breath as robots.

Those tools serve different failure modes. Robots answers “may we fetch this path at all.” Rate limits answer “how hard may we fetch the paths we are allowed.” A site can allow crawling and still need the crawler to back off. A site can disallow a path and need no rate conversation because the fetch should not happen. Readers who collapse both into “Search is slow, so Search is broken” lose the ability to tell compliance from incident.

Corrections’ robots story on www is related but not identical. The publisher once shipped a robots.txt that missed live application routes because trailing-slash Disallow lines did not match bare paths. Empty shells became fetchable. The fix used noindex headers and a narrower Disallow set so crawlers could read the header. That episode is about the publisher’s own routes and about writing a rule that matches the live URL shape. Search’s robots respect is about other people’s routes. Both stories teach the same discipline: write the rule you intend, then live with what the rule excludes. Rate limits belong to that discipline. They are not a temporary embarrassment until the index is “done.”

## Politeness is not Status red

Tracker.oernoe.com and the public Status surface answer whether named services are operational. A green Status line means the product hosts the company watches are up. It does not mean every public URL on the open web has been fetched today. It does not mean a publisher’s robots.txt has been overridden. It does not mean crawl rate limits have been lifted because a user is impatient. Treating a polite refusal or a slow revisit schedule as an outage asks Status to lie about a different machine.

How Search works also separates ranking from crawling. Content relevance, authority, link analysis, freshness, aggregated satisfaction signals, mobile-friendliness, and page speed can order documents the index already holds. They cannot rank a page the crawler never fetched. They cannot justify ignoring a robots disallow. They cannot justify melting a small host because a coverage slogan sounded better in a pitch. Privacy-preserving ranking is a claim about how stored documents are ordered and about what personal advertising files are not built. It is not a license to crawl rudely.

Continuous crawling still matters. The guide says crawlers work continuously to stay current. Continuity is not the same as unbounded concurrency against every host. A feature can be always on and still be capped. Readers who hear “24/7” and translate it into “no rate limits” are reading marketing into a sentence that already named rate limits in the next bullet.

## Distinct from finite-crawl and finite-index essays

Finite crawl is not a promise of coverage already taught that describing the crawl machine is not an SLA that your URL will be covered. That hinge was honesty versus coverage promise. This essay assumes that honesty and focuses on one mechanism inside it: rate limits and robots as intentional politeness. You can accept that crawl prose is not a coverage promise and still treat a rate-limited fetch as a defect. The refusal here is that treatment.

Finite index means missing is ordinary already taught acceptance of absence in the stored set. That hinge was missing as normal. This essay is about process ethics on the way into the set. A missing URL can be ordinary because the index is finite, because robots blocked the path, because the crawler has not revisited yet, or because a rate limit delayed the fetch. Those causes are not interchangeable. Rate-limit politeness is a cause you should want even when you dislike the delay.

How Search works guide is not the results page already taught hostname and job: documentation on www, retrieval on search.oernoe.com. Search tips are limits printed in public already taught that tip footers are public limits, not magic. This essay stays inside Search limits and centers crawler politeness rather than tip syntax or guide-versus-results category.

## Funding and privacy do not buy rude crawl

How Search works answers how the company makes money if Search does not profile users: optional premium features and contextual Google ads on selected publisher pages such as that guide. Privacy Policy section 3 and section 4 keep the split: no sale of account or search history; no search-history advertising profiles; selected substantial www pages may use AdSense; Search itself is not that inventory. Cookie Policy adds that Search queries typed at search.oernoe.com are not an input to those units.

That model does not fund infinite aggressive crawl. Expecting publisher ads on www to force faster third-party fetches is a category error. Expecting privacy-first language to mean “we ignore robots because we are the good guys” is a worse error. Privacy-first crawling in the guide is about what identifying data is not stored during the fetch and about not harvesting addresses for advertising lists. It is not a claim that politeness is optional. A privacy story that required rude crawl would not be a privacy story worth keeping.

## How a reader should use crawl politeness

When a page is missing from Search, ask which machine failed before you file a Status theory. Did the publisher disallow the path. Is the URL new enough that continuous crawling has not revisited yet. Is the index simply finite. Did a rate limit delay a fetch that will happen later. Those questions produce useful support mail. “Your crawler is broken because it did not hammer this host” does not.

When you read How Search works, hear robots and rate limits as product promises, not as temporary shame. When you operate a site you want in the index, write a robots.txt that matches live paths, and understand that asking for polite rates is asking Search to behave as documented. When you file support@oernoe.com about emptiness, bring the query, the hostname, the time, and what a known-good query returned. Do not bring a theory that Status green required unbounded crawl.

When you read Corrections’ robots episode, keep the categories straight. That was the publisher fixing its own rule so shells would not pad the site. It is evidence that the company takes robots mechanics seriously. It is not evidence that Search should stop respecting other people’s robots or rate limits.

## Why calling politeness a feature matters for support

support@oernoe.com already receives emptiness reports that mix three jobs: a Status question, a Search coverage question, and a publisher overclaim question. Rate-limit honesty shrinks that mix. If the crawler slowed because a host asked it to, the useful reply is process, not an incident ticket. If robots blocked the path, the useful reply is compliance, not a secret demotion theory. If the index is simply finite, the useful reply is the Search homepage limit, not a promise to fetch everything tonight. Treating politeness as a feature gives the desk language that matches How Search works instead of improvising coverage apologies the guide never made.

## What this essay refuses

It will not invent crawl counts, concurrency figures, or index sizes. It will not claim every public company has a Knowledge panel. It will not claim a page published minutes ago must already be present. It will not claim Status green enlarges crawl coverage or lifts rate limits. It will not walk Search chrome control by control. It will not treat the guide as the results page. It will not claim that respecting robots is optional when a site owner prefers to rank. It will not clone the finite-crawl coverage-promise essay or the finite-index acceptance essay under a new title.

Rate limits on the crawler are a feature because How Search works named them as respect, not as failure. Robots.txt and polite rates keep Search from becoming a coverage story that burns other people’s hosts. Politeness can delay a fetch. Delay is not downtime. A finite, polite crawl is the machine the company documented. Readers who need a different machine should say so clearly. Readers who need this machine should stop translating its ethics into an outage.
O

Oernoe Editorial Team

Writes for the Oernoe Journal. Questions about this article can go to the contact page.

Get in touch

Related Articles

Want to Learn More?

Explore our complete guides and knowledge base for more insights on privacy, technology, and best practices.

Browse Our Guides