Skip to content
HomeBlogContinuous crawl is not complete coverage...
Company

Continuous crawl is not complete coverage

Continuous crawl sounds like a guarantee It is not In How Oernoe Search works under Web Indexing Crawling we say our crawlers work around the clock to visit pages and stay current with new and updated content That sentence describes…

O

Oernoe Editorial Team

Writer

Published September 8, 20269 min read
Continuous crawl sounds like a guarantee. It is not. In How Oernoe Search works, under Web Indexing & Crawling, we say our crawlers work around the clock to visit pages and stay current with new and updated content. That sentence describes a schedule and a habit. It does not describe the size of every other engine’s index, and it does not promise that every public URL on the open web sits in ours.

This journal note exists because those two readings get mashed together. People hear “continuous” and hear “complete.” Operators who run small or mid-size indexes hear the same mash from reviewers who want a single comparison number. We will not invent that number. We will explain what continuous crawl actually buys, what it refuses to buy, and how to read our own guides without turning a maintenance loop into a coverage claim.

## What the guide actually says about the index

Oernoe Search keeps an index of public web pages it has fetched. Crawlers visit pages they already know about, follow public links, and come back when a page changes. The index is what lets Search return results without asking every site on the web in real time. That last clause matters. Live fetch of the whole web for every query is not how any serious engine works. You fetch ahead of time, store what you can defend, and serve from that store.

Continuous crawling is the refresh loop on top of that store. It is how a page that changed yesterday can matter tomorrow. It is how a newly linked page can enter discovery if a crawler already has a path to it. It is not a census of every host that ever answered on port 443. Discovery still depends on seeds, link following, robots rules, rate limits, and the simple fact that some sites never want a bot at all.

The same guide is blunt about comparison: Search at search.oernoe.com is a working search engine, not a Google clone and not a claim that we match every other index in size. The difference we will stand behind is narrower. Search does not use personal advertising profiles to rank results, and the publisher site on www.oernoe.com discloses Google ads when they appear on selected finished pages. Size bragging is not on that list.

## Continuous is a verb, not a scoreboard

A continuous process can still be incomplete. Your dishwasher can run continuously in a restaurant and still leave glasses unwashed if the rack never held them. A crawler that never sleeps can still miss a URL that no known page linked, that robots.txt blocked, that returned soft errors for weeks, or that lived behind a form the bot does not submit.

When we say crawlers work continuously, we mean the queue does not wait for a quarterly batch to wake up. Freshness work happens as capacity allows. Pages already in the graph get revisited. New edges get followed when we find them. Content analysis looks at page content, structure, and links to understand relevance. None of that math produces a certificate that says “equal to Engine X’s document count.”

If someone wants a scoreboard, they want a different product conversation. Index size is interesting to engineers and competitors. It is a poor proxy for whether a result set is useful for a given query. A smaller index that ranks honest pages can beat a larger index full of thin mirrors. We care about ranking without an advertising profile of the querier. We care about respecting robots.txt and crawl rate limits set by site owners. We care about not harvesting personal data during crawling. Those are the promises on the guide. Coverage completeness is not one of them.

## What privacy-first crawling does not inflate

The crawling section also says we do not store IP addresses, device information, or identifying data during crawling in the sense of building an advertising harvest. We do not extract email addresses for advertising lists, track user identities, or store behavioral signals for that purpose. We index publicly visible content and how pages link to each other.

That privacy stance does not magically enlarge the index. If anything, it removes a whole class of “coverage” that would be unethical to claim: coverage of private inboxes, logged-in views, and personal graphs. Continuous crawl of the public web is already hard. Pretending that privacy-first rules somehow invent more public documents would be dishonest. The index grows when public pages are discoverable and fetchable under the rules we publish. It does not grow because we refuse to build an ad profile of a searcher.

Readers sometimes flip the logic. They assume that if we crawl continuously and we refuse personal harvest, we must be compensating with a gigantic public crawl that rivals the largest commercial indexes. That compensation story is fan fiction. Capacity, politeness, and discovery limits still apply. Continuous does not cancel them.

## Distinct problems that look similar from far away

Several neighboring claims deserve their own essays and should not be collapsed into this one.

A finite crawl budget is not a promise of coverage. Every operator has a budget. Continuous scheduling does not make the budget infinite. Rate limits on the crawler are a feature for site owners, not a confession that we are hiding pages. Privacy-first crawl is not an infinite index. Search is not a Google-clone claim. When results look thin, the honest move is to read status and to understand what the index actually holds, not to demand a marketing paragraph about total URL counts.

This piece is only about one confusion: taking the continuous-crawl bullet and reading it as complete coverage. If you need the other distinctions, leave this URL and open those. Do not paste them into the same paragraph and call it clarity.

## How ranking sits on top of an incomplete index

Ranking factors in the same guide assume an indexed page exists to evaluate. Content relevance asks how well page content matches the query. Authority and expertise ask how trustworthy the site looks. Link analysis asks how many other reputable sites cite the page. Freshness, mobile-friendliness, and page speed are about the document as served. Aggregated satisfaction signals, when used, are processed at scale without tying data to individuals.

None of those factors can rescue a URL that was never fetched. Continuous crawl improves the chance that a changed or newly linked page enters the candidate set. It does not invent candidates from thin air. If a page is missing, the first diagnostic is discovery and fetch, not a ranking tweak. If a page is present but ranks poorly, the diagnostic moves to relevance and quality signals. Mixing those diagnoses wastes everyone’s time.

Equal treatment of queriers is also orthogonal. The algorithm treats searches equally unless the account holds custom preferences the person saved. Two people can get the same thin set for a rare phrase because the index is thin on that phrase, not because one person was profiled. Continuous crawl may thicken that set over time. It still will not match a fantasy of universal coverage.

## What operators and reviewers should ask instead

Ask whether the crawl respects robots.txt. Ask whether rate limits are honored. Ask whether the product claims to match another engine’s size. Ask whether ranking uses a search-history advertising profile. Ask where Google ads appear on the publisher hostname, and whether Search queries are sold. Those are checkable claims with public pages behind them.

Do not ask us to pretend continuous crawl equals complete coverage. Do not ask for a fake parity number against a competitor. Do not treat a missing URL as proof that the continuous bullet was a lie. Missing URLs are normal in any index that is honest about its edges.

If you publish on the open web and want a better chance of discovery, link from pages that already sit in public graphs, keep robots rules clear, and serve stable content. That is ordinary web hygiene. It is not a back-channel into our ranking. It is not a guarantee of inclusion. Continuous crawl will meet your page when discovery and capacity allow. That is the whole deal.

## How this publisher site talks about Search without chrome

www.oernoe.com is the publisher site: journal, guides, How-To, company pages, and legal. It is operated by Anoepal. Search itself lives at search.oernoe.com. Mixing those jobs produces landing pages that say nothing useful. We already had to correct homepage copy that sounded like the whole network had no tracking while this hostname was applying for AdSense. We will not correct continuous crawl by inventing coverage myths in the other direction.

Selected finished publisher pages may show Google ads. That is separate from how Search treats queries. We will not claim the whole network is free of advertising systems. We will not use journal space to sell a size fantasy. Journal articles that stay in the public index need original text about behaviors we actually run. Continuous crawl is one of those behaviors. Complete coverage is not.

## A short reading list inside our own docs

Read Web Indexing & Crawling in How Oernoe Search works for the continuous-crawl bullet and the robots and privacy-first notes. Read the common-questions section for the explicit refusal to claim size parity with every other engine. Read Search queries and Google ads for the split between Search and publisher advertising. Read About for what is live on the homepage directory and what is not sold as live. Read What we publish if you care why thin promotional stubs leave the index.

Those pages agree on a boring truth. We crawl on a continuous schedule to keep the index current. We still operate inside discovery limits, politeness rules, and capacity. Continuous is how we refresh. Complete is a different word. We will keep using the first and refusing the second.

## Closing

If you remember one sentence from this essay, remember this: continuous crawl is a maintenance loop, not a coverage certificate. The guide’s 24/7 language is about staying current with pages we can reach under published rules. It is not a promise that our index matches another engine’s size, and it is not a promise that every public URL is covered. When results are thin, diagnose the index you have. When a page is new, wait for discovery and revisit. When someone sells you continuous as complete, ask which guide they are quoting. Ours does not say that.
O

Oernoe Editorial Team

Writes for the Oernoe Journal. Questions about this article can go to the contact page.

Get in touch

Related Articles

Want to Learn More?

Explore our complete guides and knowledge base for more insights on privacy, technology, and best practices.

Browse Our Guides