The Quiet Bottleneck: Crawl Budget, Discovery and the Indexing Hub
A page that loads in your browser has not been found by anything. Around Charlotte the gap between those two states is unusually wide, because sites here multiply addresses across two states, a ring of fast-growing suburbs and a stream of housing developments that stop existing once they sell out.
The failure is silent, which is why it lasts for years. No error appears. Navigation works, the sitemap validates, pages render quickly. Meanwhile a large share of the site's addresses have never been requested by anything, and rewriting them cannot move a single figure.
Publishing puts a file on a server, and stops there
Consider a builder who launches a page for a new community in Indian Land on a Tuesday and expects it to pull searches for the neighborhood name by Friday. Nothing about publication makes that likely. Between the file existing and a person clicking it sit four separate decisions, each taken independently, each able to stall for months.
- Something has to point at the address. Until a link, a sitemap entry or a submission mentions it, the URL is private in every practical sense.
- Being known is not being scheduled. Far more URLs are on file than will ever be requested this quarter, and yours joins a waiting list nobody publishes.
- A fetch can be spent and wasted. A redirect hop, a slow response or a success code covering an empty page all consume the visit without delivering anything.
- Retention is a separate verdict. Content judged thin, duplicated or templated is fetched, assessed and dropped, leaving no evidence anywhere in your analytics.
This is why indexing sits in its own part of the workspace rather than inside the analytics screens. Those screens describe pages that already cleared everything above; a separate section for discovery is the only place the silent failures become visible.
The attention is finite and something else is using it
Each domain gets a loose ration of crawler attention. Two things set it: what the server can absorb without slowing down, and how much value the site has previously seemed to hold. A ten-page firm never bumps into the ration. A site that manufactures URLs from a template collides with it weekly, and manufacturing URLs is exactly what a service business in this metro drifts into.
The ration is spent on whatever can be reached, in no particular order of importance. Nothing walking your links can tell that one URL carries the revenue and the next one is a sort order a plugin bolted onto a listing.
Addresses that were never intended
Query parameters, session identifiers, print views, event calendars generating a page for every month ahead.
- Filter combinations multiply quickly
- Blog tag archives
- Pagination with no natural end
Fetches that return nothing
Hop-to-hop redirects, sluggish responses, and pages replying with a success code while displaying an apology.
- Community URLs redirected twice
- Empty inventory listings
- Staging copies left open
That last category deserves emphasis. A page reading "no availability in this community" while answering 200 is, from outside, a valid page containing almost nothing. Produce sixty and you have demonstrated at scale that this domain publishes empty documents — which costs you on the documents that are not empty.
Two states, a ring of suburbs, and a page for each pair
Most metros hand a service business one list of place names. This one hands over two, because the state line runs through the middle of the market. Fifteen minutes south of Ballantyne the licensing regime changes, the county changes, and — critically for anyone building pages — so do the names.
So the arithmetic assembles itself. Somebody writes down the towns: Huntersville, Cornelius, Davidson, Mooresville, Concord, Harrisburg, Kannapolis, Matthews, Mint Hill, Indian Trail, Monroe, Waxhaw, Weddington, Pineville, Belmont, Mount Holly, Gastonia. Then somebody adds the South Carolina side: Fort Mill, Tega Cay, Rock Hill, Indian Land, Lancaster, Clover, York. Somebody else lists the services. A publishing plan then exists that nobody actually wrote: one page per pair, generated by a plugin before lunch.
Two things then go wrong at once. The pages are indistinguishable except for one proper noun, so nothing reading them separates a town where three crews work daily from a town added because it sat between two others already listed. And the state line, which is a genuine business fact, gets expressed as duplication rather than as difference: the Fort Mill page and the Pineville page say the same sentence with a different name, when the honest version would explain that one requires a South Carolina license and the other does not.
That is where the allowance disappears, and it is also where a pattern gets learned. Something that fetches eighteen near-identical town pages slows down before the nineteenth — and the nineteenth might be Concord, where you actually have depth.
Subdivisions finish, and their pages do not
The second multiplier belongs to a market building houses this quickly. Developments open, sell through their phases and close. A builder publishes a page per community. A flooring company, a garage-door installer or a landscaping firm publishes one too, because that is where this year's work is. Then the last lot closes, the sales trailer leaves, and nobody deletes anything.
What remains is a body of addresses describing finished places, phases that no longer exist and inventory homes sold three years ago. They are not wrong — the neighborhood is still there — but they answer a question nobody asks, while holding the same weight in the queue as the page for the development opening next spring.
| Page type | Life expectancy | What it becomes when the project ends | Usual treatment |
|---|---|---|---|
| Community landing page | Two to four years | A neighborhood name with stale claims attached | Left live, never revisited |
| Phase or release page | Six to eighteen months | An address for a thing that no longer exists | Left live, still in the sitemap |
| Inventory or model listing | Weeks | An empty list answering with a success code | Regenerated endlessly by the CMS |
| Project or job-site write-up | Indefinite | Genuine local proof, if it names the place | Ignored, though it is the valuable one |
The last row is the exception worth protecting. A write-up naming the subdivision, the county and what was actually installed keeps earning long after the development closes, because it is the only page on the site that could not have been generated. Everything above it is a template with a date stamped on it.
What to do with places and projects that ended
Removing pages feels like giving ground, so it rarely happens and the list only grows. But the choice is not keep-or-delete. There are four dispositions, and picking the right one per address is the whole exercise.
Into an area page
One page per stretch — the Lake Norman towns, Union County, the York County side — carrying a paragraph and an honest note on travel and licensing.
- Old URLs redirect once, never twice
- Usually outranks both originals
Gone, and saying so
For closed phases and expired promotions that never earned an impression. A definite removal beats a page that quietly persists.
- Leaves the crawl queue sooner
- Nothing left to redirect badly
Reachable but not promoted
Pages a visitor may legitimately reach that should not compete: duplicate county variants, superseded campaign copies.
- Excluded from every sitemap
- Never entered into a batch
The short list that survives
Towns you genuinely staff, plus every project write-up that names a real address, get written individually with photographs and specifics.
- Reachable from the primary menu
- At the front of the submission queue
How you redirect matters unusually much here, because these sites have been reorganized several times over as the metro expanded. Folding forty community pages into six area pages produces forty redirects, and where those destinations were themselves relocated in a redesign two years back, each redirect is now a chain. Every hop is a separate request billed to the same ration: three to deliver one document, and the error detail held per address is where chains surface.
One caution about the state line. Do not fold a South Carolina page into a North Carolina one merely to shorten the list. Where licensing genuinely differs the two pages answer different questions, and a buyer in Rock Hill checking whether you can pull a permit on their side needs to find that sentence somewhere.
The sitemap is where you say which addresses count
Sitemaps are usually treated as a dump: everything the CMS knows about, written to XML on a timer. Read instead as a statement — these are the addresses I am claiming deserve your attention — the file becomes useful, and a dump claims nothing, because it contains the closed phases too.
Shape carries the other half of the value. When 900 URLs sit in one undivided file and coverage drops by a third, the report tells you a third of something is missing and nothing more. Break the file up by purpose — services, staffed towns, area pages, project write-ups, editorial — and every gap arrives labeled. Division also handles growth gracefully: taking on the South Carolina counties adds one child file rather than forcing a full rebuild you then have to inspect.
Give it the whole tree, not one file
Built for sites where URLs are produced automatically and the total surprises everyone.
- Nested files are walked for you. An index pointing at further indexes is parsed recursively three levels down, so a split structure needs no flattening before it goes in.
- The ceiling stops being a constraint. One job accepts as many as 1,000 sitemaps, far beyond anything a local service site will construct.
- Jobs queue rather than collide. Two process at a time with as many as 20 held behind them, so a whole portfolio can be pushed in one sitting and reviewed later.
Three levels exceeds what most sites will ever need and fits this shape precisely: a root index, a child file per category, the URL lists sitting below those. Structured that way, "is anything crawling the York County pages?" stops being a guess and becomes something you read. Intake accepts either an uploaded file or a plain address, so if the CMS already publishes an index there is nothing to export — hand the job that existing address and the parser descends on its own.
Pushing URLs directly, and reading what comes back
The other instrument takes addresses rather than files. A single batch accepts as many as 10,000 URLs, while each account works through 1,000 URLs a day. People mix the two numbers up constantly. One describes how much you may hand over at once; the other describes how quickly the pile is consumed.
Delivery happens through the IndexNow API, which tells GoogleBot and BingBot an address is new or has changed. The result is not a single verdict on the job: every URL keeps its own protocol showing which bot called, when it called, the status it received and the specifics of any failure. Three counters move in real time — submitted, found, failed — so you watch the batch drain rather than open a summary afterward.
The log is worth reading as a shape rather than as rows. Four recur, and each sends you somewhere different.
| Shape in the log | Most likely cause | The fix that actually applies |
|---|---|---|
| Accepted, no visit recorded | On file, but not appealing enough to request | Add links to it; raise it in the sitemap |
| A visit, then nothing after it | Requested, examined, not retained | Rewrite the document; it reads as duplicate |
| Failures clustered under one path | Something structural, not a single bad page | Check the template, the server, the redirects |
| A visit, then rows in the reports | All four steps completed | Leave it; spend the time on query data |
AutoSEO — keeping discovery running by itself
Suited to a domain that keeps emitting new URLs while nobody verifies they arrive.
- Submission stops depending on somebody remembering. Indexing is carried along with the campaign, not performed by hand every time something goes live.
- Discovery also arrives from outside. Link placement spans a partner network exceeding 230,000 sites, and each inbound link is itself a path by which an address gets found.
- Plan in weeks. Anything measurable usually starts between week four and week eight; that is the realistic timeframe to plan against.
URL lists can be dropped into Stream by the batch, convenient when the site keeps a separate list per county, and the reports and open tasks that follow appear in the same chronological feed. Longer walkthroughs live on our blog; the campaign work itself is set out under what we do.
Run the sums before a single page gets generated
Take a flooring contractor covering both states. Nine services against twenty-four towns comes to 216 pages. Six years of new construction have left fifty-eight community pages behind, and the product catalog, gallery and articles contribute roughly three hundred more: about 574 documents that all need requesting before any of them can be seen once. Bolt four sortable orders onto the catalog listings and the reachable count clears 1,700.
At 1,000 URLs a day even that site is worked through inside two days, and one batch would swallow it several times over. So the limits were never what held anything back. What holds it back is the answer you get by reading the town list aloud to whoever schedules the crews and asking which of them would take a job there next week. Eleven names survive. That converts 216 pages into 99, compresses fifty-eight community pages into six area pages plus a dozen project write-ups worth keeping, and reduces the sort orders to nothing, because they should never have been linked in the first place.
Ninety-nine defensible pages, submitted and crawled properly, outperform 1,700 that thin each other out. This is not an argument about writing standards. It follows from a fixed ration and from the expectation a crawler builds about what this particular domain tends to hold.
Questions that come up during the first batch
We submitted 400 town pages and nothing moved. Why?
Submission only handles the first step, and the blockage was almost certainly at a later one. Pages that differ by a single proper noun get requested and then dropped; sending them again alters the speed of the rejection and nothing else. Shorten the list to towns you truly serve and write each of those separately.
What should happen to pages for communities that sold out?
Sort them by whether anything specific happened there. A community page carrying real project detail becomes a portfolio entry and stays. A page that only ever advertised availability should be folded into an area page or removed outright, since it now answers a question no one asks.
Should retired pages be redirected or deleted?
Send a redirect only when a truly equivalent destination exists — for a closed phase that is normally the area page that absorbed it. With no equivalent, answer with an unambiguous removal instead of pushing everyone to the homepage. Dozens of unrelated redirects converging on one address are read for what they are.
Is it worth resubmitting the same URLs after an edit?
Yes when something genuinely changed: a rewrite, a merge, a brand new URL. No as a recurring routine over pages that have sat untouched. Each day's quota is fixed, and burning it on unchanged documents is the same as withholding it from this week's work.
Does everything belong in the sitemap?
No, and listing everything is the usual error. The file states which URLs you consider worth someone's time. Whatever you could not justify in a meeting — filtered variants, closed phases, near-identical pages either side of the border — stays out, and keeping it out is a deliberate act rather than an omission.
To learn which step is holding your own URLs, connect the domain, submit the tree in whatever state it is in today, and study the per-URL protocol before editing so much as a headline. Start a sitemap job in the Indexing Hub. The first discovery is rarely about rankings. It is usually a third of the list that nothing has ever requested — generated by a plugin during a growth spurt, still absorbing the attention that belonged to the towns where the crews actually work.