Projects
Audit
The Overview, Indexing, and Crawling views: grouped findings, Google status per page and why pages are left out, live crawl progress and crawl reports, and the path from Audit to Tasks.
Audit groups observed site problems and opportunities so you can decide what deserves a task. Existing findings remain readable on free projects; running a new audit or asking the engine to work on an issue requires a managed project.
Audit has three views, shown as tabs under the page title and under Audit in the sidebar while it is open:
- Overview lists the finding groups: what is wrong or worth doing, and whether it is already in a task.
- Indexing shows which pages Google has, by host and page group, why the rest are left out, and the Google checks behind it.
- Crawling shows a running crawl live and keeps a report for every finished crawl.
Running an audit
Managed projects receive scheduled audits, and you can choose Run audit at any time. While an audit runs, the status next to the page title shows a progress ring, the current step, the share of known URLs processed, and the fetching time left at the recent pace. Hover it for the processed and known counts, pages per minute over the last minute, and elapsed time; select it to open Crawling. The estimate uses the crawler's own pace over roughly the last minute, not the whole run, and leaves out the analysis after fetching. Discovery can raise the known-URL count while a crawl runs, so the estimate may change.
An audit runs in five steps: Robots & sitemaps, Pages, Links & issues, Findings, and Google check. While it runs, Overview shows one crawl card with the progress bar, the page being fetched, and the current step; the card disappears when the audit finishes.
Cancel stops an active audit. If a job is interrupted, it resumes from its saved crawl frontier instead of starting every page again. While an interrupted job is waiting for its automatic retry, Try again now makes that retry eligible immediately without starting a second crawl. After an audit fails or is cancelled, Resume crawl requeues the same audit and its saved frontier instead of discarding fetched pages. If no crawl checkpoint exists yet, the button says Retry audit.
Crawling
While an audit crawls, Crawling shows its steps, pages per minute with the site's typical answer time, the latest fetched addresses, each host's current pace, progress by click depth, and how large page families are being sampled. When no crawl is running, it shows the report for the newest finished crawl; open any crawl in Crawl history to see its report. A report covers sitemap coverage, pace, responses, where the time went, click depth, answer times, hosts, and sampled page families. Crawls from before these details were recorded show only what they recorded.
A crawl separates sitemap URLs known to the crawler, URLs selected for the audit, and URLs processed. Large ID-shaped route families can be sampled in two stages (50, then 200 pages) only when the fetched pages have consistent SEO structure; if the sample differs, the crawler expands to the remaining URLs within the audit allowance. Sampled siblings are counted as unverified, never as checked or issue-free. Other sitemap URLs are selected in full while they fit within the allowance; if the inventory exceeds it, selection spreads across the inventory and reserves up to one fifth of the allowance for pages found only through links. Link targets disallowed by robots.txt are not added to the fetch frontier. Unselected or unfetched URLs have not been checked for HTTP status, noindex, canonical, or Google indexing. An administrator sets the global per-audit URL allowance, which defaults to 100,000; it is a safety ceiling, not a promise that every known URL will be fetched.
A crawl is complete only when every known URL was checked. It is partial when it stopped at the click-depth limit, missed sitemap URLs, had failed or blocked URLs, found no usable sitemap, or hit a link-evidence limit. Addresses that lead outside the crawl scope, such as a redirect to a host you excluded, are skipped rather than failed and do not make a crawl partial. A partial or sampled run can confirm observed issues, but it cannot clear an older finding just because a page was not revisited.
The crawler starts with robots.txt, conventional sitemaps, and the sitemaps submitted to the project's connected Search Console property. It follows sitemap indexes, including child files without an .xml suffix. Public subdomains of the project domain may be crawled when discovered in a sitemap, link, or redirect; Project Settings can exclude crawl hosts and paths. This does not expand permissions for site-changing actions. The crawler respects its own robots rules on each host. Each host starts at 4 requests at a time, 100 ms apart, and speeds up after sustained healthy responses. It slows down on HTTP 429, 503, and retryable failures, honors a server-supplied Retry-After delay, and ramps back up after healthy responses. A slow successful page alone does not trigger backoff.
SERPclimber reads the initial HTML returned by the server. It does not run JavaScript during normal audits. An apparent empty app shell is marked as limited evidence, not proof that Google cannot render the page. Markdown and plain-text resources are fetched and checked for HTTP status, redirects, and X-Robots-Tag directives such as noindex. They are not evaluated for HTML titles, meta descriptions, headings, or links. Other non-HTML resources can be discovered but are not evaluated as HTML pages.
Indexing
Indexing gives every known page one Google status:
- In Google: a Google check says the page is indexed (confirmed), or Search Console shows impressions for it in the last 90 days (seen in search).
- Not in Google: a Google check says Google left the page out, with Google's reason.
- Kept out by the site: the page redirects, carries noindex, answers with an error such as 404, is blocked by robots.txt, or names another page as its canonical.
- Not checked yet: the page can be indexed but gets no search impressions and has no Google check. Missing impressions never mean "not in Google"; only a Google check can say that.
- Couldn't check: the crawl skipped the page (outside the crawl scope) or could not fetch it.
Pages are grouped by host and first path segment, such as /blog/*. Hosts that appear in Search Console but have never been crawled are shown separately, with a link to the crawl settings. Why pages aren't in Google explains each reason in plain words next to Google's own wording, with example pages.
Checked in Google lists Search Console URL Inspection results: Google's stored view of each page, not a live test or a complete index census. You can enter a specific in-scope URL to check it, or use Check 20 pages now to check pages that have not been checked yet. After each crawl, up to 20 pages are checked automatically, starting with indexable pages that get no search impressions. Results are reused for seven days; automatic checks are limited to 200 per day and on-demand checks to 50 per day per connected property.
Technical indexability uses what our crawler observed: access and HTTP status, redirects, noindex directives, and whether useful initial HTML was available. It is a technical signal, not proof that Google indexed the page. Search performance is supporting evidence: a page appearing in Search results during a date range was visible then, but a page with no performance row is not necessarily unindexed.
Findings and tasks
Audit groups repeated observations into issue families. Each family shows its severity, observed affected-page count, examples, explanation, and whether it is already in a task. For large groups, the listed URLs are representative examples, not an exhaustive page list. Issues on pages outside the fetched sample are not assumed to exist.
Known Search Console clicks and impressions can raise a crawl issue's task priority when an affected page has meaningful exposure. Pages absent from the bounded Search Console summary are treated as unknown, not low value. A task's immediate sparse recheck is provisional; only fresh, complete audit evidence can clear the grouped finding.
If an unusually large crawl would produce more finding groups than the engine can reconcile safely, Audit falls back to broader observed issue groups. These preserve the raw per-page evidence but are not automatically filed as one fixable task.
The list starts with Not handled yet, followed by In a task. Fixed and hidden groups can be shown when needed. Open a family to review evidence, ask AI about it, file a task with Fix it, hide it temporarily, ignore it, or mark it fixed. Tasks hold the actual work; resolving a task does not erase the audit evidence. The next audit can confirm or reopen a finding.
Checks include fetch and HTTP failures; redirects, canonicals, and noindex; titles, descriptions, HTML language, headings, thin content, image alt attributes, malformed JSON-LD, internal links to broken pages or redirects, HTTPS pages linking to HTTP, duplicate content, and crawl-depth or orphan risks. Title and description length checks are advisory because Google snippets are not determined by a fixed character count.
