SEO Experiments: Every Task Is Measured Against Its Own Baseline
Most SEO changes are never measured. In SERPclimber there is no separate experiment to set up. Every task created from evidence records the number it set out to move and what that number read before the work started. When the task is resolved, the same number is read again over a window of the same length, and the result is shown next to the task.
replaces
- Guesswork
- Untracked title edits
- Post-hoc chart reading
- Separate SEO testing tools
- Frozenbaseline at task creation# The metric, the page or query, and the starting number are fixed before anything is touched.
- Same lengthmeasurement window# The after window matches the baseline window, so the comparison is fair.
- 6 stepson every task# Found, Planned, Approved, Executed, Verified, Measured.
What SEO Experiments does
Why most SEO is never proven
The usual loop is: change a title, wait, look at a chart, and argue about seasonality. There is no control, so any explanation fits. A page that improves gets credited to the change. A page that drops gets blamed on a Google update. Neither conclusion is earned.
True split testing, where half your visitors see version A and half see version B, is not available in organic search. Google crawls one version of a URL. So the control has to come from time rather than from audience.
The task's own baseline is the control
SERPclimber does not have a separate experiments product. Instead, measurement is a property of every task.
When the audit finds a problem and it becomes a task, the task records a measurement contract: which metric it is trying to move, which page or search that metric is read on, and what the number was over a fixed window before the work started. That baseline is never recalculated later.
When the task is resolved, the engine reads the same metric on the same page or search over a window of the same length, ending on the last day the evidence covers. The outcome card shows the new number against the starting number with a small trend line. While there is not yet enough data, the card says how many days have passed instead of claiming a result.
How it works
- 1
A finding becomes a task
From Audit: a click-through gap, a striking-distance query, content decay, cannibalization, an authority gap, or an AI citation gap.
- 2
The baseline is frozen
The metric, the page or search, and the starting number over a fixed window are recorded the moment the task is created.
- 3
A plan is prepared
The engine writes a plan or a task brief with the recommended edits: the current value, the proposed value, and the reason.
- 4
You approve, or you do the work
A plan that changes the site waits for your approval unless your autonomy policy allows it. Work handed to you has a Mark done button.
- 5
The engine verifies
After the work is done, the engine checks the live site itself rather than trusting a claim. Crawl fixes are verified this way instead of measured.
- 6
The outcome is measured
The same metric is read over a window of the same length. The result is shown on the task, in the weekly or monthly report, and on the Performance chart as a task marker.
What each kind of task measures
Each finding family has an honest metric. The task carries it from the day it is filed, so nobody can pick a flattering number after the fact.
- Click-through gap — the click rate of the page
- Striking distance — the rank of the search
- Content decay — visits from search to the page
- Keyword cannibalization — visits to the page already winning the search
- Authority gap or lost links — sites linking to you
- AI citation gap — AI citations of your site
Honest about what it cannot say yet
A measurement is only shown when there is enough data to support it. Work with no honest baseline, such as a crawl or template problem, is verified against the live site instead of measured. The task rail shows which step the work has reached and what is holding it there.
- No result is shown until a full window has passed
- Crawl and template fixes are verified, not measured
- Every step on the rail is only marked reached when the evidence for it exists
- Reports say "not yet measurable" instead of claiming impact
Everything it covers
Baseline frozen at task creation
The metric, the entity and the starting number are fixed before any work starts.
Same-length measurement window
The after window matches the baseline window.
Six-step task rail
Found, Planned, Approved, Executed, Verified, Measured, each with its date.
Live verification
The engine checks the site itself after work is done, including work you did by hand.
Outcome card on every task
The new number against the starting number, with a trend line.
Task markers on the chart
Every resolved task is flagged on the Performance timeline.
Measured outcomes in reports
Weekly and monthly digests include what the measurement windows concluded.
Sort by most at stake
Tasks can be ordered by the size of their baseline, so the biggest numbers come first.
Findings reopen
If verification fails or the problem returns, the finding reopens in Audit.
Plan versions
Every version of a plan is kept, with what changed between them.
Approval before site changes
A plan that changes the site waits for your approval unless your policy allows it.
Task numbers
Every task has a number such as #42 you can use in chat, reports and commits.
Is this the same as SEO A/B testing?
It has the same goal, with the method that organic search allows. Classic A/B testing splits an audience. You cannot split Googlebot. Large publishers approximate it by splitting a set of similar pages into test and control groups, which needs hundreds of comparable templated pages to reach significance.
SERPclimber uses the other valid design: a before-and-after comparison on one page or search, with a fixed window and the task's own starting number as the control. It works on sites with dozens of pages rather than thousands. The trade-off is real: a single-page time-based comparison is more exposed to seasonality and Google updates than a large split test. That is why the baseline is frozen in advance, why Google updates are marked on the chart, and why "not yet measurable" is a permitted answer.
Frequently asked questions
It is a before-and-after comparison, which is the design available in organic search. You cannot serve two versions of a URL to Google, so the control is the task's own starting number over a fixed window rather than a parallel audience.
No. There is nothing to set up. Every task created from audit evidence records its baseline automatically, and measurement starts on its own when the task is resolved.
The after window is the same length as the baseline window. Until it has passed, the task says how many days of how many have gone by. Nothing is called a result before then.
The outcome card shows it plainly. The finding stays on record in Audit, and the engine or you can plan a different approach on the same task. Nothing is hidden or rewritten.
No. SERPclimber does not undo changes on its own. Every action plan states whether the change is reversible, conditional or irreversible before you approve it, and the previous value is recorded in the plan so you can restore it.
Crawl and template problems, such as a missing sitemap, a noindex tag or a canonical error, have no honest traffic baseline. The engine checks the live site to confirm the fix instead.
Yes. Each task has its own outcome card, the weekly or monthly report has a Measured outcomes section, and every resolved task is marked on the Performance chart.
Works with
~/serpclimber/performance/seoSEO Dashboard
Search Console shows what ranks. GA4 shows what earns. Neither one shows you what to fix next. SERPclimber puts both on one timeline, adds AI answers, AI crawler activity and your backlink profile, and keeps 16 months of Search Console history so you can see a trend instead of a snapshot.
Read more
~/serpclimber/settingsSEO Automation
Full autonomy is a big ask for something that edits your live site. SERPclimber runs one loop — audit, plan, approve, execute, verify, measure — and lets you decide how much of it runs without asking. Start with Copilot, where the engine does low-risk work on its own and asks before anything risky, external or paid.
Read more
~/serpclimber/tasksContent Generator
Generic AI writing tools start from a keyword and a tone slider. This one starts from the page you already have: the queries it wins, the queries it is losing, and the finding that made it a task. Every draft is a rewrite of a real page, aimed at a real gap, and it arrives in the task as a proposal you review before anything changes.
Read morePoint it at your site.
A project takes one domain. The first pass tells you what it found — before anything changes on your site.