GTM experiment tracker
Set the sample, outcome thresholds, and cost guardrail before a GTM test. Record actuals later and export a decision brief.
Count only eligible exposures that could produce the primary outcome. Set a stop rate below the scale rate, and a minimum sample before interpreting results. The recommendation applies your own rules; it is not a causal or statistical test.
Inputs stay in this browser tab and are not sent to the site. Use fictional or approved information, never confidential employer or customer material.
Three fictional results, three different calls
Each test uses 200 minimum eligible exposures, a 2% stop rate, a 5% scale rate, and an $80 cost-per-outcome limit. These are teaching values, not recommended benchmarks.
| Test | Exposures / outcomes | Spend | Rate / cost | Rule check |
|---|---|---|---|---|
| Focused workflow guide | 240 / 14 | $560 | 5.83% / $40 | Scale candidate |
| Broad message | 240 / 4 | $240 | 1.67% / $60 | Stop candidate: weak qualified rate |
| Expensive channel | 240 / 14 | $1,400 | 5.83% / $100 | Stop or redesign: cost limit breached |
For the broad-message failure, inspect whether eligible people saw a clear, relevant offer before changing the threshold. For the cost failure, investigate acquisition cost and outcome quality; meeting the rate rule does not excuse a broken cost guardrail.
Keep the result interpretable
- Write the exposure unit. A delivered message, a unique eligible account, and a landing-page session are different denominators. Choose one before the test.
- Count the qualified outcome once. Deduplicate repeat actions and check the agreed qualification definition. Record missing and rejected outcomes separately.
- Include a comparable cost scope. Decide whether spend includes media, production, and delivery effort. Use the same scope for the threshold and actual result.
- Record contamination. If the offer, audience, or measurement changes, mark it in the review. Split the result or restart; do not quietly merge incompatible tests.
- Close the learning loop. Save the result, the alternative explanations, the next action, and the new review trigger. A failed test can still answer a useful question.
Download blank results worksheet ↓ Open the full experiment playbook →
What the tracker does not prove
This tool applies your decision rules. It does not calculate statistical significance, establish incrementality, adjust for channel selection bias, or recommend universal conversion thresholds. A minimum sample is a planning rule; it is not a statistical power calculation. Compare alternatives with a suitable design when a causal claim matters.