Orphex Experiment Result Reviewer
A design-aware result with effect, uncertainty, guardrails, validity checks, and an evidence-bounded rollout recommendation.
When to use
After a supplied experiment has run and mature arm outcomes are available.
Bring the right data
Use your own export, or start with the template. The fictional example shows the expected shape.
View the input columns
experiment_idRequired- Defined experiment
armRequired- control or treatment
assigned_unitsRequired- Independent eligible units included under the predeclared rule
successesRequired- Unique binary primary outcomes after mature follow-up
guardrail_successesOptional- Unique guardrail outcomes where the stated eligible denominator applies
Add your business context
Supply your objectives, conversion definitions, currency, constraints, and approved brand facts once, then reuse the profile across reviews. Leave unknown values explicit.
See an example
Fictional data · an illustrative review, not a customer result.
Review the test and decide whether treatment can roll out.
View the example result
Fictional example output
No clear primary winner; rollout is not justified by this evidence. Control is 100/1,000 = 10.00%; treatment is 110/1,000 = 11.00%. Difference is +1.00 percentage point (+10.00% relative).
Under the supplied independent binary fixed-horizon assumptions, the unpooled normal approximate 95% interval for the absolute difference is −1.6867pp to +3.6867pp. It includes zero and material harm as well as benefit. This is inconclusive, not proof of equality or a 95% probability that treatment wins.
Qualification per assigned user is 4.00% control and 3.90% treatment (−0.10pp). No predeclared guardrail margin was supplied, so acceptance remains unverified. The maturity and assignment checks are supported by the fictional design; any different actual design would require reassessment.
Retain the control and decide whether a separately planned follow-up can answer the question; do not repeatedly extend this test until significance appears. No experiment applied or traffic shifted.
Before you start
Review the required context
- Supplied data with the documented task-specific columns, stable scope, and refresh/maturity context
- Business definitions and constraints relevant to the decision; see the reusable business-context reference
The method
Verify what was tested
Read the hypothesis, primary outcome/denominator, assignment unit and method, allocation, dates, maturation, exclusion rules, effect criterion, confidence/method, and guardrail thresholds. Check assignment integrity and unexpected sample imbalance without treating its cause as known; identify repeated exposures, contamination, concurrent edits, early stops, or unplanned exclusions. Analyze assigned eligible units under the supplied intention-to-treat rule when appropriate. Platform-optimized ad delivery and simple before/after reports do not establish randomization.
View SKILL.md
Orphex Experiment Result Reviewer
Evaluate the supplied experiment against its predeclared design and business decision. A higher observed metric is not automatically a statistically supported or economically acceptable winner.
Verify what was tested
Read the hypothesis, primary outcome/denominator, assignment unit and method, allocation, dates, maturation, exclusion rules, effect criterion, confidence/method, and guardrail thresholds. Check assignment integrity and unexpected sample imbalance without treating its cause as known; identify repeated exposures, contamination, concurrent edits, early stops, or unplanned exclusions. Analyze assigned eligible units under the supplied intention-to-treat rule when appropriate. Platform-optimized ad delivery and simple before/after reports do not establish randomization.
Reconcile arm counts and numerator boundaries, then show absolute rates/counts, absolute effect in percentage points, and relative effect only with a nonzero baseline. Prefer the supplied valid platform/statistical report with its method and scope. The optional helper offers only an independent-binary fixed-horizon normal approximate difference interval. It is unsuitable for small/extreme counts, clustered assignment, revenue/CPA ratios, fractional attribution credits, multiple comparisons, or sequential peeking; do not silently reuse it. Compute with a valid method or explicitly leave inferential status unavailable.
A confidence interval including zero is inconclusive, not evidence of equivalence. Equivalence/noninferiority needs a predeclared meaningful margin and appropriate analysis. Statistical support alone does not satisfy profitability or guardrails. Check mature secondary/guardrail outcomes and do not certify a guardrail without its declared rule. Adjust or disclose multiplicity and the consequences of post-hoc segment picking.
Decide within scope
Classify the result as supported benefit, supported harm, inconclusive, or invalid under the documented method; describe any unknowns. Recommend rollout only when design, primary effect, maturity, and business guardrails support it. Any follow-up needs a planned decision rule, not indefinite extension until a winner appears. Preserve the current control and keep applying experiments or shifting traffic within explicit authorization.
Official reference
Portable inputs and examples
- Read the input contract when mapping a new export or checking the example's scope and definitions. Copy the header-only CSV template when preparing data; equivalent supplied exports remain acceptable.
- Read the reusable business context only for business facts or constraints this task needs. Reuse user-supplied facts with their source/date; the template contains no default targets.
- Inspect the complete fictional input with its example output when learning the output and calculation boundaries. Never use fictional values for a real account.
State whether the result is complete, partial, or blocked for the requested decision. Link material findings to actual supplied rows/sources and separate observed metrics, hypotheses, and estimates. Lead with a short business conclusion, then evidence, uncertainty, and the next measurable check. A data export or installed skill does not authorize account changes.
For the supported arithmetic only, optionally run the bundled calculator with Python 3: python3 scripts/marketing_math.py experiment < calculation.json. Read its input mapping in the input contract before preparing JSON. It reads JSON, not CSV directly. If Python or the requested method is unavailable, show a reproducible alternative calculation or mark it unsupported; do not report an uncomputed result as verified.
Source:View on GitHub Open SKILL.md