Orphex Experiment Result Reviewer

A design-aware result with effect, uncertainty, guardrails, validity checks, and an evidence-bounded rollout recommendation.

Conversion & experimentsMITv2.0.0

When to use

After a supplied experiment has run and mature arm outcomes are available.

Bring the right data

Use your own export, or start with the template. The fictional example shows the expected shape.

View the input columns
experiment_id Required
Defined experiment
arm Required
control or treatment
assigned_units Required
Independent eligible units included under the predeclared rule
successes Required
Unique binary primary outcomes after mature follow-up
guardrail_successes Optional
Unique guardrail outcomes where the stated eligible denominator applies
Add your business context

Supply your objectives, conversion definitions, currency, constraints, and approved brand facts once, then reuse the profile across reviews. Leave unknown values explicit.

Open profile source
View the versioned resources

See an example

Fictional data · an illustrative review, not a customer result.

Try asking

Review the test and decide whether treatment can roll out.

View the example result

Fictional example output

No clear primary winner; rollout is not justified by this evidence. Control is 100/1,000 = 10.00%; treatment is 110/1,000 = 11.00%. Difference is +1.00 percentage point (+10.00% relative).

Under the supplied independent binary fixed-horizon assumptions, the unpooled normal approximate 95% interval for the absolute difference is −1.6867pp to +3.6867pp. It includes zero and material harm as well as benefit. This is inconclusive, not proof of equality or a 95% probability that treatment wins.

Qualification per assigned user is 4.00% control and 3.90% treatment (−0.10pp). No predeclared guardrail margin was supplied, so acceptance remains unverified. The maturity and assignment checks are supported by the fictional design; any different actual design would require reassessment.

Retain the control and decide whether a separately planned follow-up can answer the question; do not repeatedly extend this test until significance appears. No experiment applied or traffic shifted.

Before you start

Review the required context
  • Supplied data with the documented task-specific columns, stable scope, and refresh/maturity context
  • Business definitions and constraints relevant to the decision; see the reusable business-context reference

The method

Verify what was tested

Read the hypothesis, primary outcome/denominator, assignment unit and method, allocation, dates, maturation, exclusion rules, effect criterion, confidence/method, and guardrail thresholds. Check assignment integrity and unexpected sample imbalance without treating its cause as known; identify repeated exposures, contamination, concurrent edits, early stops, or unplanned exclusions. Analyze assigned eligible units under the supplied intention-to-treat rule when appropriate. Platform-optimized ad delivery and simple before/after reports do not establish randomization.

View SKILL.md

Orphex Experiment Result Reviewer

Evaluate the supplied experiment against its predeclared design and business decision. A higher observed metric is not automatically a statistically supported or economically acceptable winner.

Verify what was tested

Read the hypothesis, primary outcome/denominator, assignment unit and method, allocation, dates, maturation, exclusion rules, effect criterion, confidence/method, and guardrail thresholds. Check assignment integrity and unexpected sample imbalance without treating its cause as known; identify repeated exposures, contamination, concurrent edits, early stops, or unplanned exclusions. Analyze assigned eligible units under the supplied intention-to-treat rule when appropriate. Platform-optimized ad delivery and simple before/after reports do not establish randomization.

Reconcile arm counts and numerator boundaries, then show absolute rates/counts, absolute effect in percentage points, and relative effect only with a nonzero baseline. Prefer the supplied valid platform/statistical report with its method and scope. The optional helper offers only an independent-binary fixed-horizon normal approximate difference interval. It is unsuitable for small/extreme counts, clustered assignment, revenue/CPA ratios, fractional attribution credits, multiple comparisons, or sequential peeking; do not silently reuse it. Compute with a valid method or explicitly leave inferential status unavailable.

A confidence interval including zero is inconclusive, not evidence of equivalence. Equivalence/noninferiority needs a predeclared meaningful margin and appropriate analysis. Statistical support alone does not satisfy profitability or guardrails. Check mature secondary/guardrail outcomes and do not certify a guardrail without its declared rule. Adjust or disclose multiplicity and the consequences of post-hoc segment picking.

Decide within scope

Classify the result as supported benefit, supported harm, inconclusive, or invalid under the documented method; describe any unknowns. Recommend rollout only when design, primary effect, maturity, and business guardrails support it. Any follow-up needs a planned decision rule, not indefinite extension until a winner appears. Preserve the current control and keep applying experiments or shifting traffic within explicit authorization.

Official reference

Portable inputs and examples

State whether the result is complete, partial, or blocked for the requested decision. Link material findings to actual supplied rows/sources and separate observed metrics, hypotheses, and estimates. Lead with a short business conclusion, then evidence, uncertainty, and the next measurable check. A data export or installed skill does not authorize account changes.

For the supported arithmetic only, optionally run the bundled calculator with Python 3: python3 scripts/marketing_math.py experiment < calculation.json. Read its input mapping in the input contract before preparing JSON. It reads JSON, not CSV directly. If Python or the requested method is unavailable, show a reproducible alternative calculation or mark it unsupported; do not report an uncomputed result as verified.

Source:View on GitHub Open SKILL.md