Orphex Experiment Planner

An experiment brief with hypothesis, randomization, metric/guardrail definitions, feasibility, fixed horizon, and a launch boundary.

Conversion & experimentsMITv2.0.0

When to use

Before launching a test whose business effect, traffic, guardrails, and decision rule need definition.

Bring the right data

Use your own export, or start with the template. The fictional example shows the expected shape.

View the input columns
experiment_id Required
Proposed bounded experiment identifier
baseline_trials Required
Eligible independent baseline units
baseline_successes Required
Unique binary outcomes in those baseline units
weekly_eligible_trials Required
Observed feasible traffic before splitting
target_rate Required
User-supplied meaningful treatment rate, fraction 0–1
alpha Required
Supplied two-sided false-positive criterion
power Required
Supplied desired detection probability under planning assumptions
outcome_delay_days Required
Required follow-up/reporting delay after the last enrolled unit
Add your business context

Supply your objectives, conversion definitions, currency, constraints, and approved brand facts once, then reuse the profile across reviews. Leave unknown values explicit.

Open profile source
View the versioned resources

See an example

Fictional data · an illustrative review, not a customer result.

Try asking

Plan a four-week test and assess whether its requested effect is detectable.

View the example result

Fictional example output

The requested four-week horizon is insufficient for the stated planning assumptions. Baseline submit rate is 100/1,000 = 10.00%. The meaningful target is 12.00%: +2.00 percentage points, or +20.00% relative.

The optional binary normal-approximation calculator requires approximately 3,841 independent units per arm, 7,682 total, for two-sided alpha .05 and power .80. At 1,000 eligible users/week with a 50/50 split, enrollment requires about 7.682 weeks, plus 7 days for the final cohort to mature. Four weeks supplies only about 2,000 users per arm.

Design: change only form fields; assign users consistently, analyze accepted submissions per eligible assigned user at the fixed horizon, check instrumentation before enrollment, and predeclare the qualification guardrail. The approximate sample estimate is not a guarantee and is invalid for clustered users or repeated dependent trials.

Decide whether to extend the enrollment horizon, seek additional eligible traffic, or revise the user-selected effect criterion. Do not launch a four-week test with a promised winner. No experiment created and no auto-apply enabled.

Before you start

Review the required context
  • Supplied data with the documented task-specific columns, stable scope, and refresh/maturity context
  • Business definitions and constraints relevant to the decision; see the reusable business-context reference

The method

Define the estimand and design

State the one change, hypothesis, eligible population, unit of randomization, control/treatment, primary outcome numerator and denominator, attribution/follow-up, and the smallest effect worth acting on. Separate a relative lift from percentage points. Specify consistent assignment, contamination risk, excluded units, instrumentation checks, guardrails, and the analysis horizon before launch. An optimized creative delivery comparison is not randomized assignment; platform experiment types and supported campaign/goal combinations must be verified for the actual account.

View SKILL.md

Orphex Experiment Planner

Turn a specific business uncertainty into a feasible experiment brief. Preserve the user's chosen scope and meaningful effect; do not add unrelated variants or guarantee a winner.

Define the estimand and design

State the one change, hypothesis, eligible population, unit of randomization, control/treatment, primary outcome numerator and denominator, attribution/follow-up, and the smallest effect worth acting on. Separate a relative lift from percentage points. Specify consistent assignment, contamination risk, excluded units, instrumentation checks, guardrails, and the analysis horizon before launch. An optimized creative delivery comparison is not randomized assignment; platform experiment types and supported campaign/goal combinations must be verified for the actual account.

Keep outcome follow-up, conversion maturation, reporting/import delay, and attribution model/lookback as separately named definitions. A supplied seven-day outcome or reporting lag does not establish a seven-day attribution window. Preserve the supplied label and leave an unsupplied attribution rule unknown; do not fill it with a platform default when writing the brief.

For a simple independent binary outcome with two equal arms, use the optional calculator for an approximate fixed-horizon sample estimate from baseline rate, target rate, alpha, and power supplied by the user. State assumptions and round enrollment upward. Do not use that calculation for revenue/CPA means, clustered geographies, repeated users, sequential monitoring, multiple variants, or heavy-tailed outcomes; use an appropriate documented method or report that feasibility remains unquantified. Baseline zero/one or missing traffic requires different evidence, not fabricated defaults.

For a baseline-sensitivity scenario, explicitly state whether the absolute effect, relative effect, or target rate is held fixed and recompute the sample requirement under that convention. A lower baseline does not universally increase required sample: both outcome variance and the distance to the specified target matter. Do not make a directional sample or power claim from a changed baseline alone, or silently switch effect conventions. Alternative meaningful-effect criteria require the user's approval before they replace the supplied decision question.

Compare required sample with actual eligible traffic after allocation and add outcome/reporting maturation. A business deadline is not evidence of sufficient power. If infeasible, present the tradeoff: more traffic/time, a different user-approved effect criterion, or a descriptive pilot. Do not weaken the criterion silently. Predeclare what happens for inconclusive, guardrail-failing, or invalid results, including stop conditions for operational faults and a rollback reference.

Distinguish an enrollment deadline from a fully mature result deadline. For a result deadline, subtract required outcome follow-up and reporting lag from the available calendar horizon before calculating traffic needed for enrollment. Derive a traffic requirement only when that remaining enrollment window is positive. If the result horizon is no longer than the required lag, increasing traffic cannot create a positive enrollment window for a new fully mature trial. Extend the result deadline, clarify that the deadline concerns enrollment only, or offer a clearly descriptive pilot; sample sufficiency alone does not establish deadline feasibility.

Deliver a reviewable brief

Include hypothesis, assigned units, metric, guardrails, sample assumptions, dates/lag, exclusions, and the decision rule. Keep creation, edits, and rollout proposal-only without specific authorization. Check platform auto-apply settings rather than assuming a scheduled test cannot mutate campaigns.

Official reference

Portable inputs and examples

State whether the result is complete, partial, or blocked for the requested decision. Link material findings to actual supplied rows/sources and separate observed metrics, hypotheses, and estimates. Lead with a short business conclusion, then evidence, uncertainty, and the next measurable check. A data export or installed skill does not authorize account changes.

For the supported arithmetic only, optionally run the bundled calculator with Python 3: python3 scripts/marketing_math.py experiment < calculation.json. Read its input mapping in the input contract before preparing JSON. It reads JSON, not CSV directly. If Python or the requested method is unavailable, show a reproducible alternative calculation or mark it unsupported; do not report an uncomputed result as verified.

Source:View on GitHub Open SKILL.md