Hypothesis testing is a methodical comparison of two or more versions of a commercial solution (price, product card, promotions, delivery terms) to confirm or refute a specific assumption about its impact on a key business KPI.
How hypothesis testing works
Hypothesis testing is built around four elements: a clear hypothesis, a target metric, control and test groups, and statistical verification of the result. The sequence is simple: you formulate a change, run parallel variants, collect data and check whether the change exceeded a pre-specified significance threshold.
- Hypothesis: a precise statement in the format “If we implement X, then metric Y will change by Z%.”
- Metric: one primary metric (CTR, conversion to order, average order value) and a set of secondary guard metrics (cancellations, returns, margin).
- Control and test: two groups of traffic/items — without the change and with it.
- Statistics: significance testing (commonly α=0.05) and test power (commonly 0.8) to separate random fluctuations from a real effect.
Without a correct plan and calculations, a test becomes an “observation” rather than a decision-making tool.
Why a Kaspi.kz seller should run tests
Hypothesis testing answers concrete questions that determine a seller’s sales and margin. Examples of why this matters for sellers in Kazakhstan:
- Increasing a product card conversion by just 0.5 percentage points at a baseline conversion of 2–3% can yield dozens of additional orders per month for a product with 10–30 thousand impressions.
- Price optimization: lowering price by 5–10% can boost turnover but reduce margin; a test reveals the point where revenue and profit are maximized.
- Changing delivery terms or adding a “kassa/samovyvoz” option (cash desk/pickup) can reduce cancellations and returns — especially important for bulky goods in Kazakhstan’s regions.
On Kaspi sellers face traffic limits and seasonality: cards with 200–1,500 sessions per day require different testing approaches compared to top categories that get 10–50k impressions per week.
Practical testing approaches on Kaspi.kz
Kaspi platform constraints mean you usually can’t A/B everything at the user level across the whole marketplace, but there are several practical formats:
- SKU-level tests: Create two otherwise identical SKUs (or use variants of an existing SKU) and expose one to the change (price, title, images) while keeping the other as control. This works well when buyers choose among similar items.
- Time-based tests: Apply a change for a defined period and compare results to a previous comparable period, adjusting for seasonality and day-of-week effects.
- Geographical split: If your logistics and assortment allow, run the change in specific regions where you have sufficient traffic and compare to other regions.
- Search/listing experiments: Test elements of the product card — title, main image, price ending, promo flags — on subsets of products and measure CTR and conversion lift.
- Promo mechanics: Compare different promotion parameters (discount size, cashback, Buy Box priority) on cohorts of comparable SKUs or time windows.
Choose the format that minimizes cross-contamination and suits available traffic. For low-traffic SKUs aggregate similar products into a test group or focus on high-impact changes (price or visibility) to reach significance faster.
Sample size and timing: how much traffic and time are needed
Key inputs for a sample size calculation are baseline conversion (p0), the minimal detectable effect (Δ) you care about, significance level (α, usually 0.05) and power (1−β, usually 0.8). For proportions (conversion) use standard formulas or online calculators to get required observations per group.
- If baseline conversion is low (1–2%), detecting small relative lifts requires very large samples.
- For CTR or visit-based metrics, compute required impressions; for order-based metrics — required sessions or users.
- Factor in expected data loss: fraud, cancellations, returns. If you expect 10% of orders to be invalid for the metric, inflate the sample size accordingly.
Timing: never stop a test before achieving the planned sample. Minimum duration is one full business cycle (usually 1–2 weeks) to cover weekday/weekend effects. Low-traffic categories may need several weeks or months. Do not pick end dates post‑hoc based on favorable spikes.
Metrics and guardrails
Define one primary metric that directly reflects your hypothesis and a set of guardrails (secondary metrics) to ensure no hidden harm to the business.
- Primary metric examples: CTR for visibility experiments, conversion to order for sales experiments, average order value for pricing/bundle experiments.
- Guardrails: cancellations, returns, margin, refund rate, customer satisfaction indicators. For promotions add spend, number of beneficiaries and ROI metrics.
- Set thresholds for guardrails in advance (for example: cancellations must not increase by more than X%) and stop the test if they are breached.
Track uplift and absolute values: a relative improvement on a tiny baseline may be irrelevant in revenue terms.
Steps to launch a test — checklist for the seller
- Formulate a clear hypothesis: state the change, target metric and expected effect size.
- Choose primary and guard metrics and predefine thresholds for acceptance/rejection.
- Decide the test format (SKU split, time-based, region, etc.) and randomization level.
- Calculate required sample size and estimate test duration given current traffic.
- Prepare implementation: update product cards, prices, images or promo settings. Document exactly what is changed.
- Launch and monitor daily: check for data integrity, traffic leaks, and guardrail breaches.
- At the planned end, run statistical analysis, check segment results and validate with business metrics (revenue, margin).
- Document results and decide next steps: roll-out, iterate or abandon.
Common mistakes and how to avoid them
- Stopping early: leads to false positives. Predefine sample and duration.
- Multiple uncorrected comparisons: running many tests or many metrics increases false discovery; use corrections or prioritize a single primary metric.
- Cross-contamination: poor randomization (users seeing both variants) biases results — randomize at user or SKU level and fix assignment for the duration.
- Ignoring guardrails: short-term gains may hurt margin or increase returns — always monitor secondary metrics.
- Not documenting: without a clear record of what was changed, you can’t learn or replicate results.
Concrete hypothesis ideas to test on Kaspi.kz
- Price endings: “.99” vs round prices — does psychological pricing lift conversion?
- Image variation: lifestyle image vs plain product shot — impact on CTR and conversion.
- Title optimization: including brand/model vs feature-first title — effect on search CTR.
- Promo mechanics: cashback level or % discount — impact on order volume and margin.
- Delivery options: adding kassa/samovyvoz (cash/pickup) or free regional delivery — effect on cancellations for bulky goods.
- Bundle offers: cross-sell bundles vs single item discounts — change in average check and units per order.
How to use data after the test
Translate statistical results into business decisions:
- If the result is statistically significant and economically meaningful — plan a roll‑out with monitoring and operational checks (inventory, logistics, support load).
- If the result is not significant but the direction is positive — consider increasing sample size, targeting specific segments, or testing a larger effect.
- Segment analysis: sometimes effects are concentrated in regions, user cohorts, or product groups — use this to design targeted roll-outs.
- Document learnings in a hypothesis registry: what worked, what didn’t, sample sizes and dates to avoid repeating tests on the same conditions.
Tools and automation
Use available tools to plan, run and analyze tests:
- Analytics platforms for extracting traffic, sessions and conversion metrics.
- Sample size calculators and basic statistical packages (R, Python, online calculators) for power analysis.
- Internal Kaspi seller tools and A/B frameworks (if available) or automation scripts to apply changes and revert them reliably.
- Monitoring dashboards to track guardrails in near real-time and alert on breaches.
When possible, automate repetitive parts: variant assignment, metric collection and basic significance checks. This reduces human error and speeds up iteration.
In summary, hypothesis testing on Kaspi.kz is a practical, data-driven way to improve key seller metrics — but it requires careful planning, correct metrics, sufficient sample size, and attention to guardrails to turn experiments into reliable business gains.
Часто задаваемые вопросы
- How do I calculate the minimum sample size for a conversion test on Kaspi.kz?
- You need to know the current conversion (p0), the minimal detectable effect (Δ), the significance level (usually α=0.05) and the test power (usually 0.8). Use these parameters to calculate the number of observations (or conversions) per group with standard formulas for proportions or online calculators. Don’t forget to account for expected data loss (fraud, cancellations) and split traffic so you reach the sample within an acceptable timeframe.
- How long should I run an experiment for reliable results?
- Duration depends on required sample size and average daily traffic, but the minimum is one full business cycle (usually 1–2 weeks) to account for weekday and weekend behavior. Low-traffic categories may need several weeks or months; stopping earlier risks seasonal effects and random spikes. Always check metric stability before and after and avoid choosing end dates post hoc.
- Which primary and guard metrics should I set for a product card test?
- Choose the primary metric according to the goal — card CTR if the goal is traffic, conversion to order or average check if the goal is sales. Guard metrics include cancellations, order refunds, returns, margin and satisfaction indicators (if available) to avoid a “win” that harms the business. For promotions, add spend and beneficiary share metrics.
- How to prevent traffic crossover between control and test groups?
- Randomize at the user or unique session level and fix the group assignment for that user/device; for products you can randomize at the SKU level. Block public promo codes and external channels that could skew allocation, and control for repeat visits and cross‑purchases. Monitor overlap and stop the test to investigate if you detect significant leakage.
- What if the test result is not statistically significant but the effect direction is positive?
- First, assess the test power and ensure the sample was sufficient for the expected effect; you may need to increase sample size or extend the test. Check segments — the effect might be concentrated in a specific audience, which suggests a targeted experiment. If the effect is small and doesn’t cover economic costs, document the result and move to new hypotheses.