A small gap is usually noise
Every number in this guide is an example, made up for the arithmetic. Variant A shows a deal to 1,000 visitors and gets 30 orders, a conversion rate of 3.0 percent. Variant B shows a different badge to 1,000 visitors and gets 36 orders, or 3.6 percent. B looks clearly ahead, and most dashboards would happily colour it green.
Run the standard test for two proportions on those figures and the gap sits well inside what chance produces. If the two variants were identical, a gap of six orders or more would still turn up in nearly half of tests this size. Six extra orders per thousand visitors is not yet a finding.
To show a gap that size reliably, a common sample size calculation asks for roughly 14,000 visitors per variant. Most single products do not see that in a week, which is the first thing to accept about testing deals: small differences need a lot of traffic, and large differences are rarer than dashboards suggest.
How TreStack Bundles decides there is a winner
The app tests conversion rate, comparing the best variant with the second best using a two sided test for two proportions at 95 percent confidence. With three or four variants the bar is set higher, because every extra comparison is another chance of a lucky result.
It also refuses to judge early. Nothing is marked until every variant has at least 200 visitors and 10 conversions and the test has run for 72 hours. Those are floors, not targets. Passing them only means the test is allowed to look, and the example above would pass every floor and still show no winner, correctly.
When a variant does clear the bar, it gets a Significant winner badge. The app never stops a test by itself. Set winner sends all traffic to the variant you pick, and Stop A/B test lets you keep any variant, including one the statistics did not choose. The decision stays with you, which is why reading the result properly matters.
Decide the length before you start
The most common way to fool yourself is to check every morning and stop on the first good day. Results drift up and down as orders arrive, and a test watched daily will look decisive at some point purely by chance. Pick a length before you start, at least one full week so weekdays and a weekend are both in it, and read the result at the end.
Change one thing per test. A variant with a new badge, a new title and a deeper discount can win, and you will not know which change did it. Smart A/B test in TreStack Bundles is built around this: it proposes a single change, such as a badge or a timer, publishes it against your current deal, and splits traffic evenly.
Leave the split alone once the test is running. Changing it moves some visitors to another variant, which mixes the two groups you are trying to compare. Each browser keeps its variant, so a shopper who returns on the same phone sees the same deal, but one who switches to a laptop or clears their browser may see the other. That is a little noise, and one more reason not to over read a narrow gap.
More orders is not the same as more money
The winner badge judges conversion rate only. When the variants charge different prices, and in TreStack Bundles each variant charges its own price at checkout, the variant with more orders can be the one that earns less.
Take an example candle at 40 dollars that costs 16 to make. Variant A offers two for 15 percent off: 68 dollars, leaving 36 dollars of margin. Variant B offers two for 25 percent off: 60 dollars, leaving 28. With A at 30 orders and B at 36, and every order a single pair, A earns 1,080 dollars of margin and B earns 1,008. B wins on orders and loses on money.
So read three numbers side by side before choosing: conversion rate, revenue per visitor and, if you have entered cost per item in Shopify, profit per visitor. Visitors are counted only for shoppers who allow analytics, and test orders are left out, so treat every figure as an estimate for comparing variants rather than an accounting record. TreStack Bundles is coming to the Shopify App Store, and the docs below describe the test settings and how each metric is counted.