How should Cafezinho, a cute Brazilian coffee app, price its subscription? We find the profit-maximising price — then learn why measuring it is the hard part.
Why I built this
In the microeconomics class of my INSEAD Executive MBA, we worked through the classics — demand curves, elasticity, the profit-maximising price. But every example was a physical thing: oil, aluminium, Apple Watches, physical units shipped. I’m a PM. I wanted to do two things the coursework didn’t: take those methods to a software product, and finally try A/B testing — a validation method I’d thought about for years and never actually run.
So I made up an app to do it on. “Cafezinho” is a fictional app, which is exactly the point: with an invented product I can invent the data too, stress-test different scenarios, and still keep it grounded in reality.
How do textbook pricing methods actually apply to a software product?
How would you design and use an A/B test to validate the price you chose?
Fictional app, illustrative numbers, real method. I’m working in the open — every figure here is one I’d defend or revise out loud. — Sally
Cafezinho sells a monthly subscription. Charge too little and they leave money in the tip jar; charge too much and customers walk. Somewhere between free and absurd sits a price that does the most work. Part 1 finds it. Part 2 asks the harder question: once you have a candidate price, how would you even test whether it's better — and can a business this size support that test yet?
The founder wants to jump the price from R$14.9 to R$24.9. Before I touch a slider or run a test, I work through three questions in order — price first, then whether a premium tier earns its build, then whether any of it can be proven with an experiment. Here's the path, and where each move lives on the page. Move 1 spans most of Part 1, move 2 is the Roaster question that closes it, and move 3 is all of Part 2.
What it is. A cute Brazilian coffee app (iOS & Android). The free Beans Log lets you track the beans you've bought, rate them, and remember what worked. The subscription unlocks the Recipe Suggester — a weekly "dial-in" (grind, dose, water, time) tuned to the exact beans and gear you own.
Costs. ≈ R$2.5 per subscriber each month (payment fees + servers + recipe content), plus ≈ R$ 4,000 / month fixed (hosting, tools, the founder's time). Fixed cost doesn't move the best price — it just decides whether Cafezinho is in the black.
The question. The founder thinks R$14.9 is too low and wants to jump to R$24.9. Is that right — and how would you ever prove it? That's what the rest of this page works out.
All figures illustrative; the pricing results below are computed from this setup.
We're setting the price of Cafezinho's paid tier — the Recipe Suggester — and optimising for one thing: monthly profit. Drag the price. Four business metrics move at once — and they don't agree on where the peak is.
Demand model is illustrative — linear demand, marginal cost R$2.5/sub.
Each metric peaks somewhere else on the same curve. Picking which one to chase is a leadership call — not something the maths decides for you.
Below the unit-elastic line (E = 1), a price rise grows revenue. Above it, customers flee faster than price climbs. The profit peak always lands in the elastic region — to the right of E = 1.
The single price is set — now the last move on pricing is to tune it per segment. Not every drinker reacts the same way. The rule of thumb: cut price where demand is springy, keep it where demand is sticky — always measured at the baseline you just set. (Notice the café-pros barely flinch — that stickiness is exactly what the Roaster build question turns on.)
Once you've set the single best price (the baseline), layer a tiered menu on top of it. Different drinkers will pay different amounts; the menu lets willingness-to-pay self-select. You design the ladder, customers sort themselves. The top rung — a premium Roaster tier — isn't built yet; whether it's worth building is exactly the question the next section takes on.
It looks like the café-pros' higher willingness-to-pay leaves money on the table on a higher tier, but should the founder build this right now? The short answer is: on the economics alone, it’s a long shot. The café-pro niche is small, so the extra contribution a premium tier earns is capped at about R$ 212/month — too thin to ever repay even a lean build. (A purely economic read: it assumes a lean, iterative build and leaves go-to-market and marketing out of the lens — either could shift the picture.)
So, back to Q1 — how do textbook pricing methods apply to a software product? Like this: with marginal cost near zero, the classic variable-cost test goes almost silent (every subscriber clears it by a mile), so the interesting decisions relocate — to the fixed and avoidable costs of what you choose to build, and to whether you can even measure that a price is better. The methods still hold; software just moves where the hard call lives.
This demand curve was estimated. In real life you have to measure it — usually with an A/B test. Exploring how I'd run one, I worked through a five-step sequence — and found a common trap sitting inside it.
I work these five checks in order. Each one is a gate: it either earns the right to move to the next, or it tells me to stop. You don't power a test you haven't framed, you rehearse the plumbing with an A/A before a real A/B rides on it, the same A/A data then shows you exactly what peeking would fabricate, and you don't read a result whose split you never checked. One honest note: Cafezinho stops at the Step 2 gate today — so read Steps 3–5 as the playbook you'd run once traffic clears it, walked through here on A/A data so you can see how each check behaves.
Skip this and "did it work?" has four answers, not one — you end up keeping whichever metric flatters the result after the fact.
Before touching any data, I name the single effect the test has to detect and pin it to one primary metric. The move from R$14.9 towards R$24.9 implies paywall conversion falls from about 8% to about 5% — a three-percentage-point drop. That 3-point drop is the bar every step below is measured against: it's the effect the power check must be able to see, and the size of win peeking would fabricate on data where the true difference is actually zero.
I check power first for a blunt reason: a test that can't see the effect is worthless whatever it reports. Run it underpowered and a "no difference" result hasn't cleared the price — the test was simply blind. Trust that and you either ship a worse price believing it's safe, or you burn the whole experiment — and every later check earns you nothing, because they can't rescue a test that was blind from the start.
Power is the chance a test actually spots a real effect when one exists. At 150 visitors/arm this test reads at only ~18% power — it would miss that 3-point drop roughly four times in five. To reach the 80% power we want — at a two-sided α of 5% (95% confidence) — we'd need ~1,059 visitors/arm — about 7× today's traffic. So for Cafezinho today: not testable.
Once Cafezinho reaches ~1,059/arm, this exact test becomes trustworthy: 80% power, and a tight confidence interval instead of the noisy one you'd get today. How it grows to that traffic — marketing, growth, funnel work — is a separate plan, out of scope here.
Skip the dress rehearsal and you can't tell a genuine A/B win from broken plumbing — you'd be trusting a result you never calibrated.
To pressure-test the setup honestly, we run an A/A test — two groups that get the exact same offer (same price, same everything). The true difference between them is zero by construction, so any "significant" result is a false alarm. Run it before any A/B and it's a dress rehearsal: with no real effect to chase, it tells you whether the instrumentation is honest. It catches broken plumbing and sample-ratio problems (the split-ratio gate below), and it calibrates your real-world false-positive rate. A clean A/A should call "no difference" about 95% of the time.
We start with an A/A, not an A/B: in an A/B, group B is a real change. Here both groups are identical — so we know the honest answer is "no difference", and can confirm the instrumentation is sound before a real A/B ever depends on it.
Watch card ① — the tester who stops the moment it "looks significant." On a test with no real effect, it can still light up. That's not a fluke; it's the trap the next step measures.
Peek at the numbers and stop the moment they look good, and you'll ship phantom wins — burning engineering time and momentum chasing "improvements" the full test would never have found.
That same A/A data exposes the trap. On a true null — no real effect anywhere — peeking means stopping the first time a result crosses "significant". Take one look and you carry the 5% error you signed up for; take many looks and each is a fresh roll at a false alarm, so they compound. Run it 200 times and that habit inflates the honest 5% rate to ~22%: more than one "win" in five is a mirage. The fix costs nothing — pre-register the sample size N and the primary metric up front, then read the result once, at planned N. Do that and the false-alarm rate stays right where it should: ~5%.
Skip this and a broken randomiser turns any "effect" you measure into an artefact of who landed where — not the price at all. A note on order: I teach it last because it's easiest to picture once the A/A has shown you a clean split — but on live data it's the first thing to check, before you read a single result.
Sample-ratio mismatch (SRM) is when a test you designed to split 50/50 arrives lopsided. It's the last gate before you read anything, because it protects every other result: if the wrong people landed in the wrong arm, the "effect" could just be broken randomisation or logging. A healthy ~50/50 split (left) means the plumbing held — read on. A skew like 58/42 on a 50/50 design (right) means something's broken: stop, fix the randomisation or logging, and re-run before you diagnose a single number.
Five checks, in order — each one either earns the right to run the test or tells us to stop. Here's the ledger: what I checked at each step, and what it told me for Cafezinho.
Not at today's size. The power check says the test is effectively blind until traffic grows ~7×, and even once it isn't, peeking would still manufacture false wins unless you pre-register and read once. The honest playbook: start from the modelled price (Part 1), roll it out to new sign-ups rather than secretly charging identical users two prices, pre-register what you'll measure, clear the SRM gate, and watch the guardrails — cancellations, refunds, lifetime value — over time. An experiment supports the decision; it rarely makes it for you.