SK the record / projects / the coffee experiment
> a pricing experiment, in two movements.

The Coffee
Experiment

How should Cafezinho, a cute Brazilian coffee app, price its subscription? We find the profit-maximising price — then learn why measuring it is the hard part.

PROJECT · PRICING & A/BMETHOD · MICROECONOMICS + STATSSTATUS · LIVEWHEN · 2026

Why I built this

My pricing class taught me how supply-and-demand curves set prices for physical commodities. I wanted to point those methods at software.

In the microeconomics class of my INSEAD Executive MBA, we worked through the classics — demand curves, elasticity, the profit-maximising price. But every example was a physical thing: oil, aluminium, Apple Watches, physical units shipped. I’m a PM. I wanted to do two things the coursework didn’t: take those methods to a software product, and finally try A/B testing — a validation method I’d thought about for years and never actually run.

So I made up an app to do it on. “Cafezinho” is a fictional app, which is exactly the point: with an invented product I can invent the data too, stress-test different scenarios, and still keep it grounded in reality.

Q1

How do textbook pricing methods actually apply to a software product?

Q2

How would you design and use an A/B test to validate the price you chose?

Fictional app, illustrative numbers, real method. I’m working in the open — every figure here is one I’d defend or revise out loud. — Sally

> The one idea

One price isn't an answer. It's a point on a curve.

Cafezinho sells a monthly subscription. Charge too little and they leave money in the tip jar; charge too much and customers walk. Somewhere between free and absurd sits a price that does the most work. Part 1 finds it. Part 2 asks the harder question: once you have a candidate price, how would you even test whether it's better — and can a business this size support that test yet?

Receipts
R$22
profit-max price
4 metrics
4 different “best” prices
~22%
false positives if you peek
The through-line

One founder's question, answered in three moves

The founder wants to jump the price from R$14.9 to R$24.9. Before I touch a slider or run a test, I work through three questions in order — price first, then whether a premium tier earns its build, then whether any of it can be proven with an experiment. Here's the path, and where each move lives on the page. Move 1 spans most of Part 1, move 2 is the Roaster question that closes it, and move 3 is all of Part 2.

1 Part 1
Find the price that does the most work
Goal
the paid tier's profit-maximising price — then tuned by segment
Method
fit the demand curve, then marginal revenue = marginal cost
2 Part 1
Weigh whether a premium tier is worth building
Key question
is a premium "Roaster" tier worth building?
The test
whether the extra contribution ever repays the one-time build
3 Part 2
Establish whether this case can even be A/B tested
Why
Part 1's demand curve was estimated; in real life you measure it — usually an A/B test
Key question
can Cafezinho's traffic even see the effect?
Meet Cafezinho — the setup behind the numbers

What it is. A cute Brazilian coffee app (iOS & Android). The free Beans Log lets you track the beans you've bought, rate them, and remember what worked. The subscription unlocks the Recipe Suggester — a weekly "dial-in" (grind, dose, water, time) tuned to the exact beans and gear you own.

Free · Beans Log
Track & rate the beans you buy. Fills the funnel; costs almost nothing to run.
R$14.9/mo · Recipe Suggester
Weekly dial-in recipes + everything in Free. This paid upgrade is what we're pricing.
~9,000
free users
~720
paying subscribers
R$14.9
current price / mo
~8%
paywall conversion
~6% / mo
subscribers cancel
one flat price
no tiers or discounts yet

Costs. ≈ R$2.5 per subscriber each month (payment fees + servers + recipe content), plus ≈ R$ 4,000 / month fixed (hosting, tools, the founder's time). Fixed cost doesn't move the best price — it just decides whether Cafezinho is in the black.

The question. The founder thinks R$14.9 is too low and wants to jump to R$24.9. Is that right — and how would you ever prove it? That's what the rest of this page works out.

All figures illustrative; the pricing results below are computed from this setup.

Part 1
Find the price
1Model the demand curve
2Find the one best price
3Pressure-test with elasticity
4Tune by segment
5Weigh the Roaster build
Step 1 · Model the demand curve

We're setting the price of Cafezinho's paid tier — the Recipe Suggester — and optimising for one thing: monthly profit. Drag the price. Four business metrics move at once — and they don't agree on where the peak is.

Interactive · price explorer
FX illustrative
Total profit (R$/mo)2.3k4.6k6.9kR$14R$28R$42Monthly priceR$22 optimum
Monthly price R$22
Conversion
5.9%
Revenue
R$ 11,623
Profit
R$ 6,302
Profit / sub
R$19.5
Price elasticity here
1.12
elastic
profit-max always sits in the elastic zone

Demand model is illustrative — linear demand, marginal cost R$2.5/sub.

Held constant here: churn. Every figure is per month. The real prize is lifetime value (LTV) — a higher price that nudges up cancellations can win this month yet lose over a customer's life as they get tired of paying. We've kept this to a single-month view on purpose. Including churn → LTV is the next layer of modelling, and beyond the scope of this case study.
Step 2 · Find the one best price

Four metrics. Four different "best" prices.

Each metric peaks somewhere else on the same curve. Picking which one to chase is a leadership call — not something the maths decides for you.

R$14R$28R$42Monthly price →ConversionR$0RevenueR$20.8ProfitR$22Profit / subR$42+
Maximise conversion and you let consumers keep their surplus. Maximise profit-per-sub and you price everyone out. The most interesting for our business case — revenue and total profit — live close together but not at the same place.
Step 3 · Pressure-test with elasticity

Elasticity: how hard does demand push back?

Below the unit-elastic line (E = 1), a price rise grows revenue. Above it, customers flee faster than price climbs. The profit peak always lands in the elastic region — to the right of E = 1.

Elasticity |E|1234+E = 1 · unit elastic · R$20.8InelasticElasticR$14R$28R$42
The brown dot is elasticity at your price — it lands on the red E = 1 line only at the revenue-max price (R$20.8). Now: R$22 → E = 1.12 (elastic). And because a positive marginal cost pushes the profit peak past the revenue peak, it sits to the right of E = 1, out in the elastic zone, at R$22 and not the founder's R$24.9.
Step 4 · Tune by segment

Discount the elastic. Hold the line on the inelastic.

The single price is set — now the last move on pricing is to tune it per segment. Not every drinker reacts the same way. The rule of thumb: cut price where demand is springy, keep it where demand is sticky — always measured at the baseline you just set. (Notice the café-pros barely flinch — that stickiness is exactly what the Roaster build question turns on.)

Students
Tight budgets, lots of options
stickyspringy
E ≈ 2.8
Discount hard
Hobbyists
Care about coffee, price-aware
stickyspringy
E ≈ 1.13
Hold near optimum
Café-pros
A work tool, not a treat
stickyspringy
E ≈ 0.6
Keep price firm
How do you give just students a lower price? You fence it — verify who qualifies (a student email or ID check, e.g. via a service like SheerID) so the discount reaches students only. Without a fence, a "student price" quietly becomes everyone's price and the segmentation collapses.
Step 5 · Weigh the Roaster build

Don't pick one price. Build a menu.

Once you've set the single best price (the baseline), layer a tiered menu on top of it. Different drinkers will pay different amounts; the menu lets willingness-to-pay self-select. You design the ladder, customers sort themselves. The top rung — a premium Roaster tier — isn't built yet; whether it's worth building is exactly the question the next section takes on.

Free
Free
Tasters & students. Costs nothing, fills the funnel.
Brew (Paid tier)
✓ recommended
R$22
The everyday drinker — baseline paid tier, set to the profit-max price.
Roaster (Premium paid tier)
not built yet
R$30.6
Café-pros who barely flinch at price — a premium tier that doesn’t exist yet.
Should Cafezinho build the Roaster tier?

It looks like the café-pros' higher willingness-to-pay leaves money on the table on a higher tier, but should the founder build this right now? The short answer is: on the economics alone, it’s a long shot. The café-pro niche is small, so the extra contribution a premium tier earns is capped at about R$ 212/month — too thin to ever repay even a lean build. (A purely economic read: it assumes a lean, iterative build and leaves go-to-market and marketing out of the lens — either could shift the picture.)

So, back to Q1 — how do textbook pricing methods apply to a software product? Like this: with marginal cost near zero, the classic variable-cost test goes almost silent (every subscriber clears it by a mile), so the interesting decisions relocate — to the fixed and avoidable costs of what you choose to build, and to whether you can even measure that a price is better. The methods still hold; software just moves where the hard call lives.

Part 2
Estimate that curve

This demand curve was estimated. In real life you have to measure it — usually with an A/B test. Exploring how I'd run one, I worked through a five-step sequence — and found a common trap sitting inside it.

1Frame the hypothesis & metric
2Power it — enough traffic?
3Validate with an A/A test
4Don't peek — pre-register
5Check the split (SRM)

I work these five checks in order. Each one is a gate: it either earns the right to move to the next, or it tells me to stop. You don't power a test you haven't framed, you rehearse the plumbing with an A/A before a real A/B rides on it, the same A/A data then shows you exactly what peeking would fabricate, and you don't read a result whose split you never checked. One honest note: Cafezinho stops at the Step 2 gate today — so read Steps 3–5 as the playbook you'd run once traffic clears it, walked through here on A/A data so you can see how each check behaves.

Step 1 · Frame the hypothesis & metric

What is the one effect this test must see?

Skip this and "did it work?" has four answers, not one — you end up keeping whichever metric flatters the result after the fact.

Before touching any data, I name the single effect the test has to detect and pin it to one primary metric. The move from R$14.9 towards R$24.9 implies paywall conversion falls from about 8% to about 5% — a three-percentage-point drop. That 3-point drop is the bar every step below is measured against: it's the effect the power check must be able to see, and the size of win peeking would fabricate on data where the true difference is actually zero.

Step 2 · Power it — enough traffic?

Can the test even see that 3-point drop?

I check power first for a blunt reason: a test that can't see the effect is worthless whatever it reports. Run it underpowered and a "no difference" result hasn't cleared the price — the test was simply blind. Trust that and you either ship a worse price believing it's safe, or you burn the whole experiment — and every later check earns you nothing, because they can't rescue a test that was blind from the start.

Power is the chance a test actually spots a real effect when one exists. At 150 visitors/arm this test reads at only ~18% power — it would miss that 3-point drop roughly four times in five. To reach the 80% power we want — at a two-sided α of 5% (95% confidence) — we'd need ~1,059 visitors/arm — about 7× today's traffic. So for Cafezinho today: not testable.

Traffic today / arm150
Traffic needed to run the test / arm — 80% power1,059

Once Cafezinho reaches ~1,059/arm, this exact test becomes trustworthy: 80% power, and a tight confidence interval instead of the noisy one you'd get today. How it grows to that traffic — marketing, growth, funnel work — is a separate plan, out of scope here.

Step 3 · Validate with an A/A test

Is the instrumentation honest before a real A/B rides on it?

Skip the dress rehearsal and you can't tell a genuine A/B win from broken plumbing — you'd be trusting a result you never calibrated.

To pressure-test the setup honestly, we run an A/A test — two groups that get the exact same offer (same price, same everything). The true difference between them is zero by construction, so any "significant" result is a false alarm. Run it before any A/B and it's a dress rehearsal: with no real effect to chase, it tells you whether the instrumentation is honest. It catches broken plumbing and sample-ratio problems (the split-ratio gate below), and it calibrates your real-world false-positive rate. A clean A/A should call "no difference" about 95% of the time.

We start with an A/A, not an A/B: in an A/B, group B is a real change. Here both groups are identical — so we know the honest answer is "no difference", and can confirm the instrumentation is sound before a real A/B ever depends on it.

Interactive · A/A test
…then watch both stopping rules below judge the same data.
Arm A · same offer
n = 0
Arm B · same offer
n = 0
95% CI for the difference (B − A)
no difference (0)
① Peek & stop
stops the instant it looks significant
② Run to planned N
judged once, at the end

Watch card ① — the tester who stops the moment it "looks significant." On a test with no real effect, it can still light up. That's not a fluke; it's the trap the next step measures.

Step 4 · Don't peek — pre-register

Does stopping early manufacture false wins?

Peek at the numbers and stop the moment they look good, and you'll ship phantom wins — burning engineering time and momentum chasing "improvements" the full test would never have found.

That same A/A data exposes the trap. On a true null — no real effect anywhere — peeking means stopping the first time a result crosses "significant". Take one look and you carry the 5% error you signed up for; take many looks and each is a fresh roll at a false alarm, so they compound. Run it 200 times and that habit inflates the honest 5% rate to ~22%: more than one "win" in five is a mirage. The fix costs nothing — pre-register the sample size N and the primary metric up front, then read the result once, at planned N. Do that and the false-alarm rate stays right where it should: ~5%.

Interactive · 200 A/A trials
Now run it 200 times — on data with no real effect
Peek & stop on "significant"
false positives — should be 5%
Pre-register, check once at N
false positives — right on target
No real difference exists — every "win" here is a false positive.
Step 5 · Check the split (SRM)

Did the 50/50 assignment actually arrive 50/50?

Skip this and a broken randomiser turns any "effect" you measure into an artefact of who landed where — not the price at all. A note on order: I teach it last because it's easiest to picture once the A/A has shown you a clean split — but on live data it's the first thing to check, before you read a single result.

Sample-ratio mismatch (SRM) is when a test you designed to split 50/50 arrives lopsided. It's the last gate before you read anything, because it protects every other result: if the wrong people landed in the wrong arm, the "effect" could just be broken randomisation or logging. A healthy ~50/50 split (left) means the plumbing held — read on. A skew like 58/42 on a 50/50 design (right) means something's broken: stop, fix the randomisation or logging, and re-run before you diagnose a single number.

Healthy✓ 50 / 50
Broken⚠ 58 / 42

What each step bought us

Five checks, in order — each one either earns the right to run the test or tells us to stop. Here's the ledger: what I checked at each step, and what it told me for Cafezinho.

1 Step 1 · Frame the hypothesis & metric
Checked
Named the one effect the test must see, and pinned it to a single primary metric before looking at any data — so "did it work?" has one honest answer, not four.
It told us
The move from R$14.9 towards R$24.9 implies paywall conversion falls ~8% → ~5%: that three-percentage-point drop is the bar everything else is measured against.
2 Step 2 · Power it — enough traffic?
Checked
Asked whether Cafezinho even has the traffic to detect that three-point drop at 80% power and 95% confidence — the gate that decides whether a test is worth running at all.
It told us
Not close — at 150/arm the test reads at ~18% power and needs ~1,059/arm (~7× today's traffic). Today the test is effectively blind, so a "no difference" result would prove nothing.
3 Step 3 · Validate with an A/A test
Checked
Ran a dress rehearsal on a true null — two identical arms, where "no difference" is the answer by construction — to prove the instrumentation before a real A/B rides on it.
It told us
The 5% false-alarm rate we allow is the flip side of correctly calling "no difference" ~95% of the time at planned N. A clean A/A means a later A/B's "significant" is signal we can trust, not noise.
4 Step 4 · Don't peek — pre-register
Checked
Compared stopping the moment a result looks significant against reading once at a pre-declared N — measured live across 200 identical-arm trials.
It told us
Peeking inflates the 5% false-alarm rate to ~22% (watch it happen in the sim above); read once at planned N and it stays ~5%. The fix is free — pre-register N and the metric, then read once.
5 Step 5 · Check the split (SRM)
Checked
Checked the 50/50 assignment actually arrived 50/50 — a test of the randomisation itself, not of the outcome.
It told us
A healthy ~50/50 means read on; a 58/42 skew means the randomisation or logging is broken and any "effect" is an artefact of who landed where. Clear this gate before you read any A/B result.

So can Cafezinho just A/B its way to the price?

Not at today's size. The power check says the test is effectively blind until traffic grows ~7×, and even once it isn't, peeking would still manufacture false wins unless you pre-register and read once. The honest playbook: start from the modelled price (Part 1), roll it out to new sign-ups rather than secretly charging identical users two prices, pre-register what you'll measure, clear the SRM gate, and watch the guardrails — cancellations, refunds, lifetime value — over time. An experiment supports the decision; it rarely makes it for you.

Recommendations for Cafezinho
  1. 1 Raise the price now. Move the paid tier off R$14.9 and towards the profit-max R$22 — the “Brew” baseline. The current price is simply leaving money on the table.
  2. 2 Roll it out to new sign-ups carefully. Don’t charge existing users two different prices. Pre-register what you’ll watch, then track the guardrails — churn, refunds, lifetime value — for a few months before you call it.
  3. 3 Hold Brew as the core tier; shelve Roaster. The build can’t pay for itself at today’s café-pro numbers — see the go/no-go above. It’s a someday-tier, not the main event.
  4. 4 Don’t trust an A/B test yet. At today’s traffic it’s underpowered — it can’t see the effect. Decide from the model now, and revisit experimentation once growth gets you to roughly 7× today’s traffic.
The Coffee Experiment · a Cafezinho case study · oisally.com All figures illustrative · purple & teal politely declined ☕