Insights / Measurement and paid acquisition

ICE and RICE Scoring: How to Prioritise Growth Experiments, With a Worked Example

How ICE and RICE scoring work, when to use each, a worked example with a real backlog, and how to run the winning test so that the result means something.

ICE and RICE scoring — Gufy insights.

ICE and RICE are two simple ways to rank ideas so that you test the most promising one first. ICE scores Impact, Confidence and Ease. RICE scores Reach, Impact, Confidence and Effort. Use ICE to choose this week's growth experiment and RICE when ideas reach very different numbers of people.

This guide explains both, works through an example, and then covers the part most guides leave out: how to run the winning test so the result tells you something.

Why score at all

Every startup has more ideas than time. Without a method, the test that gets run is the one the loudest person likes, or the most recent idea. Scoring does not make the decision for you. It makes you say out loud why you believe an idea will work, which is where most weak ideas fall over.

ICE scoring

Popularised by Sean Ellis for growth teams. Score each idea from 1 to 10 on three questions.

Factor Question A 10 looks like
Impact If it works, how much does it move the goal? Could double the number you are trying to move
Confidence How sure are you that it will work? You have data or a previous test that points this way
Ease How little time and money does it take? One person, one day, no developer

ICE score = (Impact + Confidence + Ease) ÷ 3. Some teams multiply the three instead, which punishes a low score on any one factor. Either works if you are consistent.

RICE scoring

Developed at Intercom for product decisions.

Factor How to score it
Reach Number of people affected in a set period, such as visitors per month
Impact 3 = massive, 2 = high, 1 = medium, 0.5 = low, 0.25 = minimal
Confidence 100% = high, 80% = medium, 50% = low
Effort Person-months of work

RICE score = (Reach × Impact × Confidence) ÷ Effort.

Which to use

ICE RICE
Time to score Minutes Longer; needs real numbers
Best for Weekly growth experiments Product roadmaps and larger pieces of work
Strength Fast, easy to explain Accounts for how many people are affected
Weakness Subjective Slower; false precision if the inputs are guesses

A third method, PIE (Potential, Importance, Ease), from Chris Goward, is built for choosing which pages to test and works the same way.

For an early-stage startup choosing what to test next, ICE is usually enough. Add Reach when one idea touches every visitor and another touches a small corner of the product.

A worked example

A B2B startup is running ads to a landing page with a booking calendar. Clicks are coming in. Bookings are not. The goal is booked demos. The team lists five ideas.

Idea Impact Confidence Ease ICE
Make the landing page headline repeat the promise in the ad 8 8 9 8.3
Test a short page (headline, one line, calendar) against the long page 7 6 8 7.0
Move the qualifying question to after the booking step 7 7 6 6.7
Use the ad platform's built-in lead form in place of the landing page 6 5 9 6.7
Rebuild the page with customer logos, case studies and product images 6 4 3 4.3

How the scores were reached matters more than the totals.

  • Headline match scores high on confidence because there is evidence. The ad promises one specific saving. The page opens with a different idea. A visitor who clicked for one promise and lands on another has to work out whether they are in the right place.
  • The qualifying question scores well on confidence for the same reason. In this example, session recordings showed several visitors filling in their name and email, reaching a question about what kind of business they were, and leaving. That is evidence, not opinion.
  • The full rebuild scores low on ease and confidence. It might work. It is also the most expensive idea, and a previous test suggested the opposite: a page with more detail produced fewer bookings than a stripped-back one.

That last point is a disagreement we see on almost every project. One person is sure a plain page looks unfinished and needs logos, testimonials and images. Another has seen tests where less information booked more meetings. Neither can know which is right for this audience. The honest answer to "is this the best landing page?" is "I do not know yet, and neither do you". So both versions go into the queue as hypotheses, in order of score, and the cheap one runs first.

Running the test so the result means something

Scoring picks the test. These rules make it worth running.

  1. Write the hypothesis first. "If the headline repeats the ad's promise, more visitors will book, because they will know they are in the right place." If you cannot write the "because", you will not learn anything when it wins or loses.
  2. Change one thing at a time. If you change the headline, the layout and the form together and bookings rise, you will not know why. Mix two changes and the hypothesis is muddled.
  3. Decide the length and the budget before you start. With a small budget, run fewer versions. Spreading a modest daily spend across seven pages gives each one too few visitors to judge.
  4. Agree what counts as a win. Bookings, not clicks. Qualified bookings, if you can tell.
  5. Look at behaviour, not just totals. Session recordings and form analytics show where people stop. That is often where the next idea comes from.
  6. Keep a log. A screenshot of each version of the ad and the page, the dates, the spend and the result, in one place. In three months you will not remember what you tested.
  7. Judge it against the right goal. A test that produces one well-qualified meeting can be a success for a product where each customer is worth a great deal.

For how much a real test costs, see Marketing Budget for a Tech Startup Without Revenue.

What to do with the result

  • It won. Keep it, record why you think it worked, and move to the next idea.
  • It lost. That is normal. Most tests do not produce a winner. Record it and lower your confidence in similar ideas.
  • It was unclear. Either run it longer or accept that the difference is too small to matter and move on.

Then re-score the backlog. A result changes your confidence in the ideas that are left.

Common mistakes

  • Scoring everything 7. Force a spread. If every idea is a 7, the scores are not doing any work.
  • Confidence based on enthusiasm. Confidence should come from evidence: data, recordings, a previous test, customer conversations.
  • Scoring alone. Have two or three people score separately, then compare. The disagreements are the useful part.
  • Treating the score as the decision. It is a way to rank ideas and expose reasoning.
  • Never testing the low-ease, high-impact idea. Quick wins are attractive, but some of the biggest gains take real work. Schedule one of those alongside the easy ones.
  • Testing without a goal. Tie every experiment to the number the company is steering by. See North Star Metric.

Common questions

What is ICE scoring?

ICE scoring ranks ideas by three factors, each scored from 1 to 10: Impact (how much it would move your goal), Confidence (how sure you are) and Ease (how little time and money it needs). Average or multiply the three and run the highest-scoring ideas first. It was popularised by Sean Ellis for growth experiments.

What is the RICE scoring formula?

RICE score = Reach × Impact × Confidence ÷ Effort. Reach is the number of people affected in a set period, Impact is scored on a fixed scale from 0.25 to 3, Confidence is a percentage, and Effort is measured in person-months. It was developed at Intercom for product roadmaps.

What is the difference between ICE and RICE?

RICE adds Reach and measures Effort in real time units, so it is more precise and takes longer. ICE is quicker and more subjective. Use ICE for a small team choosing between growth experiments each week. Use RICE when ideas affect very different numbers of people or when you need to justify a roadmap.

How many experiments should a startup run at once?

As many as your traffic and budget can give a clear answer for, which for most early startups is one or two. Splitting a small budget across many tests means none of them gets enough data to read.

How long should a growth experiment run?

Decide before it starts, based on how many visitors or leads you need to tell the versions apart. A few days is enough to spot something badly broken. A reliable comparison usually needs at least one to two full weeks, so that weekdays and weekends are both covered.

Where Gufy fits

Gufy runs this process for SaaS and mobile-app startups: a scored backlog of experiments across ads, landing pages and conversion, onboarding and lifecycle messaging, each with a written hypothesis and an honest reading of the result. A growth workshop is a quick way to build the first backlog with your team.

If you have plenty of ideas and no way of choosing between them, bring the list.

Book a 30-minute growth call

Book a 30-minute growth call

Loading the calendar…