14  Scenarios: what if…?

Everything up to this point has been about building a model as a tool. This lesson is about using it.

Learning Objectives

By the end of this lesson, you will be able to:

  • make use of a valid simulation model for decision support,
  • define the context: who has to decide, what can be controlled, and what is the context,
  • formulate a scenario set that is comparable, contrasting, plausible, and small enough to interpret,
  • analyse and interpret simulation output against a baseline to derive an insight,
  • state what a scenario result allows you to claim, and express the uncertainty of your claim.

I refer to Shell (2008, p. 8) for a simple and elegant definition of a scenario:

A scenario is a story that describes a possible future.

A valid model as such is not an answer to anything. The valid model is the basis, but it produces insight only when we use it to for “what if…” questions about the modelled system (Figure 14.1).

Figure 14.1: The workflow of working with models. Model development builds the tool; scenario analysis uses it.

14.1 Asking the right question

Take these five questions:

  • Would additional doors have saved lives in that fire?
  • Do more dogs get the flock into the pen faster?
  • What happens to motorway traffic as autonomous cars take over?
  • How does street infrastructure impact bicycle safety?
  • Do a few heavy storms erode more soil than a year of drizzle?

Two of these are decisions that somebody has to make. Three are research questions. In terms of the workflow, it makes no difference: in all cases the model is asked to compare alternatives, and in all cases the alternatives have to be worked out before anything is run. This scenario-analysis workflow is the subject of this lesson.

In the first lesson we listed the situations in which simulation modelling can answer questions that cannot be answered otherwise. Each of the five questions is one of them (Table 14.1):

Table 14.1: The five questions of this lesson, and the reason each of them needs a model.
Question Why we cannot simply try it out
More doors Too dangerous and too unethical to experiment with
More dogs Observable, but not repeatable often enough to learn from
Autonomous cars The future does not yet exist to be measured
Open streets Expensive to build and expensive to undo
Heavy storms Slower than any experiment we could run

Working out a good scenario question is harder than it sounds, and it is where most scenario studies go wrong.

For a good example of a well-posed question consider the fire-evacuation model (Figure 14.2). On 1 January 2026, 41 people died in the bar Le Constellation in Crans-Montana. Investigators have reported that the main door opened inwards, that an emergency exit in the basement was kept locked, and that a renovation had narrowed the staircase on which most of the victims were found. Criminal proceedings are ongoing, so these are findings under investigation and not an established cause.

Figure 14.2: Three scenarios from Rino Lovreglio’s evacuation reconstruction of the Crans-Montana fire (Lovreglio, 2026) in a bar, as shown in a Tagesschau report (Schweizer Radio und Fernsehen, 2026): the same room with one, two and three exits. Drag the slider to move all three to the same moment after the alarm.

14.1.1 Levers and context

The question behind these scenarios is a question about the lever: would more accessible exits have let people out in time?

Levers are things that can be controlled and changed by a decision maker. In scenarios we want to test what would happen, if we changed a lever. There is also a long list of things that the decision maker can’t control: how many people were inside at the time the fire braked out, how fast the fire spread, how people behaved once they understood what was happening, the weather outside. This is the context.

This gives us the distinction that scenario design rests on (Table 14.2).

Table 14.2: Levers and context. Confusing the two is the most common flaw in a scenario study.
Levers Things the decision maker can change. Levers become the dimensions of the scenario set.
Context Things they are exposed to but cannot change. Context becomes the uncertainty the scenarios have to span.

Levers and context are not properties of the world, but of the decision. Crowd size is context for a building authority and a lever for an event organiser. The same model can serve both, with different scenarios.

Sometimes the thing we care about most is not a lever at all. Nobody decides how quickly autonomous cars appear on the motorway; it happens to us. A model of mixed traffic can still be enormously useful, but the scenario set has to be built the other way round: the share of autonomous vehicles spans the uncertainty, and the levers, if any, are dedicated lanes, headway regulations or a mandated communication standard. Studies of this kind are called exploratory: they ask what the world might do to us. Studies built on levers are called normative or policy scenarios: they ask what we might do about it. Most real studies cross the two, running each policy against several futures.

Exercise: levers or context?

Back to the fire. Seven things that matter for how many people get out alive - sort them from the point of view of the authority that writes the fire safety regulation. One of the seven belongs in neither box: work out which, and why.

A question phrased without this distinction tends to be unanswerable. “What will erosion look like in 2050?” asks a model to predict, which it cannot do, and names no decision. “Which of these two grazing regimes holds up better in a year of heavy storms?” names a lever, names a context, and can be answered.

Exercise: Who decides? What changes? How to compare?

A scenario question has four parts: someone who has to decide, something that this person can change (lever), something they cannot (context), and something they would measure (indicator) for decision support. Assemble one for the herding model. Not every combination is a question that a model can answer - find out which are, and why the others are not.

14.2 Designing a scenario set

You now have a question. A scenario set turns it into a small number of stories that can be compared. Work along the following principles to design your scenario set.

14.2.1 Start from a baseline

Every scenario set needs a reference: the world as it is, the current practice, business as usual. Without it, “more dogs” has nothing to be more than. The baseline is the most frequently forgotten scenario and the one that carries the most interpretation.

Note

Define the baseline scenario first, and describe it in one sentence of plain language.

14.2.2 Make the contrast large enough

The lesson on parameterisation established that stochastic models return a range of outcomes, not a single value. A scenario difference that is smaller than that range is invisible, no matter how many runs you do.

14.2.3 Hold everything else equal

This is the principle that makes a scenario set an experiment rather than a collection of runs. In the erosion example (Figure 14.3), the two rainfall scenarios differ in the distribution of rain over the year, and are carefully designed to deliver the same total amount of water. Because of that, any difference in the results can be attributed to the timing. Had the heavy-rain scenario also delivered more water, the comparison would have been worthless.

Figure 14.3: Two scenarios that differ in one respect only: a few heavy rain events (right) against constant drizzle (left), summing to the same amount of rainfall.

14.2.4 Keep them plausible

A scenario that nobody believes produces a result that nobody uses. Plausibility is not a statistical property; it is established by talking to people who know the system, and it is closely related to the face validity tests of the validation lesson.

14.2.5 Keep them few

Two or three scenarios should tell the story. Each additional dimension multiplies the number of simulation runs, exactly as it did for parameters. More importantly, each additional dimension halves your audience’s ability to hold the comparison in their head, and a comparison nobody can follow is not a result.

14.2.6 Make sure everything is set before you start

Which indicators you will record, how many runs per scenario, and how large a difference would have to be before you would call it meaningful. Decide on all three before hit the simulation button.

Depending on the model, the runs may take a while, and there is time enough for a coffee, and some larger models you even need to leave running for a couple of days. It can get frustrating, if you realise in the end that you forgot to adjust this little detail in the code, and you have to redo the whole thing. It’s worth taking your time to double-check everything before you start with the scenario simulations.

14.3 Reading the answer

Now you are ready to run the scenarios. But before you do so, commit to an expectation.

The value of a simulation experiment lies almost entirely in the distance between what you expected and what you got, and that distance cannot be measured after the fact. Once we know an outcome, we unconsciously adjust our memory of what we predicted so that it agrees with what happened. Fischhoff (1975) called this hindsight bias. Only a prediction that you explicitly formulated and wrote down before the run is evidence.

14.3.1 The answer is encoded in patterns

We model space, so the answer usually is a spatial or spatio-temporal pattern rather than a number. Compare the paths the sheep take under each scenario, or where the flock tends to split, or which corner of the field the dogs never manage to clear. Figure 14.4 show an automated feature extraction of the flock’s convex hull over time, which could track the herd’s centroid trajectory, differences in hull complexity across the herding phases, and how the flock’s distance to the gate evolved over time. Visualised differences of this kind are frequently more convincing to a practitioner than any statistics.

Figure 14.4: Spatial herding patterns that allow pattern comparison between scenarios and the observed baseline.

14.3.2 Look at distributions, not at means

The value of an indicator from a single run is one draw from a distribution, and the mean of many runs is only one property of that distribution. Plot the full distributions of your indicator, one curve per scenario, and compare them to fully understand the results for this indicator.

What you are looking for is overlap. Two curves whose means differ but which sit almost on top of each other describe scenarios you could not tell apart in practice. Two curves that barely touch describe a difference that would be visible to a shepherd on a hillside.

The sandbox in Figure 14.5 has nothing to do with sheep and its numbers are invented. Drag the two sliders and watch what happens to a difference when the world is noisy.

Figure 14.5: When is a difference between two scenarios large enough to see? The left slider moves the scenarios apart, the right one adds run-to-run variability. The numbers are invented.

14.3.3 Effect size before significance

This one may sound surprising: statistical significance is not a strong argument! With enough runs, almost any difference becomes statistically significant, because you control the sample size yourself. This makes significance a much weaker criterion in simulation than in field research, and it makes the size of the effect the thing that matters.

The erosion study makes the point sharply. The mean erosion depth was 2.3 mm under drizzle and 2.4 mm under heavy rain, and the difference was highly significant. It was also 0.1 mm. The difference in eroded area, by contrast, was both significant and large enough to matter on a hillside. Same experiment, same p-values, entirely different practical weight.

Your herding results will most likely show a trade-off: the scenario that pens the flock fastest is not the one that treats the sheep most gently. Which matters more is not a question the model can answer. It is the shepherd’s question, and the modeller’s job is to lay the trade-off out clearly rather than to resolve it.

To compare the distributions of two scenarios formally, the family of t-tests is the standard tool. A t-test asks whether the means of two samples differ more than would be expected by chance. In R:

t.test(scenario_a$time_to_pen, scenario_b$time_to_pen)

The null hypothesis is that there is no difference between the means. A p-value below 0.05 is conventionally taken as evidence against it. R uses the Welch variant by default, which does not assume that the two samples have equal variance, and this is usually what you want for simulation output.

Report the p-value if you must, but report the two means, the spread, and the size of the difference always. And remember that the number of runs is your choice, which means the p-value is partly your choice too.

14.3.4 Robustness beats precision

A precise answer that holds only for one parameter setting is worth less than an imprecise one that holds for all plausible settings. Run your scenario comparison again with the model parameterised slightly differently, using the sensitivity analysis technique from the parameterisation lesson. If the ranking of the scenarios survives, you have a result worth reporting. If it flips, you have learned something more important: the answer depends on a quantity nobody has measured, and that is what your study should say.

14.3.5 Insights arise from the unexpected

Now take out the prediction you wrote down.

If the model confirmed it, you have gained a little confidence and not much knowledge. If it contradicted it, you have gained the thing the model was built for.

The erosion study is a worked example of exactly that. The hypothesis was that a few heavy rain events produce less erosion, but deeper. The simulations showed more erosion, and only marginally deeper. The hypothesis had to be rejected and reformulated (Figure 14.6).

Figure 14.6: The revised hypothesis: heavy rain results in more, and marginally deeper, erosion.

In the herding model, my guess is that you will find something similar. Two dogs will beat one comfortably. Four dogs may well be worse than two, because they get in each other’s way and the flock splits. If that happens, resist the temptation to call it a bug. Non-linear behaviour of this kind is what emerges from interaction, and it is precisely the class of result that a bottom-up model exists to produce.

14.4 What you may claim

A scenario result is a statement about the model, not about the world. What permits us to make a claim about the real world is the validation work of the previous lesson, but it has its limits.

The first limit is the domain of applicability. The traffic model was validated against ten minutes of cars counted from a webcam, in traffic that was one hundred percent human-driven. The scenario asks about traffic that is 80% autonomous. Nothing in the validation data supports that, and no amount of further counting could, because the situation does not exist yet. This is not a flaw peculiar to that study. It is the normal condition of scenario analysis: the interesting question is outside the range in which the model could be tested. What allows us to extrapolate is not the data but the credibility of the represented processes, which is why so much of this module has been about getting the mechanisms right rather than the fit.

The second limit is equifinality. Your scenario produced a pattern, and other model structures could have produced the same pattern for other reasons. “This scenario produced that outcome” is not the same claim as “this cause produces that effect”.

The third is that uncertainty must be inseparably communicated together with the result. A number quoted without its range is received and further distributed as a crisp fact. In the first lesson we met the fisheries of the Adriatic, where uncertainty in the input parameters was set aside on the way into policy, and as a consequence the ecosystems were overfished.

This matters most where the results are used. The Crans-Montana reconstruction went out on the evening news while criminal proceedings against the operators of the bar and against municipal officials were under way. A model like this one is read by a court, by a regulator and by bereaved families. The honest form of its finding is not “two more doors would have saved those 41 people”. It is something narrower and more useful: under the assumptions of this model, the additional exits brought the evacuation time below the time that was available. The gap between those two sentences makes for the ethical use of a powerful tool.

Which brings us to communicating the insights gained from scenario analysis to non-modellers and decision makers. Show two or three scenarios, never ten. Show ranges, not points. Support the verbal description with one graphical abstract of the scenarios themselves, as in Figure 14.3, because it is intuitive and it travels.

14.5 Three things to remember:

  • Design the scenario set before you run it, including what would count as a difference worth acting on.
  • Compare scenarios against a baseline, holding everything else equal, and keep the set small enough to interpret.
  • The model answers about the model world. Say what it cannot tell you as clearly as what it can.

Step by step, and always beyond the reach of the experiment we would rather have run, scenario analysis contributes to a better understanding of the processes that shape our complex world, forwards as well as backwards in time (Figure 14.7).

Figure 14.7: Scenario analysis works into the future as well as into the past.

In the very first lesson we asked why anybody would simulate. This is the answer: because some questions can only be answered by a model, and because asking them well is a skill of its own.

References