Archived post from legacy Decis reporting

This might feel a little heavy for a Sunday email but I wanted to share this note on DCDR’s accuracy ASAP. (You can skip to the end to see the results if you don’t want to wade through the maths.)

And if you are very impatient, the bottom line is that the baseline model outperforms random guesses without returning any wholly inaccurate results, which is a solid start.

Why We’re Tracking DCDR’s Accuracy

The accuracy of DCDR's stability assessments is very important, primarily to the user but also to the project's viability. After all, if the assessments aren't better than random guesswork, we aren't adding any value.

So, I want a way to assess and share the accuracy of the output.

However, as obvious as sharing the accuracy of your assessment might sound, it's not something we see a lot.

This makes sense when you think about it: forecasting is hard, and the chances of you getting a straight-A report are pretty low. So, the smart thing to do is not to publicize how accurate your results are. Or, if you do, you only share those instances where you got it right.

But that means the assessment provider doesn't have much skin in the game, making it too easy to brush over mistakes.

So, for better or worse, I've added an evaluation scores module to DCDR. This reviews an assessment at the end of the assessment period and compares the result to the forecast. That way, I can track the effectiveness of the models over time, and users can have confidence in the products they are using.

The Brier System for Scoring Forecasts

But how do you score these kinds of things? After all, there are three possible outcomes:

  1. The results matched the forecast: wholly accurate forecast

  2. The situation didn't change when it was supposed to: inaccurate forecast

  3. The situation went in the opposite direction: wholly inaccurate forecast (and misleading)

Next, we need a way to score these results where an opposite result -- which could have serious consequences -- is exaggerated and 'punishes' the overall score.

Luckily, Glenn W. Brier has already done the hard work here by developing the Brier Score to assess the effectiveness of things like weather predictions.

"...a strictly proper score function or strictly proper scoring rule that measures the accuracy of probabilistic predictions. For unidimensional predictions, it is strictly equivalent to the mean squared error as applied to predicted probabilities."

But to put it another way, the score takes the likelihood you ascribe to each outcome and then produces a value for the actual outcome relative to your confidence or likelihood level. But, because of the mean square function, the value of a confident right answer is exaggerated, as is the value of a confident wrong answer. Hedged -- e.g., 33/33/33 answers -- don't 'reward' correct answers as much as this is getting close to random chance (which we will come to in a moment as this is important).

Importantly

"the lower the Brier score is for a set of predictions, the better the predictions are calibrated" (Wikipedia)

So, an exact match is 0, and a completely wrong answer is 1 or 2, depending on how you use this system.

This is how the SuperForecasters at the Good Judgement project track their accuracy, so it's a well-recognized approach. (See point 4 here for more.)

The DCDR Evaluation

How does this apply to the three possible outcomes from DCDR:

  1. The results matched the forecast: wholly accurate forecast

  2. The situation didn't change when it was supposed to: inaccurate forecast

  3. The situation went in the opposite direction: wholly inaccurate forecast (and misleading)

I have pretty high confidence in the assessment pairs and expect these to reflect the situation in a location of the corresponding stability/instability. Therefore, assuming the assessment model matches these correctly*, I expect the forecast to be correct most of the time.

(* Pure LLM performance is a separate issue and will be assessed and scored separately.)

There's still the chance that the situation doesn't change and a much slimmer chance (in my estimation) that things go in the opposite direction. Based on this confidence, I ascribe the following likelihoods to each outcome, giving us the Brier scores.

  1. The results matched the forecast: 65% - 0.215

  2. The situation didn't change when it was supposed to: 30% - 0.915

  3. The situation went in the opposite direction: 5% - 1.415

You might think that my ascribing a 5% likelihood of heading in the opposite direction is trying to downplay this outcome. But remember, a low number is better in a Brier system, so 1.415 definitely 'punishes' the system.

I added the scoring module at the beginning of February and will start running monthly but will move up to weekly over time. That will help identify if there's a sweet spot for the forecast window, suggesting that I need to tighten the near-term window.

Before looking at the initial performance, we need to talk about lucky guesses.

What About Dart-Throwing Monkeys?

Sometimes, we hear about systems based on random guesses, dart-throwing monkeys, or stock-picking octopi that beat the experts. And there are cases where a lucky, uninformed guess will be right. Similarly, there will be cases where a highly informed, well-thought-out forecast is wrong.

So, how do we determine if DCDR or Mr. Bubbles is a better bet as an assistant to the CRO?

If we put random guesses to the test using a Brier score, a random selection that gives us each result 1/3 of the time gets a score of 0.667.

So 0.667 is the score to beat, right?

Not so fast.

In our application, these results would contain an equal number of wholly accurate, inaccurate, and wholly inaccurate forecasts.

However, in the case of country stability, where an opposite forecast could be catastrophic, a wholly accurate forecast doesn't cancel out a wholly inaccurate forecast.

This goes to Nassim Taleb's idea of 'extremistan' and fat tails -- where the consequences are highly exaggerated at the extreme ends of the scale -- the concept of ruin where the average doesn't matter: once you're ruined, you're ruined.

Therefore, because of this extreme downside, the dart-throwing monkey's 0.667 score is very different from a score in the same range that does not include so many wholly incorrect results.

So, I'm not saying you shouldn't employ a dart-throwing monkey to assist your Chief Risk Officer, but if you do, just remember that the quality of the output might not be what it first appears. (Plus, they cost a fortune in bananas.)

So, how did DCDR do? Did it outperform Mr. Bubbles?

DCDR's First Report Card

I tested this on the first set of countries, which had two near-term assessments, far enough apart to run a comparative test over ~30 days. That left me with 20 sets of results.

At 0.631, this result is immediately better than the average or random results.

This is a good start, but what about the quality of the results: how many of these were wholly wrong? (Remember, an opposite forecast is potentially catastrophic to the user.)

Even more encouragingly, of the 20 results, none were in the opposite direction.

This is a significant improvement in quality over the 0.667 score of random selection, even though these ‘scores’ are very close.

Overall, 40% of the forecasts were correct, but the remaining results were heavily skewed toward a pessimistic forecast, where the situation remained the same.

In some ways, that's not terrible guidance - it's better to tend towards caution or the worst case - but from an analysis perspective, it still means I can improve the analysis model quite a bit.

Next Steps

I aim to flip this breakdown to 60% correct and 40% incorrect (maintaining zero wholly incorrect results) in the next eight to 10 weeks. That gets the score closer to 0.5, which would be a pretty good accuracy rate, keeping in mind that this is all being done automatically with no human intervention: we are only assessing the model performance right now.

I'll share these results in the newsletter as we go and find a place to post them in the app.

But for now, this is a good start and it’s highly gratifying to see that DCDR is both fast, and reliable.

If you're not already a user, why not see for yourself? Go to dcdr.io to start a free trial