Skip to content
BIZENIUS

The supervisory early warning system review checklist: ten areas it is tested on

BIZENIUS Advisory Team · Last updated: 28 August 2026

Written and reviewed by the BIZENIUS advisory practice — senior practitioners from risk, treasury, finance and supervision.

Ten areas where a supervisory early warning system is tested — by its own board, by a bank that deteriorated without being flagged, and by a peer review — what a sound answer looks like in each, and the symptom that gives a weak one away.

In short

  • Supervisory early warning systems are rarely found to be wrong in one decisive way. They decay: thresholds stay still while the banking system moves, the response ladder is never exercised, and each annual review refreshes the figures rather than re-examining the reasoning.
  • The ten areas below are where a system is most consistently tested. Each is stated as what a sound system can show, followed by the symptom that gives a weak one away.
  • Use is the single most diagnostic area, and the only one that cannot be improved by improving the system: name a supervisory decision it changed.
  • The pattern across all ten is that the defects are found by asking for specifics — a date, a name, a decision, a dismissal reason — rather than by reading the framework document.
  • Run the ten against a single institution that deteriorated in the last two years. Most of what is wrong will surface in that one case.
On this page
  1. 1. Data quality at collection
  2. 2. Stated reporting lag
  3. 3. Method transparency
  4. 4. Threshold derivation
  5. 5. The error the authority accepts
  6. 6. The response ladder
  7. 7. Dismissals recorded
  8. 8. Back-testing against real cases
  9. 9. Independent challenge
  10. 10. Use
  11. Using the list

Supervisory early warning systems are rarely found to be wrong in one decisive way. They decay. Thresholds stay still while the banking system moves beneath them, the response ladder is never exercised, the analyst who built the method leaves, and each annual review refreshes the figures rather than re-examining the reasoning that produced them.

The ten areas below are where such a system is most consistently tested — by its own board, by a peer review, and most unforgivingly by a bank that deteriorated without being flagged. Each is stated as what a sound system can show, followed by the symptom that gives a weak one away.

1. Data quality at collection#

**Sound:** returns are validated at entry, reconciled across forms, and late or revised submissions are tracked by institution. **Symptom:** data quality is described as an ongoing initiative, and nobody can say which institutions revise most.

2. Stated reporting lag#

**Sound:** the delay between a condition arising in a bank and its appearance in the system is measured and written down, per input family. **Symptom:** the lag is treated as zero, so a signal is discussed as though it described today.

3. Method transparency#

**Sound:** someone currently employed can explain how a score is produced and which inputs drive it. **Symptom:** the method is understood by one person, a departed contractor, or a vendor.

4. Threshold derivation#

**Sound:** each threshold has a stated basis and a stated expected flag rate. **Symptom:** thresholds were set once at launch, and the share of institutions flagged has drifted without anyone deciding it should.

5. The error the authority accepts#

**Sound:** the balance between missing a deteriorating bank and disturbing a sound one is a recorded governance decision, revisited deliberately. **Symptom:** the operating point is wherever the modelling defaults left it, and nobody has been asked to approve it.

6. The response ladder#

**Sound:** each supervisory state carries a stated consequence, and someone can name a month in which that consequence occurred. **Symptom:** the ladder exists in the framework document and every institution is supervised at the same intensity regardless of its state.

7. Dismissals recorded#

**Sound:** alerts closed without action carry a recorded reason, and the pattern is reviewed. **Symptom:** the alert log shows closures without reasons, so nobody can distinguish a miscalibrated system from an ignored one.

8. Back-testing against real cases#

**Sound:** every institution that deteriorated in recent years has been run back through the system to establish when it would have been flagged. **Symptom:** back-testing is described as planned, and the last real case was reviewed narratively rather than against the system.

9. Independent challenge#

**Sound:** a function that did not build the system reviews it and has standing to say it should not be relied on. **Symptom:** the system is validated by the unit that owns it, and no review has ever recommended against using an output.

10. Use#

**Sound:** supervisory decision records cite the system, and at least one decision in the last year was changed by it. **Symptom:** the system is current, maintained, reported on, and absent from every record of a supervisory decision actually taken.

Using the list#

The pattern across all ten is that the defects are found by asking for specifics — a date, a name, a decision, a dismissal reason — rather than by reading the framework document. A documentary review will pass a system that fails eight of these.

Run the ten against a single institution that deteriorated in the last two years rather than against the system in the abstract. Most of what is wrong will surface in that one case, and it will surface as a sequence of dates rather than as an opinion.

BIZENIUS works through this review with supervisory teams, and delivers the underlying material as a programme for authorities building the capability in-house. Both are available in English and French; scope is agreed with the authority and confirmed on enquiry.

Frequently asked

How do you review a supervisory early warning system?

By asking for specifics rather than reading the framework document, because a documentary review will pass a system that fails on use. Ten areas carry most of the diagnostic weight: data quality at collection, whether the reporting lag is measured rather than assumed to be zero, whether the method can be explained by someone currently employed, how thresholds were derived, whether the balance between the two error types is a recorded governance decision, whether the response ladder has ever been exercised, whether alert dismissals are recorded with reasons, whether real deteriorations have been run back through the system, whether independent challenge exists, and whether any supervisory decision was actually changed by it.

What is the strongest single sign that a supervisory early warning system is weak?

No supervisory decision has ever been changed by it. A system can be current, well maintained, regularly reported on and entirely absent from every record of a decision actually taken — at which point it is a reporting exercise rather than a supervisory instrument. The closely related sign is a response ladder that exists in the framework document while every institution is supervised at the same intensity regardless of its assessed state. Both point at the same underlying gap: the analytical work was built and the link to supervisory action was not, and that link is the cheapest part of the system to construct and the part no vendor supplies.

How often should a supervisory early warning system be reviewed?

Annually at minimum, and additionally whenever the banking system it watches changes materially — a wave of consolidation, a new category of institution, a shift in funding structure — because thresholds derived from one population stop describing a different one. The more useful question is what a review consists of. A sound review record shows reasoning revisited: why this threshold given the current population, whether the flag rate is still what was intended, whether the drivers of the score have drifted. A weak one shows consecutive versions differing only in their figures, which is the signature of a refresh that updates numbers without re-examining what produced them.

Who should carry out the review — the analytics unit or someone independent?

The unit that runs the system can work the whole list, and should, because most of the ten are answered from records it already holds. Two areas benefit from someone outside it: threshold derivation and use. Both require judging work the unit itself produced, and both are where a confident internal answer is hardest to challenge from inside. A practical arrangement is for the owning unit to run the ten, with those two tested by internal audit, by a supervision department that did not build the system, or through peer review with another authority — which has the additional benefit of showing how a comparable system is calibrated elsewhere.

More where this came from

Browse the full resources hub, or subscribe in the footer for occasional substantial pieces.

BIZENIUS

Speak to an expert

Tell us where you stand — an expert replies within one business day.

Phone *
Area of interest
+ Add a message or details (optional)

We only use your details to respond to your enquiry. See our Privacy Policy.