CPS 230 Tolerance Levels: Can You Defend the Number?
9 MIN READ
CPS 230 has been in force since 1 July 2025. What changed six weeks ago is that the last of the scaffolding came down.
Two things happened on 1 July 2026. Transitional arrangements for service provider contracts signed before CPS 230 commenced expired — every material arrangement now has to meet the standard or be documented as an exception. And APRA’s targeted amendments, finalised on 30 April 2026, began to apply.
Most of the commentary since has been about contracts. Contract remediation is visible, finite work with a lawyer attached to it, so it gets attention.
The harder obligation has been sitting there since day one, and it is not a contract. It is a number.
CPS 230 requires your board to approve tolerance levels for disruption to each critical operation. Not a policy about tolerances. Not a framework for thinking about tolerances. A specific figure — the maximum period a business service can be unavailable, the maximum data loss you can absorb, the minimum resources needed to keep operating — which your directors have challenged and signed.
Here is the question worth asking before your next board pack goes out: where did that number come from?
A tolerance level is a quantitative statement whether you treat it as one or not
APRA’s Prudential Practice Guide CPG 230 is specific about what goes into setting a tolerance. Entities are asked to consider the maximum amount of time a business service can be unavailable before the impact is unacceptable, the maximum time allowed to recover the information assets behind that service, the maximum data loss the business can tolerate, and the minimum people and resources required to keep operating.
How long the business service can be down before the impact becomes unacceptable. Measured in hours.
How long you have to restore the information assets sitting behind that service. Measured in hours.
How far back you can reconstruct, which then drives backup frequency. Measured in records or elapsed time.
The people, information assets and infrastructure needed to keep operating. The floor that makes the other three achievable.
Every one of those is a measurement. Three of the four are measured in units — hours, records, headcount — and the fourth is the resourcing floor that determines whether the first three are achievable.
The word “unacceptable” is doing the heavy lifting. Unacceptable to whom, and above what dollar figure? A four-hour outage in payments is not unacceptable because four hours feels long. It is unacceptable because somewhere past that point the accumulated cost — customer remediation, regulatory exposure, funding disruption, reputational damage, the cost of standing up a manual process — crosses a threshold the board is unwilling to carry.
That threshold is risk appetite. Risk appetite is expressed in dollars. So a tolerance level is the point where a duration curve crosses a dollar line.
If you have never drawn either curve, you have not set a tolerance. You have nominated one.
The board did not approve a recovery time objective. It approved a statement about how much loss the institution is prepared to absorb, expressed in hours.
The 24-hour problem
Under CPS 230, a disruption to a critical operation beyond its approved tolerance level triggers a notification to APRA as soon as possible, and in any case within 24 hours. That sits alongside the separate 72-hour obligation to notify a material operational risk incident.
The 24-hour clock is the interesting one, because of what it implies about the number.
Notifying a tolerance breach is an admission with two parts. The first is operational: something broke and stayed broken longer than we said it would. That part is uncomfortable but normal — outages happen, and a standard that assumed otherwise would be useless.
The second part is analytical, and it does not surface for weeks. Once the incident is closed and the supervisor comes back with questions, someone has to explain how the tolerance was arrived at, whether it was ever achievable with the recovery capability actually in place, and why the dependency that caused the breach was not accounted for when the figure was set.
APRA’s supervision programme for CPS 230 runs prudential reviews across a subset of entities in 2025–26, another subset in 2026–27, and moves to business-as-usual supervision in 2027–28. Between 2025 and 2027 APRA has said it will engage in heightened supervision where a material event occurs. A tolerance breach is precisely the kind of event that moves an entity into that population.
And there is a capital consequence attached. Where an entity identifies material weaknesses in its operational risk management, APRA expects to be kept informed of remediation progress — and for banks and insurers, expects additional capital to be held until the remediation is complete. Weak tolerance-setting is not a documentation finding. It has a price.
Why most tolerance levels are guesses
Having watched risk quantification frameworks get built and adopted — and, at Defence, watched a methodologically sound one get rejected because the tooling was too painful to use — the failure pattern here is familiar. The number gets set by the fastest available method, and the fastest available method is a room full of people.
The workshop anchors on the first figure spoken
Six or seven people from operations, technology and risk get an afternoon. Someone opens with “we could probably live with four hours.” Everyone adjusts around that anchor. The final figure is a round number in a spreadsheet cell with no stated assumptions behind it, which means there is nothing for a director to challenge — and CPS 230 explicitly asks directors to challenge these.
The number is often circular
The most common tolerance-setting error is quiet and hard to spot: the tolerance ends up equal to whatever the current recovery capability can deliver. If the disaster recovery runbook says four hours, the tolerance becomes four hours.
That inverts the logic. The tolerance is supposed to express what the business can absorb. The recovery capability is supposed to be engineered to meet it. When the tolerance is set to match the capability, the standard has been satisfied on paper and the gap it was designed to expose has been hidden instead.
Worse, it makes the number fragile. A tolerance set at the edge of current capability breaches the first time anything goes slightly wrong.
The dependencies you cannot see are excluded
This is where APRA’s 30 April 2026 industry letter lands hardest. Reporting on a targeted review of large banks, insurers and superannuation trustees, APRA found that governance, risk management, assurance and operational resilience practices were not keeping pace with AI adoption — and identified the widest gap between current practice and regulatory expectation in third-party and supply chain risk. The letter noted an over-reliance on vendor presentations and summaries without sufficient examination of the impact on critical operations.
That is a quantification finding dressed as a governance finding.
Your critical operation depends on a material service provider. That provider depends on fourth parties CPS 230 also expects you to identify. You cannot run a scanner across any of it. You cannot instrument it. You get a SOC 2 report, an architecture diagram and a slide deck.
So when the tolerance is set, that entire dependency chain is either excluded from the analysis or represented by an assumption nobody wrote down. The chain is then the single most likely cause of the breach that triggers the 24-hour notification.
What a defensible tolerance looks like
The difference is not sophistication. It is whether the working exists.
| Question a supervisor will ask | Workshop number | Modelled number |
|---|---|---|
| How was this figure derived? | Management judgement | Loss distribution crossing stated appetite |
| What loss does breaching it represent? | Not calculated | A dollar range with confidence bounds |
| What assumptions is it built on? | Undocumented | Listed, versioned, individually testable |
| Is it achievable with current capability? | Assumed — it was set to match | Measured; the gap is an explicit output |
| How were unscannable dependencies handled? | Excluded | Estimated through structured expert input |
| What happens if threat conditions change? | Re-run the workshop | Re-run the model, compare the deltas |
The method that produces the right-hand column is not exotic. It is the same approach that underpins FAIR and every quantified operational risk model used in insurance: decompose the critical operation into what it depends on, express each dependency as a probability range rather than a point estimate, simulate the outcome many thousands of times, and read the answer off the resulting distribution.
Two features of that approach matter specifically for CPS 230.
It gives you a curve, not a point
Once you have the full distribution of disruption durations and their associated losses, the tolerance is no longer a nominated figure. It is the point where cumulative modelled loss crosses the board’s stated appetite. The board is now approving something it can interrogate: not “is four hours right?” but “are we comfortable carrying the loss that sits beyond this line, given how often we expect to be past it?”
It has an answer for the things you cannot measure
The dependencies that break tolerances — a provider’s platform, a fourth party two steps down the chain, an AI service embedded in a vendor product — are exactly the ones no telemetry reaches. Structured expert elicitation, with the accuracy of contributors tracked over time, produces a defensible probability range for those dependencies instead of a blank in the model. An estimate with stated uncertainty is a governance artefact. An omission is a finding.
This is the part of the market where our own work sits, and it is why we lead with the assets you cannot scan rather than the ones you can. In a CPS 230 context, the unscannable dependency is usually where the tolerance actually fails.
Five things worth doing before your next board pack
Pull the derivation for one tolerance level
Pick your most material critical operation and ask for the working behind the number. If what comes back is a workshop attendance list, you have found your gap — and you have found it before a supervisor did.
Test each tolerance for circularity
Compare every tolerance against the recovery time your capability actually delivers. Where they match exactly, the tolerance was almost certainly reverse-engineered from the capability. Reset it from business impact and let the gap be visible.
Convert the hours into dollars
For each critical operation, put a range on what an outage costs per hour — customer remediation, manual workarounds, regulatory and contractual exposure. Until that range exists, “unacceptable impact” is an opinion.
Model the dependency chain, including what you cannot scan
Map each critical operation to its material service providers and known fourth parties. Where you have no visibility, estimate deliberately and record the assumption. Given APRA’s April findings, expect this to be examined.
Give the board something it can challenge
CPS 230 asks directors to challenge and approve tolerances. A single figure cannot be challenged. A distribution, a stated appetite line and a list of assumptions can be. That is the difference between a board that approved a number and a board that exercised oversight.
The bottom line
CPS 230 did something the Australian regulatory environment had mostly avoided until now: it required a board to sign a specific, quantitative statement about operational loss, and then attached a 24-hour reporting obligation to being wrong about it.
That is a quantification mandate. It does not use the word, and it never will — regulators write outcomes, not methods. But you cannot defend a tolerance level you did not derive, and you cannot derive one without modelling the loss it is supposed to cap.
Up from 46% in 2025. GuidePoint 2026 State of Cyber Risk Management.
AI-enabled breaches average around US$6M. IBM Cost of a Data Breach 2026.
Up from 20% a year earlier — a dependency almost no tolerance model includes.
Boards are asking for financial framing because financial framing is what they can act on. The regulators are asking for the same thing in different vocabulary. And the losses those tolerance levels are meant to be calibrated against keep growing.
The entities that come through the 2026–27 supervision cycle comfortably will not be the ones with the most conservative tolerances. They will be the ones who can explain, on demand, exactly how each figure was reached, what it assumes, and what changes when the assumptions do.
Every number in a board pack should be able to survive the question “how did you get there?” CPS 230 has now made that question mandatory, on a 24-hour clock.
Can you defend your tolerance levels?
CyQuantiFi models cyber and operational loss in dollars — including for the dependencies your scanners cannot reach.
Related reading
- Prudential Standard CPS 230 Operational Risk Management — APRA
- New insights for ensuring compliance with APRA’s CPS 230 — Corrs Chambers Westgarth
- APRA calls for a step-change in AI-related risk management and governance — APRA, 30 April 2026
- Everyone’s Quantifying Cyber Risk Now. Almost No One Is Checking the Inputs. — CyQuantiFi
- Risk quantification on the CyQuantiFi platform
