
Two 737 MAX aircraft crashed within four and a half months of each other, on 29 October 2018 and 10 March 2019, killing 346 people. The fleet was then grounded for 619 days, and the legal consequences ran until a federal judge dismissed the remaining criminal charge in November 2025, seven years after the first accident.
It is studied as an operational risk case because almost nothing about it was an accident in the ordinary sense. The exposure was created by a commercial decision, carried into the design by an engineering accommodation, concentrated onto a single sensor, and hidden from the people who would have had to respond to it. Every one of those is a control point, and each one is recognisable inside a bank.
The sequence matters more than any single date, because the interval between the two accidents is where the operational risk lesson sits.
| Date | Event |
|---|---|
| 29 October 2018 | Lion Air Flight 610 crashes shortly after take-off from Jakarta. 189 people are killed. |
| 10 March 2019 | Ethiopian Airlines Flight 302 crashes shortly after take-off from Addis Ababa. 157 people are killed, bringing the total to 346. |
| 13 March 2019 | The Federal Aviation Administration grounds the 737 MAX fleet. Several other regulators had already done so in the preceding days. |
| 18 November 2020 | The FAA lifts the grounding after 619 days, subject to design changes and new pilot training requirements. |
| 7 January 2021 | Boeing enters a deferred prosecution agreement with the United States Department of Justice, paying over 2.5 billion dollars in total. |
| 5 to 6 January 2024 | A door plug separates from Alaska Airlines Flight 1282 in flight, with no fatalities. The FAA grounds 171 MAX 9 aircraft the following day. |
| May 2024 | The Department of Justice determines that Boeing breached the compliance obligations of the 2021 agreement. |
| 29 May 2025 | A non-prosecution agreement is reached, worth about 1.1 billion dollars. |
| 6 November 2025 | A federal judge dismisses the criminal charge on the basis of that agreement. |
Read the first three rows again. Four and a half months separated the two accidents, and the aircraft kept flying through them. That interval is the part of the case that belongs to risk management rather than to engineering, because by then the firm had a loss event, an investigation, and a known failure mode, and the fleet was still in service.
The 737 MAX was a re-engining of an airframe first certified in the 1960s. The new engines were larger and more efficient, and fitting them under a wing designed for smaller ones meant mounting them further forward and higher. That changed how the aircraft behaved at high angles of attack.
The design response, according to the accident investigations and the regulatory reviews that followed, was a software system called the Manoeuvring Characteristics Augmentation System, or MCAS, which trimmed the horizontal stabiliser nose-down when it judged the angle of attack to be too high. The purpose was to make the new aircraft handle like the old one.
The commercial reason for wanting that is the part worth dwelling on. If the MAX handled like earlier 737s, it could share their type rating, and airlines could move pilots onto it without simulator training. Simulator training is expensive and slow, and avoiding it was a genuine selling point against the competing aircraft. A cheaper transition for the customer was, in effect, a design constraint on the engineer.
The exposure was not created in the engineering department. It was created by a commercial commitment, made upstream, that the engineering department then had to satisfy. Operational risk almost always enters this way. A promise is made about cost, or speed, or continuity of service, and the constraint it imposes is inherited by people who did not make it and who are not measured on it.
The investigations into both accidents found that MCAS acted on the reading from a single angle of attack sensor. If that sensor gave a faulty reading, the system responded to a condition that was not occurring, and pushed the nose down against the pilots who were trying to raise it.
Aviation design is built on redundancy, and a single input driving an authority of that kind is the sort of arrangement redundancy exists to prevent. Two further conditions turned a design vulnerability into an accident: MCAS was not described in the flight manual, and pilots were not trained on it, so the crews encountering it had no name for what was happening to their aircraft.
Note what that chain implies for measurement. There is no single cause to attach a probability to. There are four control points, each of which would have broken the sequence, and the event required all four to fail together. Operational risk failures are almost always shaped this way, and it is why control self-assessment asks about layers rather than about causes.
Aircraft certification in the United States has long relied on delegation. The regulator authorises employees of the manufacturer to perform specified certification work on its behalf, because the manufacturer holds the technical expertise and the regulator does not have the staff to duplicate it. The arrangement is not a secret and it is not unique to aviation. Model validation, internal audit and internal ratings-based credit models all run on some version of it.
The difficulty with delegated assurance is structural rather than personal. The people performing the check are paid by the organisation being checked, and they sit inside its schedule pressure. Independence in that setting depends on reporting lines and on culture, not on job titles, and both of those are difficult to observe from outside.
This is the closest the case comes to a directly transferable lesson for a financial institution. A bank that lets the desk that owns a model choose who validates it has built the same arrangement, and will discover the same thing about it under the same conditions: when the calendar is tight and the finding is inconvenient.
Reading this case as a story about one aircraft programme. The transferable content is the governance pattern: a commercial commitment made upstream, an engineering or modelling accommodation made downstream to satisfy it, an assurance function that is not independent of either, and a disclosure gap that leaves the last line of defence, in this instance the flight crew, without the information required to respond.
The standard taxonomy classifies operational loss events by cause. This one is not a rogue trader, not a system outage, and not external fraud. It is a failure in the design, delivery and management of a product, which is the category most often skipped in a risk taxonomy discussion because it sounds like someone else’s problem.
| Dimension | This event | Why it matters for measurement |
|---|---|---|
| Event type | Failure in product design, delivery and process management | Rarely well populated in internal loss data, so the modelled distribution is thin exactly where the tail sits |
| Frequency | Very low | A firm may never have observed one, so the empirical frequency is close to zero |
| Severity | Extreme, and spread across several years | A single event can exceed the sum of every other operational loss the firm has recorded |
| Correlation | Strongly correlated with strategic and reputational risk | Treating operational risk as a standalone silo understates the total |
| Time to crystallise | Seven years and counting from the first accident | A loss reserved in one year keeps being revised for a decade, which distorts backtesting |
Consider a manufacturer, unrelated to this case, with 40 years of internal operational loss data. Its largest recorded loss is 80 million dollars and its annual operational losses average 25 million. A product design failure then produces a loss of 20 billion dollars. What does the loss history say about the capital that should have been held?
Answer: the loss history does not contain the answer, and no amount of statistical refinement extracts it. This is the standing limitation of loss distribution approaches to operational risk, and it is the reason scenario analysis exists as a separate input rather than as a cross-check. Scenarios are the only place where a loss the firm has never suffered can enter the calculation, and they are worth something only if the workshop is allowed to write down a number that embarrasses the business plan.
A risk manager reading only the settlement figures would draw the wrong picture of the loss. The legal penalties are the visible part and the smaller part.
| Component | Reported figure | When it landed |
|---|---|---|
| Deferred prosecution agreement | Over 2.5 billion dollars, including a criminal penalty of 243.6 million, about 1.77 billion of compensation to airline customers, and a 500 million victim fund | January 2021 |
| Non-prosecution agreement | About 1.1 billion dollars, including a 487.2 million criminal penalty, 444.5 million to a beneficiaries fund for the families, and 455 million of investment in compliance, safety and quality | May 2025, dismissed by the court in November 2025 |
| Grounding, production halt and delivery delays | Widely estimated at around 20 billion dollars of direct cost | 2019 onward |
| Lost orders, financing and reputational effects | Estimated by analysts at up to 60 billion dollars of indirect cost | Not fully crystallised |
The estimates in the last two rows are analyst figures rather than reported ones and should be read as orders of magnitude. Even so, the shape they describe is unambiguous: the fines are a rounding error against the operational consequences. The 619-day grounding stopped deliveries, and deliveries are when an aircraft manufacturer is paid.
The 2024 door plug event belongs in the same account for a different reason. Nobody died, and by the standards of the earlier accidents it was a minor event. Its significance is that it arrived while the firm was operating under a compliance agreement arising from the first losses, and the Department of Justice subsequently determined that the compliance obligations of that agreement had been breached. That is what turned a settled matter back into an open one four years later.
Case study questions on this material rarely ask for the timeline. They ask which control failure a described scenario most closely resembles, or which category of operational risk event a set of facts belongs to, or why an internal loss distribution understates a firm’s exposure to it. The examinable content is the pattern, not the chronology.
Five points survive the transfer from aviation to finance.
A commercial promise is a risk decision. When a commitment about cost, speed or continuity is made before the technical work is scoped, the constraint it creates is real and someone downstream will absorb it. Risk functions that review products after design see this too late.
Redundancy is worth its cost precisely on the inputs that seem reliable. A single input driving an automated action with real authority is an exposure whether that input is a sensor, a price feed, a rating, or a single approver.
Assurance that reports to the business it assures is not assurance. Delegated checking works while the calendar is comfortable and stops working when it is not, which is exactly when it is needed.
The last line of defence has to be told what it is defending against. The flight crews were the final control and were not informed of the system they were fighting. The equivalent in a bank is an operations team executing a process whose failure modes have never been explained to it.
And the interval between the first loss and the second is the part to study. There was a loss event, an investigation and a known failure mode, and the fleet continued flying for four and a half months. Whether a firm can stop something profitable while it does not yet know how bad the problem is, on incomplete information and with the commercial cost immediate and certain, is the question this case actually asks. Most risk frameworks are written as though that decision is procedural. It is not, and the interval is where the framework is tested.
MCAS, the Manoeuvring Characteristics Augmentation System, trimmed the horizontal stabiliser nose-down when it judged the angle of attack to be too high. It was added because the MAX carried larger engines mounted further forward and higher than on earlier 737s, which changed the handling at high angles of attack. Making the new aircraft behave like the old one allowed it to share the existing type rating, so airlines could transition pilots without simulator training.
Because the loss required a sequence of control failures rather than a single technical fault. A commercial commitment created a design constraint, the design response concentrated authority on one sensor with no redundancy, the assurance process ran through employees of the firm being assured, and the system was left out of the manual and out of pilot training. Each of those is a control that would have broken the chain, and all four are recognisable inside a financial institution.
The FAA grounded the fleet on 13 March 2019 and lifted the grounding on 18 November 2020, a period of 619 days. Direct costs are widely estimated at around 20 billion dollars, with analyst estimates of indirect costs running as high as 60 billion. The legal settlements, over 2.5 billion dollars in January 2021 and about 1.1 billion in May 2025, are the visible part of the loss rather than the large part.
Boeing entered a deferred prosecution agreement in January 2021. In May 2024 the Department of Justice determined that the compliance obligations of that agreement had been breached, and a proposed plea agreement submitted in July 2024 was rejected by the court in December 2024. A non-prosecution agreement worth about 1.1 billion dollars was reached on 29 May 2025, and on 6 November 2025 a federal judge dismissed the criminal charge on that basis.
Because the model is fitted to losses the firm has already suffered, and an event of this kind is far outside that range. A manufacturer with 40 years of data whose largest recorded loss is 80 million dollars has an empirical maximum of 80 million, while the realised loss is 250 times that figure and equivalent to 800 years of its typical annual losses. No percentile of a distribution fitted to that history reaches the realised number, which is why scenario analysis is a separate input rather than a cross-check.
It is the practice of a regulator authorising employees of the manufacturer to carry out specified certification work on its behalf, because the technical expertise sits with the manufacturer. The weakness is structural rather than personal: the people performing the check are paid by the organisation being checked and sit inside its schedule pressure, so independence depends on reporting lines and culture rather than on job titles. Model validation and internal audit inside a bank raise the same question.
That the interval between the first loss and the second is where a risk framework is tested. By March 2019 there had already been an accident, an investigation and a known failure mode, and the fleet was still flying. Whether a firm can suspend something profitable on incomplete information, when the commercial cost is immediate and certain and the risk is not yet quantified, is a judgement that most frameworks describe as procedural and that almost never is.
Loading comments...
Add your Thoughts: