Skip to content
STOIntelligence
All insights

The furnace was in a mode nobody had named

Tom MankowskiFounder & PrincipalSeptember 20, 20266 min read

Six Minutes: the Monaca sequence, reconstructed from the CSB report. Times approximate.

On June 4, 2025, Furnace 5 at Shell Polymers Monaca was coming back from a coke trap cleanout. At 2:14 p.m., an engineer meant to open one 36-inch isolation valve and opened the other. Minutes later both were open. Cracked gas from the furnaces still running backflowed into the Furnace 5 firebox, and about six minutes after that first command it reached the lit pilots. The blast ruptured the firebox wall. Shell estimated $95 million in damage and the furnace was down for seven months. Nobody was hurt, but 15 people were evacuated and a contractor had to be rescued from the elevator next to the furnace.

The U.S. Chemical Safety Board published its final report on September 16. Most coverage will focus on its two headline issues: reliance on administrative controls and a poorly designed operator interface. Both are real. Read it as a turnaround professional, though, and a different pattern jumps out.

Every safeguard asked what mode the furnace was in. None had an answer.

That afternoon the furnace was isolated, pilots lit, and partway back to service. That state did not exist on any document. Look at how each layer of protection was tied to an operating mode:

  • Alarms. Firebox pressure, carbon monoxide and methane alarms were suppressed by design in pilot-only mode to cut nuisance alarms. The carbon monoxide alarm would have fired about 90 seconds before ignition.
  • Procedure. The start-up procedure begins after isolation is removed and assumes a cold furnace. The first pilot is not lit until step 98. With pilots already burning, there was no procedure for the path the crew was on.
  • Exclusion zone. Site policy clears nonessential people during start-ups and abnormal situations. Coming out of a maintenance isolation counted as neither, so no zone was set.
  • Engineered protection. The licensor supplied a sequenced key interlock and a SIL 2 backflow trip at the local panel. The panel logic was built around operating modes like decoking, so it could not remove a double isolation, and those interlocks never applied in that state.
  • Hazard analysis. The 2023 hazard analysis revalidation identified reverse flow into an offline furnace as a multiple-fatality scenario and credited a methane alarm and a start-up procedure. On the day, the alarm was suppressed and the procedure was not in use.

Five layers, one shared blind spot. That is not five failures. It is one failure, repeated: nobody had defined the transient as a mode.

The job aid only worked for people who did not need it

The task had been done successfully 19 times. That record hid a defect: the job aid's steps were out of sequence. The engineer assigned that morning had never done the task, followed the aid, and stalled. A colleague who had done it before spotted the problem immediately: the safety system had to be in debug mode. Switching modes refreshed the screen back to the top, where the furnace-side valve sat. The three valve tags differed only in their last digit: 511, 512, 513.

Nineteen successes did not prove the task was safe. They proved experienced people had been quietly compensating for it.

The alarm that fired was read as confirmation

When the wrong valve moved, an unexpected-state alarm fired about five minutes before ignition. The console operator knew the plan was to open the tower-side valve, saw a nearly identical descriptor, acknowledged it, and moved on. Good pre-job communication created the exact expectation that made the warning look like good news.

One reasonable call quietly erased the map

The site left the pilots on to protect the refractory after a previous damage event. That is a defensible asset-care decision. But it departed from the isolation plan without the required deviation review, and it pushed the job off the only start-up procedure they had. Add an induced draft fan left in speed control after a firmware update during the outage, and the furnace re-entered service in a configuration nobody had reviewed as a whole.

In scheduling terms, this was a logic failure, not an activity failure. The individual tasks got done. The ties between them (who acts, from what state, with which protections live) were never drawn, and the error cascaded downstream the same way broken logic does in a schedule network.

The fix was already on a list

Reprogramming the panel to handle double isolation had been identified as a project. It sat on an improvement list behind higher priorities. The hazard analysis team judged engineered controls "grossly disproportional" to the risk reduction. The eventual bill was $95 million and seven months of a cracking furnace.

Before your next return to service

This is not one company's problem. Every site has transitions that live between procedures. Here is the test worth applying to any turnaround or outage return-to-service plan:

  1. Map states, not steps. For each major equipment item, list every state it passes through from full isolation to normal operation, including the awkward ones: pilots lit, one valve open, fans in manual. If a state has no name, it has no protection.
  2. Fill five cells for every state: governing procedure, live alarms, active interlocks, energy and ignition sources, exclusion zone. Every blank cell is a finding.
  3. Treat plan deviations like scope changes. Leaving pilots lit, running a fan in a non-normal mode, bypassing a safety system: each gets the same review as a scope add, with every downstream step rechecked.
  4. Qualify people for the transition, not the task. Ask who has done this exact move, from this exact state, before. If the answer is nobody on this shift, that is a readiness gap, not a staffing detail.
  5. Test job aids on someone who has never used them. If a job aid only works for the experts, it does not work.
  6. Give transitions their own lines in the schedule. A single "return to service" milestone hides the riskiest hours of the event. Break it out, assign owners, attach readiness criteria.
  7. Pull the someday list. Any deferred project that would add an engineered safeguard to a transient state deserves a fresh look before the next event, not after it.

We built a self-check for the first two: the Transient Test. Pick one piece of equipment, walk its states, and see which cells come up blank. It runs in your browser and sends nothing anywhere.

Turnarounds are planned around the work. Incidents cluster around the moves between the work. That is where our readiness and assurance reviews look first.

Source: U.S. Chemical Safety Board, Furnace Explosion and Fire at Shell Polymers Monaca, Investigation Report No. 2025-05-I-PA, September 2026.

Want a defensible read on your next event?

Start a conversation