What the Hell Happened (and what can I do about it?)
Something happened this week that changed the conditions your organization is operating in. Most leaders will find out too late.
What the Hell Happened? is a weekly 30-minute briefing for executives in healthcare, higher education, manufacturing, logistics, transportation, and construction. Every episode takes the week's most consequential events — policy shifts, system failures, supply chain disruptions, regulatory changes — and works through what they mean for the people making decisions at the top.
Three segments. Every episode.
Readiness — what happened and why it matters. Resilience — what it reveals about the assumptions your organization is running on. Advantage — one specific action before Monday.
Hosted by Mike McCracken, founder of Southwind Planning Solutions, with decades at the intersection of emergency management and private sector operations.
On a July evening in 2024, a piece of equipment failed on a transmission line in Northern Virginia. The grid's protection system responded correctly. So did the protection systems inside a cluster of nearby data centers. Within 82 seconds, roughly 1,500 megawatts of demand vanished from the grid — not because anything malfunctioned, but because two systems, each behaving exactly as designed, produced a combined outcome nobody had modeled.
This week, Mike breaks down the July 2024 Northern Virginia load-loss event — and the NERC and FERC actions still unfolding around it in 2026 — through what he calls the ownership gap: the space between two individually correct systems that nobody was ever assigned to own.
Using the Business Lifelines™ framework, this episode covers:
Why "each system worked" doesn't mean "the organization is safe"
Why an unknown condition — Gray — is more dangerous than a known failure
Where ownership gaps hide in healthcare, manufacturing, and logistics
A 20-minute R.I.S.E. exercise to find and close your own ownership gaps before they find you
www.southwindplanning.com
www.bluetogray.beehiiv.com
SPEAKER_00
At approximately seven o'clock on a July evening, a piece of equipment failed on a two hundred and thirty thousand volt transmission line in northern Virginia. The electrical power grid's protection system responded. And so did the protection systems inside a cluster of data centers that were nearby. Both of these systems did exactly what they'd been built to do. Within seconds, roughly 1,500 megawatts of electrical demand disappeared from the power grid. Nobody deliberately disconnected 1,500 megawatts of power. That outcome emerged from the way that these two systems responded to each other. And some version of this may already be possible inside of your organization, because it won't look like a failure until the interaction has already produced the outcome. I'm Mike McCracken and this is What the Hell Happened. Sometimes it's a policy shift. Sometimes it's a system that fails. Or it could be that a market moves. Often leaders find out too late. The signal was there, but nobody turned it into something they could act on before the moment passed. What the Hell Happened is a weekly briefing for leaders who would rather be ready. I'm Michael Crackett, and I've spent over 30 years working in situations where planning gets interrupted by reality. This program looks at not just what the hell happened, but also what can be done about it. Here's what actually happened in order. At 7 p.m. Eastern Time on July 10th, 2024, a lightning arrestor failed on that 32 kilovolt line. This failure produced a fault that the line couldn't clear on its own, so it locked it out. The fault and the automatic reclosing sequence produced six faults and six voltage depressions over the next 82 seconds. The electrical power grid's own protection system did its job every single time. It detected each fault. It cleared it and it moved on. But something else was watching these same six voltage depressions at the same time. Some of the facility controls at nearby data centers. These systems were built to protect very expensive, very sensitive equipment, and they were configured to count electrical disturbances. Three within a minute could trigger a transfer to backup power. And once that's triggered, it stays there until a person switches it back by hand. Six depressions in 82 seconds tripped that logic. Roughly 1500 megawatts of data center load left the grid immediately. This did not happen because a system had failed. It actually happened because the facility's own protections did exactly what they were supposed to do. About 1260 megawatts of power didn't come back to the grid for hours. Both the frequency and the voltage on the power grid rose. This sequence of events created a scenario where the light electrical loading combined with a surplus generation pushed both the grid speed or the frequency and the grid pressure or the voltage into high territory. This forced the operators to take direct manual action to stabilize the entire system. The operators had to manually pull the capacitor banks offline just to bring the voltage back down. If you read that back, you'll notice what is missing and what's not missing. What's not missing here is a single equipment failure that wasn't cleared correctly on either side. There were two systems that they were each doing exactly what they were built to do, and they produced a result that the space between the two systems was actually never built to handle. Hang on to that in your mind for a few minutes and we'll come back to exactly what kind of space that was and why that matters more than you'd expect. This isn't a story that I'm dragging back up because nothing better happened this week. It's actually live and going on right now. In May of 2026, the North American Electric Reliability Corporation, or NERC, which is a nonprofit organization whose responsibility it is to maintain the reliability the entire North American power grid. NERC issued what's called a Level 3 Essential Action Alert. Level three is NERC's highest alert tier, and issuing one of these alerts requires approval from NERC's own board of trustees. This alert actually got that approval on April 16th and it was issued on May the 4th. All registered entities had until August 3rd, which is just two weeks ago, to answer 33 questions about where they actually stand on managing this type of behavior. Meanwhile, on july sixteenth, the Federal Energy Regulatory Commission ordered the NERC to file a mandatory reliability standard covering these loads by the end of this year, with more to follow in early 2027. So this isn't a case study that I randomly pulled out of an archive somewhere. The clock is still running on this case, and the FERC didn't wait for the NERC answers to arrive before it took action. That order came more than two weeks before the response deadline even hit. One more thing that's worth knowing, and then I'll move on. The NERC is also watching a second, separate problem. This problem includes and looks at large computing loads. Specifically, these are the rhythm of AI training jobs cycling between full power and near idle. These can produce real oscillations in the same frequency range that the grid's own natural rhythms live in. Basically what that means in short is that these AI data centers don't just consume a steady baseload power. They can actually behave like massive electrical subwoofers vibrating at the exact pitch that shakes the entire physical grid. This isn't what happened in Virginia. This is a different mechanism, and I talk more about it in my blue to gray newsletter, and that's available at Beehive if you'd like a deeper version. I bring this up because it's the same warning from a different angle. These loads don't behave like the static lines on a spreadsheet that they were modeled after. The broader pattern is very familiar. This is where a system gets studied, it's connected and commissioned, and then it's treated as though its behavior will stay static from that point on. These two facilities got modeled as passive numbers on an interconnection study somewhere. But in practice, they're actually volatile, autonomous systems that are wired directly into the physical backbone of the country's power supply, and nothing in the record suggests that the connection point itself was ever anyone's specific job to monitor or watch. So here's the question to consider going into the next few minutes of this episode. Consider what's running inside of your organization right now exactly the way it's supposed to. Whose interaction with something else that's also running exactly right? Has ever once anybody been assigned to watch that function to see what those interactions are putting out? That's what's happened here and that's what's still unfolding. What it might mean for you is a different question. If you're listening to this episode and you're in healthcare, higher education, logistics, or construction, you might be thinking, I don't run a power grid and I don't own a data center, so what does this have to do with me? Well what you're actually likely running is a similar type of architecture within your own organization. You've probably got at least two systems or processes or vendors that each make a correct decision within their own rules. But the harder question is whether there is actually anyone that has responsibility or accountability when both of these act at the same moment. When that moment comes, you might find out that nobody has claimed that space. And that's probably not because of anyone's failure. It's most likely because no one's ever asked that question. And that's the bridge. So now let's get specific about why this might happen. Let's talk about resilience and what resilience tells us. Two systems are getting built. They both get reviewed and approved on their own terms. Each one performs exactly as it's intended every single day. And the organization's confidence rests on each system separately because that's how each one got evaluated in the first place. Nobody actually sits down and asks a question of what happens when both of these systems act at once. This isn't an oversight. It's just that no one's ever been assigned with that task to take a look at this outcome. A disruption finally forces this question to be asked, and only then does anyone discover that the space between these two never had an owner. This is where my concept of business lifelines actually earns its keep. You can use it to find before anything breaks which lifeline has nobody actually watching it or monitoring it, especially where two systems meet. Let's look at the incident in Northern Virginia through that lens. These systems didn't malfunction. What the larger system hadn't adequately modeled was what their protection decisions would produce together. Or in management terms we would call this an over ownership gap. Decision authority and governance was clearly assigned on each side of these fence lines, one inside the utility and one inside of each facility. There's nothing in the record that suggests that anyone had that job across the connection between the two. Coordination had the same shape. There was a defined handoff on each side, but no evidence of one where they actually met. This isn't a paperwork problem, it's an ownership problem. And it's precisely the kind of gap that this framework is built to catch. If you're actually running it as a live question instead of filing it as an exercise. You remember that space that I asked you to hold on to back when we were talking about readiness? Here's what it was a gap like that where nobody was watching and nobody was responsible. And that's what produces what the business lifelines call a gray status. This is the most dangerous status in the entire model. And that's not because it's merely unclear, it's because it means that the information hygiene is already broken down and that nobody's maintaining real situational awareness. A black status on a lifeline is an outright failure, and this is at least something that you can see and respond to. But that gray status means that you can't see the problem well enough to know that you even have one. The situation in Northern Virginia was never trending toward a failure that anyone would have caught in time. Viewed through these business lifelines, the scene between the two systems was effectively gray from the start. The individual systems appeared understood, but their combined condition was actually not. So let's give this a name because you're going to need to spot it again after today. Let's call this the ownership gap. It's not a system that broke or information that quietly went stale. It's a space between two systems that were each working fine, but that nobody ever had an assignment or a task to own. That's a different question than the one that we asked in the last episode. The question here is not whether a source was still trustworthy. The question here is whether anyone had the responsibility or the ownership for the space where the two trustworthy things meet. In healthcare, picture a scheduling platform that fills a shift correctly based on availability. Now picture a separate credentialing platform or system that correctly blocks a clinician that the scheduling platform just assigned because that person's licensing renewal hasn't cleared yet. In this instance, neither system reports a failure. The schedule still shows that the position was filled, but the floor actually has a vacancy with no one to fill that position in real life. The question is now who owns that combined outcome? Who's responsible for the vacancy on the floor? You can apply this same idea to manufacturing or logistics or transportation. For instance, picture an inventory system that correctly pulls back orders because of its own demand signal, while at the same time a carrier or a lane allocation system correctly reprioritizes capacities because of its own signal. Each one of these decisions on its own is completely defensible. However, the combination leaves your materials stranded on a dock somewhere, and nobody's actually accountable for that combination. Only the two systems that did exactly what they were designed to do. The organization that catches this type of gap early doesn't necessarily have better technology. What it most likely has is someone whose job it is who's been assigned to monitor and close that gap. So here's the takeaway from all this, and it's not a new framework. You've already got frameworks. This is a way to run it against a gap that you've probably never gone looking for. And here's what you can do about it. This is your advantage. Don't go back to your office tomorrow and buy another monitoring dashboard. And don't write a fifty page manual that nobody's going to read. This isn't a visibility problem that you fix with more reporting. This is an ownership problem and it needs an owner, not a document. Here's something that will actually work. Design a twenty minute session with your leadership team and run it through the RISE system. That's rank, identify, shape, and execute. Let's start with rank. First, pick the operational outcome that you want to look at. Not the system, the outcome. The outcome that would hurt the most if two of your systems behaved correctly and still produce this outcome wrong together. Then you identify, take that outcome and run it through these four questions. First, what can change? Which two systems, processes, or vendors independently shape this outcome without either one ever seeing the other's decision? Second, how will we know? How will we know what actually tells you the combination went wrong when neither system on its own would ever report a failure? And third, who can act? Who has the standing authority over the combined outcome? Not either system alone, and can they move without waiting on a committee to make a decision? And fourth, can we recover and reconcile after the outcome? Once it's caught, can the work keep going on in a controlled state and does the recovery restore the actual record and not just the technology? Next we'll move on to shape. Decide on purpose who owns this space between the two systems, who has visibility into it, and who has the authority to step in and is responsible for putting things back together afterward. If the honest answer in your room is nobody, then that's not a failed exercise. That is actually going to be your finding. And that brings us to execute. Now name a person and then actually test it. That does not mean to have a conversation about the plan. Run an actual live pathway test to find out whether the ownership that you just assigned on paper actually holds up under pressure. If you're in healthcare, name someone who owns the combined outcome when your scheduling system and your credentialing systems interact. Don't identify who owns each system. Identify who owns the intersection and outcome of those systems. If you're in manufacturing or logistics or transportation, maybe you'd want to identify your two most interdependent automated systems. Write those down and write down who's accountable for what they produce together, not just what each one produces on its own. If this examination and discussion has surfaced an ownership gap, don't turn that into another finding on a list. Make an assignment and test it and find out whether that authority worked before you need it. To wrap things up, think back to Northern Virginia. In that situation, the transmission protection worked. The data center protection worked. But what didn't exist was anyone responsible for what happened the moment both of those correct decisions landed in the same place at the same time. The grid isn't in trouble because the technology failed. It's in trouble because two correct systems met in a space that nobody was watching. They'll come from everyone doing their job correctly, with nobody accountable for what those correct decisions add up to. The system doesn't have to fail for you to lose control of the outcome. That's what the hell happened for this week. Now let's do something about it. Remember readiness, resilience, and advantage. You can check out the show notes to find links to the blue to gray newsletter for the deep dive on Beehive. And you can also find the website www.southwindplanning.com for more information about finding a signal integrity assessment that can examine some of these issues within your organization. I'd also love to hear from you, so drop me a message or send me an email and let me know what you think and let me know of any ideas you have for this or any upcoming episodes. I'm Mike McCracken. Thanks for listening, and I'll talk to you next time. What the Hell Happened is produced by South Wind Planning Solutions LLC. If this episode was useful, the Blue to Gray newsletter goes deeper into this and many other topics each week. You can subscribe for free on Beehive. The link is in the show notes. If you are sitting on an assumption that you have not verified in longer than you'd like to admit, or maybe there are other challenges where reality has interrupted your planning, there's a starting point for that in the show notes as well. You can find us on Apple Podcasts, Spotify, or anywhere else at QListen Podcast. You can also find me on LinkedIn and on our website at www.southwindplanning.com. I'm Mike McCracken. Thanks for listening, and I'll see you next time.