What the Hell Happened (and what can I do about it?)
Something happened this week that changed the conditions your organization is operating in. Most leaders will find out too late.
What the Hell Happened? is a weekly 30-minute briefing for executives in healthcare, higher education, manufacturing, logistics, transportation, and construction. Every episode takes the week's most consequential events — policy shifts, system failures, supply chain disruptions, regulatory changes — and works through what they mean for the people making decisions at the top.
Three segments. Every episode.
Readiness — what happened and why it matters. Resilience — what it reveals about the assumptions your organization is running on. Advantage — one specific action before Monday.
Hosted by Mike McCracken, founder of Southwind Planning Solutions, with decades at the intersection of emergency management and private sector operations.
What the Hell Happened (and what can I do about it?)
Episode 1: Nobody Had the Button
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
What the Hell Happened? — Episode 1: Nobody Had the Button
Host: Mike McCracken
Sometimes it's a policy shift. Sometimes it's a system that fails. Or it could be that a market moves. Often leaders find out too late. The signal was there, but nobody turned it into something they could act on before the moment passed.
What the Hell Happened is a weekly briefing for leaders who would rather be ready. I'm Mike McCracken, and I've spent over 30 years working in situations where planning gets interrupted by reality. This program looks at not just what the hell happened, but also what can be done about it.
At some point in your career, maybe already, something is going to go wrong in your organization. Not a small thing, a real thing. And when it does, the question won't be whether someone saw it coming. The question will be whether that information reached the person with the authority to act on it, and if it reached that person in time for it to matter.
Readiness
I'm going to open each episode with a single moment from a story about a problem or a failure that an organization has experienced in real life. Crisis and threats to an organization can appear at any time and they appear in many different ways. Sometimes they're obvious and dramatic, but sometimes they're more subtle. We're going to take a look at examples of each and figure out just what the hell happened, and what can we do to keep it from happening again.
We'll look at a moment where the status said that operations were normal and everything was well managed. There was nothing to worry about. It was a blue sky day. But the reality was actually that the skies were gray. They had been gray for a while, they'd just been overlooked. That's called the blue skies illusion.
The operations plan was 59 pages long. The plan identified two people who were authorized to stop the show. But nobody could reach either one of them in time. Nobody else was authorized to push the button.
Does your organization have a button? Does everyone in your organization know who holds it? Do they know how to contact that person if the button needs to be pushed? These are the questions we're going to look at and discuss in this episode.
Today we're going to walk through exactly what happened at Astroworld, not the tragedy, but the organizational sequence of events that made it possible. By the end of this segment, you'll see that it wasn't a concert problem. It was a leadership and information problem. A version of it could exist in your organization right now.
Before we get into exactly what happened, I want you to hold on to one question. How does critical information get from the front lines of your organization to the person who has the authority to act on it? And have you ever actually tested that pathway to see if it's connected or if it works? Keep that question in your head because in this episode you're going to see exactly what happens when the answer is no.
Let's look at the incident itself. It occurred at NRG Park in Houston, Texas. There were about 50,000 people in attendance at a music festival called the Third Annual Astroworld Festival. The headliner was Travis Scott, a hip hop and rap artist. A mass casualty event was declared after the concert had started at approximately 9 PM. By the time it was all finished, ten people had died and there were hundreds injured. There was a crowd surge and people could not move and they could not breathe.
This didn't start at 9:30 PM. It started hours earlier. The perimeter gates were breached earlier in the day several times. Security and operations treated each breach as an isolated nuisance. Nobody connected the dots and nobody asked what the pattern might mean. By the time the crowd surge began, the safety margin was already gone. The risk didn't appear suddenly. It had been accumulating all day long.
Here's the sequence of events and what failed in what order. At around 9:30 PM first responders began to receive injury reports. There was information moving around at ground level, but it wasn't reaching the decision authority. The show continued for approximately 40 more minutes. Firefighters had phone numbers for the production team, but not a direct line. Those phone numbers also fail under mass crowd signal load, which is common at large events where thousands of people are on their cell phones simultaneously. Nobody with the authority to stop the show knew what was happening on the ground.
There was a 59 page operations plan and a detailed chain of command involved in the planning of this event. According to that plan, only two people held the authority to stop the show. One was the executive producer for Live Nation, and the other was the festival director, also with Live Nation. Travis Scott was the artist, but he was not identified in the plan as a stop authority. The Houston police chief was on record saying we don't hold the plug. The plan did not address crowd surge scenarios. The result was that everyone had a role, but nobody had the button.
After it was all over, fingers started being pointed. Every party pointed to another. The police, the production, the artist. A crowd safety expert during the civil trial that followed found that any key decision maker should have been able to initiate a show stop through a simple button to push. The button existed on paper, but in practice it was unreachable. The grand jury declined to indict Travis Scott in 2023, but the organizational failure was fully documented.
This is about looking at what was occurring and how it might apply to any organization to prevent things from happening, whether it's a large scale life or death situation or another type of business failure. This was not a concert production problem. It was an organizational design problem where the warning signs were visible all day. They just weren't connected. The information existed, the authority existed, but the pathway between the information and the authority did not exist.
Ask yourself right now, if something started to go wrong in your organization today, who has the authority to make the first call? How would they find out that it was time to make that call? What's the pathway? Not generally speaking, but specifically. This wasn't a Travis Scott problem and it wasn't a Live Nation problem. It was a leadership and information problem, and a version of it may exist in your organization right now.
Resilience
In this next segment we're going to take a look at why these things happen, not just at Astroworld, but in every organization that has ever had a crisis it didn't see coming. There are three mechanisms at work here. Once you see them, you'll recognize them in your own organization.
Disasters don't happen because one thing breaks. They happen because leadership and the front line are operating in two different realities. The gap between those two realities can grow over time, quietly and invisibly. At Astroworld, the front lines were seeing gate breaches, building crowd pressure, and injuries beginning to occur. At the leadership level, the dashboard in the command post was still showing a well managed event, until news finally reached them that the reality was different. The decay was already advanced before anyone with authority knew it had started.
The event's medical and security plans relied on a centralized dispatch system, but when the surge began, frontline personnel couldn't communicate with the production trailer. Signals dropped and critical data was trapped in manual pathways. This is called the delta of decay. It's the widening distance between what is actually happening and what leadership believes is happening.
What does this look like in a corporation or other organization? C-suite executives and senior leaders often make high stakes decisions based on reported data. But that data passes through middle management before it reaches the top, and under pressure, negative signals often get filtered. Not out of malice, but out of instinct. Supply chain friction, a cyber vulnerability indicator, an operational near miss. Each one gets softened slightly before it goes up the chain. When the quality of the information pathway degrades, middle management often filters out the negative frontline signals to maintain the illusion of stability. By the time the C-suite realizes the signal is compromised, the delta of decay has already triggered a systemic failure. Leadership may be looking at a green dashboard while the skies have already turned gray.
At Astroworld, security saw the perimeter breaches in the morning. Each one was logged, managed, closed out, and treated as normal. Nobody asked what the pattern meant at scale. This is normalized deviance. When small failures become acceptable, they become invisible. In your organization it could be a system outage that happens once a quarter and gets patched. Maybe it's a vendor who consistently misses the service level agreement and always has a good reason. Or a process that requires a workaround that everyone knows about but nobody documents. Each one is a gate breach. The safety margin is shrinking and nobody is connecting the dots.
Astroworld had a 59 page operations plan. The plan was not the problem. The problem was the gap between the plan and the operational reality. Possessing a documented plan does not equal operational capability. If the people who need to execute the plan can't find the button under pressure, the plan is simply fiction. This applies to your business continuity plan, your crisis communications protocol, your IT disaster recovery documentation. The question is never whether the plan exists. The question is whether your people can execute it without looking for instructions that may not be there.
This all comes down to how information is managed. There's an information management matrix that I use in training, and it helps sort out what to do with information, how to find it, and how it should move through an organization so it reaches the right place. When information comes into the organization, the first question to ask is whether it belongs to you. Do you have the authority to act on it? If yes, handle it, document it, and close the loop. If no, then whose information is it? Do you know who owns it? If you do, do you know how to get it to them? If you don't know who owns it, what do you do with it? Do you hold it, protect it, document it? You don't let it disappear. The loop needs to be closed. How do you document what you did, and who needs to know what you did?
At Astroworld, nobody had a clear answer to any of those questions. When the answer to whose information is this is silence, the information usually goes nowhere.
So we've looked at three mechanisms. The delta of decay, the widening gap between what's happening and what leadership knows. Normalized deviance, when small failures stop registering as warnings and become routine. And the paper plan illusion, the difference between having a plan and having a tested pathway. Every one of these was present at Astroworld, and every one of them is measurable before a crisis finds you.
The organizations that handle this well don't have better people. They have a smaller delta of decay and they actively measure and monitor it. They treat near misses as data, not nuisances. And they test whether information can actually reach the authority it needs to reach before the moment it's needed. The difference between a plan and a tested pathway is the difference between paper and practice.
Advantage
Here's something you can take into your next leadership meeting. One simple exercise, 20 minutes with your leadership team. It will tell you more about how your organization actually functions than any dashboard you'll watch or any report you'll read this week.
Pick one realistic scenario for your organization. Something more than a fire drill. A key system goes dark, a critical vendor can't deliver on deadline, a serious incident happens after hours when management has gone home. Gather your leadership team and put three questions on the table.
First, if a major operational failure or cyber incident happened today, does your frontline manager have the explicit preauthorized mandate to shut down operations immediately? Without waiting for an executive vote. Without needing to reach two specific people whose phone numbers may or may not work under load. If the answer is no, or if nobody is sure, you have the Astroworld situation. An authority problem that just hasn't had a date put on it yet.
Second, how do you know that the data you use to monitor your operational health is accurate? Not reported, but accurate. Are you auditing the integrity of your information pathways, or are you trusting a single reporting chain that might get overwhelmed or filtered? If a problem were developing two levels below your current visibility, how long before it reaches you? At Astroworld that gap was 40 minutes. What is your organization's gap?
Third, where is your organization assuming perfect performance from a vendor, a system, or your infrastructure? What happens to that assumption when that node goes down? How fast does performance degrade and who is most likely to notice it first? Your most dangerous dependencies are the ones that have never failed, because nobody has ever had to find out what happens when they do.
Take those questions and run them against your scenario. Who knows whose information it is when something goes wrong? Is that documented somewhere people can actually find it? Has that pathway been tested? Actually tested, not assumed. Document what you find, even if the finding is a gap, because a documented gap is the beginning of a fix. An undocumented gap is a liability waiting for something to happen.
Three questions. The authority question: does your front line have the mandate to act? The masking metric: how long before a real problem reaches you? And the delta of decay: what are you assuming works that you have never tested?
Run the exercise and see what surfaces. If this episode has surfaced a concern, that's the right reaction. That kind of discomfort is the signal working correctly. The next structured place to look is the Signal Integrity Assessment, designed to surface exactly this kind of gap and find a crisis before it finds you. The link is in the show notes.
Close
At Astroworld, a 59 page plan existed with two people authorized to stop the event. Neither could be reached in time, and warning signals had been visible all day. Leadership and the front line were operating in different realities until the gap closed violently. That gap is the delta of decay, and it can be managed with disciplined information monitoring.
So what can you do about it? Take 20 minutes. Ask three questions. Find that gap before that gap finds you.
Next week on What the Hell Happened, we'll take a look at another case and discuss what the hell happened and what can we do about it.
What the Hell Happened is produced by Southwind Planning Solutions, LLC. If this episode was useful, the Blue to Gray newsletter goes deeper into this and many other topics each week. You can subscribe for free on beehiiv. The link is in the show notes. If you're sitting on an assumption you haven't verified in longer than you'd like to admit, or there are other challenges where reality has interrupted your plans, there's a starting point for that in the show notes as well. You can find us on Apple Podcasts, Spotify, or anywhere else you listen to podcasts. You can also find me on LinkedIn and on our website at southwindplanning.com. I'm Mike McCracken. Thanks for listening, and I'll see you next time.