What the Hell Happened (and what can I do about it?)
Something happened this week that changed the conditions your organization is operating in. Most leaders will find out too late.
What the Hell Happened? is a weekly 30-minute briefing for executives in healthcare, higher education, manufacturing, logistics, transportation, and construction. Every episode takes the week's most consequential events — policy shifts, system failures, supply chain disruptions, regulatory changes — and works through what they mean for the people making decisions at the top.
Three segments. Every episode.
Readiness — what happened and why it matters. Resilience — what it reveals about the assumptions your organization is running on. Advantage — one specific action before Monday.
Hosted by Mike McCracken, founder of Southwind Planning Solutions, with decades at the intersection of emergency management and private sector operations.
On May 18, 2012, Facebook went public, and Nasdaq expected the opening to produce the largest IPO order volume in its history. What it got instead was a system that couldn't finish calculating the opening price. Engineers pushed through an untested workaround, and the exchange declared the problem solved. Trading was open. The tape was printing. From the outside, everything looked fine.
It wasn't. For the next two hours, thousands of trade confirmations never went out. Firms across Wall Street had placed orders and had no way to verify whether those orders had gone through. At 12:01 that afternoon, the CEO of a major market-making firm put it to Nasdaq's CEO in three words: we are all trading blind.
This week, Mike breaks down the Facebook IPO failure through what he calls the Unknown Status Problem, the doctrinal cousin of a common operating picture failure in emergency management. It's the gap between what an organization believes is true and what is actually true, and why that gap is more dangerous than a visible failure. A known problem tells you to fix it. An unknown status tells you to go find out, and most organizations aren't built to notice that second condition exists at all.
The episode also covers how UBS turned an unknown status into a documented $349 million loss by taking a second action to compensate for a first one it couldn't confirm, and closes with a Monday-morning audit any leader can run in under a minute: find one decision in your organization standing on a time-bound assumption right now, and ask who owns coming back to it if that assumption expires.
Frameworks referenced: Business Lifelines™, the Unknown Status Problem, Green/Gray status distinction
Companion content: This episode pairs with this week's Blue to Gray deep dive, expanding on the decision-shelf-life concept and the UBS compounding-loss mechanism.
There are usually three answers that you can get when you ask about the status of something that's very important. Either it's okay or it's not okay, or maybe you just don't know. Most organizations are pretty comfortable with the first two. If it's okay you just keep going. If it's not okay, then you should do something about it. But both of those answers give you a next move. It's that third answer that causes trouble, because I don't know doesn't usually look like a problem the way that the other two do. There's no red light there, there's no alarm, and nobody calls to say that something failed. Sometimes they say the dashboard just hasn't updated, or maybe the confirmation check never came back, or the person who normally checks in hasn't checked in yet today. Maybe the shipment was supposed to arrive, but nobody's reported that it didn't. None of that looks like an emergency. Oftentimes it looks like just another day. And somewhere along the way, without anyone deciding to do it on purpose, an organization starts treating one fact as if it were a different one. The fact is that we don't know anything is wrong. What gets treated as true instead is that we know that everything is okay. These aren't the same statement. On may eighteenth, twenty twelve, some of the biggest financial firms in the world got a very expensive demonstration of the difference. This is what the hell happened. Today we're going to talk about the unknown status problem.
SPEAKER_00
Businesses fail the same way disasters do. They just haven't given it a name yet. What the Hell Happened is a podcast that applies years of emergency management doctrine to the failure showing up in ordinary businesses. It's hosted by Mike McCracken, a field-tested emergency manager with more than 30 years of experience and who still trades professionals in the discipline today.
SPEAKER_01
Facebook was just going public. If you remember back in 2012, this wasn't just another IPO. This was the biggest and most watched public offering in many years, and NASDAQ had expected it to produce the largest opening order volume in their exchanges history. So keeping the mechanics simple because this isn't an episode about how the stock exchanges work. Before the trading opened on an IPO, the buy and sell orders pile up, and NASDAQ's system takes all of them, calculates the price where the most shares can trade, and matches everyone up, and then opens continuous trading. Now that calculation normally takes one or two milliseconds. NASDAQ had tested that system to handle up to about 40,000 orders. The problem is the Facebook IPO generated more than four hundred and ninety-six thousand initial orders. At ten fifty eight, the lead underwriter had already asked NASDAQ to push the opening back for five minutes. In hindsight, that was a sign that something wasn't behaving normally. At eleven oh five, NASDAQ tried to run the opening cross anyway, and it didn't work. The system was designed to recalculate the price every time an order changed or got canceled, which is normal, except the cancellations kept arriving faster than the system could finish each recalculation. It would start over and get partway through and then get interrupted again, and again and again, and it never caught up. From the outside this looked like a straightforward technology problem on one of the most anticipated IPOs in history. Senior NASDAQ officials convened what they called a code blue call. This was an internal emergency line for exactly this type of a situation. Engineers on that call had identified the loop and they came up with a workaround. They planned to switch to a backup matching engine with several lines of code controlling a validation check that was pulled out entirely. Nobody had tested that specific workaround for this specific situation, however, and according to the SEC's later investigation, the people who were on that code blue call didn't yet know what actually had caused the validation error in the first place. They had a fix for the symptom, but they had not confirmed or understood what the cause of the problem actually was. At eleven twenty five a Nasdaq executive approved it anyhow. That was the first moment of real uncertainty in this story, but it's not the most important one. Organizations make decisions based on incomplete information constantly. They have to. We're not interested in the easy hindsight take that they should have waited. What matters here is what happened afterwards. At eleven thirty the modified system ran the cross and it worked. Seventy five point seven million shares traded at forty two dollars a share, and continuous trading began. Facebook was open and the problem was solved. Except it wasn't. The process that handled the IPO had quietly fallen about nineteen minutes behind the actual time. The cross that printed at eleven thirty wasn't built on orders that came in through eleven thirty. It was actually built on the orders that only ran through around eleven eleven. There were more than thirty eight thousand marketable orders that were placed in that gap that never made it into the opening print. Over thirty thousand of them were simply stuck in the system. Nobody knew it yet, not NASDAQ or the firms that were trading on the results. In emergency management, emergency managers have a term for what had just split apart here. They call it the common operating picture. It's the shared and current understanding of what's actually true, and that's what drives and determines everyone's decisions. But for those two hours, NASDAQ actually had two operating pictures running at once. One of them said that the trading was normal. The other one was real also, but it was invisible, and it had already taken nineteen minutes of stale information and was falling further and further behind by the second. Only one of these was actually true, and almost nobody knew which one. However, there was a clue. NASDAQ's own systems had projected about eighty two million shares would probably trade. The actual print was a seventy five point seven million, which was a gap of six point three million shares. NASDAQ's chief economist had noticed this and understood that something was still wrong. According to the SEC, NASDAQ could have run a real time check, and that would have shown that the cross contained nothing that was submitted after about eleven eleven. That information wasn't unknowable. The problem was that it just hadn't been checked, and that's a very different problem than not being able to know about something at all. One of these issues called for managing uncertainty, but the other one is a choice that someone made, whether anyone recognized it as one or not. Then it even got worse. Within seconds of the trade opening, execution confirmations had stopped reaching the firms that had placed the orders. An execution confirmation answers the simple question, did the order actually go through? You can forget the Wall Street terminology for a second here, but imagine telling someone to buy one hundred thousand shares of a stock for you. Now the stock's trading and the price is moving, but you have no way to find out whether you actually own these shares yet. Did the order get filled? Did it get canceled? Or is it still sitting somewhere waiting and you don't know? And until you know what had happened to that first transaction, every decision that you make about your next move gets more complicated because you're making it on top or based around facts that you can't confirm. So at 1135, NASDAQ executives got together to talk about whether to halt the trading in Facebook altogether. But they decided not to. The continuous trading appeared to be running fine everywhere else, on NASDAQ and on other exchanges, and NASDAQ believed that the confirmation problem would clear itself up within a few minutes. Now we have to be careful here because it would be very easy fourteen years later to say that the obvious answer was to halt the trading and everything would have been fine. The evidence doesn't actually support that. What matters isn't whether eleven thirty five was the wrong call, it's the condition that that call was made under and what happened to that condition after the call was made. Engineers tried to restore the confirmations, but it didn't work, and so they tried again, and it still didn't work. The few minutes that NASDAQ had expected the fix to take turned into longer stretches of minutes. And then the details that matter the most in this entire case. NASDAQ never went back and revisited their original decision. And that wasn't because someone had checked and it was still fine, it was because nobody had checked it. The decision was built on the fact or the assumption that this will probably resolve itself shortly, and it was kept standing long after shortly had quietly come and gone. But nobody owned the problem or checked to see that it had been fixed. At twelve oh one, the CEO of a major marketing firm put it into five simple words in an email to NASDAQ's own CEO. He said we are all trading blind. He asked whether the trading should stop so the firms could figure out their actual exposure. By then the firms had been waiting thirty six minutes for their confirmations that still weren't coming, but the trading continued anyway. This is where an unknown status stops being passive and starts compounding the problem. If you don't know whether your first action actually went through, then what do you do next? Repeat it or cancel it? Or do you hedge it? Do you wait? Whatever you choose, you're not choosing that based on what had actually happened. You're making that choice based on what you believe had happened, which at that point is absolutely not the same thing. UBS found this out directly. They later reported a loss of three hundred and forty nine million Swiss francs that were tied to the IPO, because their pre-market orders were not confirmed for hours, and their systems entered them again and again, trying to make sure that the trades that their clients wanted actually went through. NASDAQ eventually filled all of those trades, both the original orders and the repeated ones, and UBS ended up holding far more Facebook stock than any client had actually ordered, and they were exposed to a position that nobody had intended to take. What this means in plain language is that they took an action and they couldn't verify that it actually happened, so they took a second action to compensate for the one that they couldn't confirm, and then they later found out that the first action had actually gone through the whole time. This wasn't a Wall Street problem, it's an operation systems problem, and it shows up everywhere. Send a supplier order twice because the first confirmation never came back and both orders got filled. Or resubmit a payment because no one can tell you if the first one cleared, and now two payments clear. Dispatch a second crew because the first crew hasn't reported in yet, and now there are two crews working on the same job with each one assuming that it's the only one that's been dispatched. These are different industries but the same architecture. The state of the first action is unknown, and the second action depends on it anyway. NASDAQ had its own version of this. For more than two hours engineers kept working the problem while the exchange's own call center took complaint after complaint from firms that were trying to figure out exactly what they owned. Then at 149 that afternoon, this was almost two and a half hours after the cross, the confirmations finally went out, and visibility came back all at once. And when it did, it didn't create new consequences. It revealed the ones that had already been quietly stacking up inside the two hours that nobody could see. There were more than thirty thousand stuck orders that got released into the market. This added roughly three million shares to the sell side, and it coincides with a fast drop in the Facebook's share price. In the middle of untangling all this, NASDAQ discovered something about itself. The workaround from that morning had left it holding a short position of more than three million Facebook shares that were worth about one hundred twenty nine million dollars. They had expected that position to be small, but it wasn't. For those two hours, NASDAQ had been running on its own believed operating picture the whole time, and it hadn't known that that picture was wrong either. The SEC later found serious problems in NASDAQ's technical design and their testing and controls and their decision making process. In fact, NASDAQ agreed to a ten million dollar penalty. At the time this was the largest penalty that was ever imposed on a stock exchange. The part that matters to me the most though, is not the technology failure. It's the decision that was made at eleven thirty five and it never got a second look. That's where we find the real gap. It's not that nobody tested the backup. Somebody made a time bound call. This will probably be fine shortly and no one is ever assigned to check back on it, whether shortly it actually arrived or not. These decisions have a shelf life, but nobody at NASDAQ owned watching that clock tick. There's a good chance that you've got a version of this similar operation running in your organization right now. Some call may have been made recently, maybe it was one that said we expected this to clear up soon, or it should be back online shortly, or maybe someone said we'll know more by the end of the day. It was probably made in a reasonable fashion with the information available at the time. Nobody's questioning whether that call made sense when it was made. The question that's worth asking is what happens to it after the decision's made. Does anyone go back to check to see if the window that it was built on had quietly passed? Or does it just keep standing by by default because nothing forced anyone to take a look at it again? So this week don't run a drill, run an audit instead. Just a short one, maybe ten minutes or so. Find one active decision in your organization that currently stands on a time bound assumption. This doesn't have to be anything dramatic. It could be as simple as a vendor timeline or an IT fix, or a staffing gap that you're covering until. And then ask yourself out loud in the room, who's responsible for coming back to this if the expectation behind it doesn't hold true? Don't ask who gets notified if it fails outright, because that part's usually covered. And there's an alarm that triggers that. But who's on the hook to notice that the clock ran out and that nobody actually checked? If nobody in the room can answer that within thirty seconds, then you found your gray area. And you found it before it actually cost you anything. Because that's what actually separates the two answers that most people treat as the same thing. Something has failed tells you to fix it. But we no longer know if what we're relying on is still true, tells you a different story. It tells you to go look. Most organizations build around the first one. Very few are built for the second. On may eighteenth, 2012, Facebook opened. The tape kept printing, and from the outside the whole thing looked fine. But underneath it for two hours some of the largest trading firms in the world couldn't answer the most basic questions in their business. What do we actually own right now? This wasn't because anyone had lied to them or because the market had crashed. It was because a status that was supposed to be current had quietly gone stale, and the decision that were built on the top of it never got checked again. Someone finally noticed and said we are all trading blind, and by the time he said it, the decision that put them there had already been standing unexamined for the better part of an hour. If you don't know, the status is not green. The status is actually gray, and gray status only stays survivable for as long as somebody's actually watching it. This means that the real question isn't whether your organization has an unknown running somewhere around within it right now, because it almost certainly does. The real question is whether there is anyone who's actually assigned to notice that. That's it for this week. I hope you enjoyed this episode and this conversation. If you did, drop a note in the comments and let me know about it. If there are things that you'd like me to explore more, let me know that as well. If you'd like to learn more about this incident, check out my deep dive. It's available in the Blue to Gray newsletter. The link is in the show notes. I'm Mike McCracken and we'll see you next time. What the hell happened is produced by Southwind Planning Solutions LLC. If this episode was useful, the Blue to Gray newsletter goes deeper into this and many other topics each week. You can subscribe for free on Beehive. The link is in the show notes. If you're sitting on an assumption that you have not verified it more than you'd like to, or maybe there are other challenges where reality has interrupted your planet, there's a starting point for that in the show notes as well. You can find us on Apple Podcasts, Spotify, anywhere else at TListen Podcasts. You can also find me on LinkedIn and on our website at www.southwindplanets.com. I'm Mike McCracken. Thanks for listening, and I'll see you next time.