Beyond the Hackathon: Building a Continuous Innovation Engine
You’ve seen how the hackathon ends. Two days of pizza, whiteboards, and real energy. A dozen teams demo something clever. Leadership claps. Somebody wins a gift card. Then everybody goes back to their real jobs on Monday, and within a week the winning idea is a Slack thread nobody reads anymore.
That’s not a knock on your people. The ideas were often good. The problem is what happened after the applause.
When researchers went looking for what becomes of hackathon projects, the answer was sobering. In a study of nearly 12,000 of them, only about 7% showed any activity six months later. The rest just stopped. Not because they were bad. Because there was nowhere for them to go.1 Steve Blank, the entrepreneur and teacher who helped start the lean startup movement, has a name for this: innovation theater. You adopt the rituals of innovation, the hackathon, the idea portal, the “innovation day,” and you mistake the ritual for the outcome. The motion feels like progress. It isn’t.2
Underneath that is a harder problem. Most companies aren’t short on ideas. They’re short on an engine to carry an idea from “interesting” to “in market.”
The hackathon is an event. Innovation is a system.
Run a system like an event and you get exactly what you’d expect. A burst of activity, then nothing.
The engine, not the event
Think about what an event is. It has a start date, an end date, a budget line, and a room. When it’s over, it’s over. That’s fine for a conference, but it’s terrible design for the one capability that decides whether your company still matters in five years.
An engine is different. An engine runs continuously. Fuel goes in one side, motion comes out the other, over and over, without stopping. A continuous innovation engine has three chambers, and ideas move through them on a loop that never ends: ideation, evaluation, and experimentation. You generate possibilities. You evaluate which ones are worth a bet. You run cheap experiments to see if the bet pays. The results feed the next round of ideas, and the loop turns again.
This is well-studied ground. Researchers describe the companies that keep winning as “ambidextrous,” able to run the core business and explore new ones at the same time. They talk about “dynamic capabilities,” the standing ability to sense what’s shifting, seize what’s worth seizing, and reshape the organization to capture it.3 The common thread is simple. None of it is an event. These are capabilities you build and keep. You don’t schedule them for a Tuesday in Q3.
And notice where the failure actually lives. It isn’t idea generation. McKinsey’s research on what separates the high-growth innovators from everyone else lands on an unglamorous point. Innovation is, at its core, a resource-allocation problem, not a creativity problem. The weak link is the messy middle: picking the right ideas, funding them, and moving people and money toward them when the moment calls for it.4 The hackathon is great at the front of the loop and silent on the rest. That’s why it goes nowhere.
Fund it like a portfolio, not a budget
So why does the middle break down? Follow the money.
Most companies fund innovation the way they fund everything else. Once a year, in a budget cycle, with a business case that projects three years of returns for something nobody has tested yet. It’s a strange ritual when you say it out loud. You’re asking for a confident forecast about the most uncertain work in the building, then locking the number in for twelve months.
The bigger problem isn’t speed. Standard net-present-value math rejects exactly the bets you most want to make, because a high-potential, high-uncertainty idea looks terrible on a spreadsheet that demands precision it can’t yet provide. Annual budgeting doesn’t just move slowly. It filters for the wrong things.
The fix is to stop thinking like a budget owner and start thinking like an investor. A venture capitalist doesn’t write one check and walk away. She makes a small bet, watches what the team learns, and decides whether to write a bigger check or stop. The money moves in stages, and each stage buys information. Every idea really needs two instincts, applied in turn. The innovator’s instinct, bold enough to imagine it. And the investor’s instinct, disciplined enough to ask whether it’s worth funding. The engine is where those two instincts meet.
The model has two moving parts. The first is metered funding. You fund an idea in increments, and the team only gets the next increment when it shows real learning. You’re not betting the farm on a pitch. You’re buying the right to see the next data point.5 The second is the growth board. Instead of an annual planning committee, a small group of senior leaders meets regularly to review live bets and decide which ones get more fuel and which ones get retired. It’s the same discipline a venture capitalist uses, run inside your own walls.6
What makes those moving parts work isn’t the mechanics; it’s trust. You’re replacing permission with evidence. Teams don’t lobby for budget. They earn the next stage by learning something real. Companies already run this. Adobe’s Kickbox program hands any employee a red box with a thousand-dollar prepaid card and no approval process, on the bet that a thousand small experiments teach you more than one committee-approved project. Amazon makes every new idea start as a written press release and FAQ for a product that doesn’t exist yet, so the bet gets pressure-tested in plain prose before a dollar is spent, then hands it to a team small enough to feed with two pizzas. When the Dutch bank ING reorganized into small standing squads, it swapped the annual plan for a quarterly review where every team reports what it learned and where the money should go next. It went from five or six big launches a year to shipping every two to three weeks.7
Different mechanics, same core move. The money follows the learning, and it moves on a cadence measured in weeks and quarters, not years.
The hard part: evaluation without killing the good ones
This is where most engines seize up. It’s worth slowing down on, because it’s the part nobody has fully solved.
If money moves in stages, someone has to decide at each stage whether the idea lives or dies. Get that wrong in one direction and you fund noise, pouring good money into a pitch that sounded exciting and tested flat. Get it wrong in the other direction and you pull up the seedling before it can grow, killing the awkward early idea that would have become the whole business if you’d given it one more cycle. Both failures are expensive. Most organizations manage to commit both.
Let’s start with the kill-too-early problem, because it’s the sneakier one. Clayton Christensen, the Harvard professor whose book The Innovator’s Dilemma explained how disruption topples good companies, co-wrote a piece with a title that says it all: “Innovation Killers.” His argument was that the financial tools we trust most, discounted cash flow and NPV, are quietly biased against innovation. They measure a risky new bet against a default assumption that the current business will keep humming along safely if you do nothing. That assumption is almost never true. Compared against that rosy fiction, the new bet always looks worse.8
Run every idea through that filter and you will, with perfect discipline and clean spreadsheets, reject your own future.
The fix isn’t to evaluate less. It’s to evaluate differently. Rita McGrath, a Columbia Business School professor who studies strategy under uncertainty, and her colleague Ian MacMillan gave us the cleanest tool for this back in 1995. They call it discovery-driven planning. Instead of funding a plan and hoping the assumptions hold, you make the assumptions the plan. You write down everything that has to be true for the idea to work, rank those things by how much they’d hurt if they’re wrong, and spend money to test the scariest ones first.9
In practice it looks like a simple map. Plot every assumption behind an idea on two axes: how important it is, and how much evidence you actually have for it. Say your team wants to launch an AI-powered feature. The assumption “customers will trust an AI-generated answer enough to act on it” is both critical and, right now, unproven. Top-right corner. That’s the first thing you test, maybe by putting a rough prototype in front of ten customers next week, not by greenlighting a six-month build. Most of those small tests are supposed to fail, because a cheap failure this week saves an expensive one next quarter.10 The point is always the same. You’re not trying to be right about the idea. You’re trying to buy evidence faster than you burn cash.
Now the other failure, funding the loser too long. Back in 1976, the organizational psychologist Barry Staw ran a study he called “Knee-Deep in the Big Muddy.” He found that people escalate their commitment to a failing course of action, and they escalate hardest precisely when they were the ones who chose it. Decades of research since have said the same thing. A meta-analysis pulling together thirty years of studies confirmed that sunk costs, personal responsibility, and ego threat all reliably push people to throw good money after bad.11 It’s human wiring. It’s why “let’s give it one more quarter” quietly becomes three years and a dead product.
The engine’s answer to both failures is the same, and it’s almost boring. Decide the kill-or-continue criteria before you start, not in the meeting where everyone’s ego is on the line. When a team gets metered funding, the next tranche is tied to a milestone the group agreed to in advance. Did the scary assumption survive contact with reality, or not? That one discipline protects the good idea from a nervous spreadsheet and protects the company from the sunk-cost spiral at the same time.
Run it on cadence, because most experiments fail
Once you’re evaluating this way, a surprising thing becomes obvious. You need volume.
The most rigorous look at this comes from Ron Kohavi and Stefan Thomke, two researchers who documented how the biggest technology companies run online experiments at scale. The numbers are humbling. At Google and Bing, only 10 to 20% of experiments produce a positive result. Across Microsoft, it’s roughly a third positive, a third flat, a third actively negative. These are among the best experimentation teams on earth, and most of their ideas are wrong. Their response isn’t to get smarter about picking. It’s to run more. Each of these firms runs well over ten thousand controlled experiments a year.12
Airbnb’s team once tested about 250 ideas and found 20 that moved the metric, a better than 90% failure rate. Those 20 lifted booking conversion around 6%, worth a fortune. Or take the single most valuable idea in Bing’s history, an engineer’s tweak to how ad headlines displayed. It sat in a backlog for six months, tagged low priority, until someone finally ran the experiment. It raised revenue 12%, more than a hundred million dollars a year in the US alone. Nobody’s judgment caught it. The experiment did.
That’s the case for cadence. If most ideas are wrong and you can’t reliably tell which in advance, then the rate at which you can cheaply test them becomes your real edge. Most progress comes from hundreds of small improvements, not a few big swings. The team that learns fastest wins, not the one with the best hunches. So shorten the loop. A cycle you run weekly beats one you run monthly, not because speed is a virtue on its own, but because every turn is another chance to learn before the money runs out.
Knowing whether the engine is actually running
If you build this, how do you know it’s working? Not by counting ideas generated. That’s a vanity metric. It measures the fun part and ignores whether anything shipped.
Measure the engine, not the enthusiasm. The useful metrics are the leading ones. In your live experiments, is customer behavior actually moving the way the idea predicted? That tells you something months before the financials do. A blunter, old-school measure still works too: the share of revenue coming from things you launched in the last few years. 3M built a culture around keeping that number near 30%. They later let it slide into the low teens, though, which is its own lesson. This is an engine you have to keep fueling, not a switch you flip once.
And watch the balance of your bets. The old three-horizons habit is a useful gut check. Put most of your effort on the core, some on adjacent growth, a little on the genuinely new. With one caveat: in fast markets, the “long-term” horizon can show up next year, so don’t starve it.13 I’d add a fourth horizon to the picture, and it’s the one most portfolios miss. Beyond the bets you’re actively planning sits everything you’re not yet watching closely enough, the signals that haven’t earned a place on any horizon yet. The three horizons are where you plan. The fourth is where you watch. A healthy engine does both, on the same Monday.
What happens Monday
Innovation isn’t a day you put on the calendar. It’s an engine you build and keep running. Ideation feeds evaluation. Evaluation feeds experiments. Experiments feed the next idea. And the whole thing is funded in small stages by people willing to kill their own darlings when the evidence says to.
The hackathon isn’t the enemy. It’s just the spark. The real question, the one that separates the companies pulling ahead from the ones running in place, is what happens to that spark on Monday. Does it hit an engine that carries it forward, a place to be evaluated, funded a little, tested cheaply, and either advanced or retired with a clear conscience? Or does it hit a wall, get a round of applause, and die in a Slack thread?
You already generate the ideas. The only question left is whether you’ve built the engine to do something with them. So what’s the next idea in your loop, and what happens to it when the clapping stops?
Notes
-
Alexander Nolte et al., “What Happens to All These Hackathon Projects? Identifying Factors to Promote Hackathon Project Continuation,” Proceedings of the ACM on Human-Computer Interaction (CSCW), 2020. A companion empirical analysis of roughly 11,900 Major League Hacking hackathon repositories on GitHub found only about 7% showed any commit activity six months after the event. ↩
-
Steve Blank, “Why Companies Do ‘Innovation Theater’ Instead of Actual Innovation,” Harvard Business Review, October 2019. ↩
-
On ambidexterity: Charles O’Reilly and Michael Tushman, “The Ambidextrous Organization,” Harvard Business Review, 2004. On dynamic capabilities: David Teece, “Explicating Dynamic Capabilities: The Nature and Microfoundations of (Sustainable) Enterprise Performance,” Strategic Management Journal, 2007. ↩
-
McKinsey & Company, “The Eight Essentials of Innovation.” A recurring research finding that innovation performance hinges on resource allocation, not idea volume. ↩
-
Metered funding and innovation accounting are developed in Eric Ries, The Lean Startup (2011) and The Startup Way (2017). ↩
-
David Kidder and John Geraci, “To Innovate Like a Startup, Make Decisions Like a Venture Capitalist,” Harvard Business Review, 2018. See also Alexander Osterwalder and the Strategyzer team on funding “explore” projects through metered investment tied to validated evidence. ↩
-
“ING’s Agile Transformation,” McKinsey Quarterly, 2017 (interview with Bart Schlatmann and Peter Jacobs). ↩
-
Clayton Christensen, Stephen Kaufman, and Willy Shih, “Innovation Killers: How Financial Tools Destroy Your Capacity to Do New Things,” Harvard Business Review, 2008. ↩
-
Rita Gunther McGrath and Ian MacMillan, “Discovery-Driven Planning,” Harvard Business Review, July 1995. ↩
-
The two-axis assumptions map is from David Bland and Alexander Osterwalder, Testing Business Ideas (2019). The discipline of small, frequent assumption tests is central to Teresa Torres, Continuous Discovery Habits (2021). ↩
-
Barry M. Staw, “Knee-Deep in the Big Muddy: A Study of Escalating Commitment to a Chosen Course of Action,” Organizational Behavior and Human Performance, 1976. Confirmed and extended by Dustin Sleesman et al., “Cleaning Up the Big Muddy: A Meta-Analytic Review of the Determinants of Escalation of Commitment,” Academy of Management Journal, 2012. ↩
-
Ron Kohavi and Stefan Thomke, “The Surprising Power of Online Experiments,” Harvard Business Review, 2017; and Stefan Thomke, Experimentation Works (2020). The Airbnb and Bing figures are drawn from this body of work. ↩
-
The Three Horizons model comes from Mehrdad Baghai, Stephen Coley, and David White, The Alchemy of Growth (1999). Steve Blank’s caution about its broken time assumption is in Harvard Business Review, 2019. On the payoff of a standing sensing capability, see René Rohrbeck and Menes Etingue Kum, “Corporate Foresight and Its Impact on Firm Performance,” Technological Forecasting and Social Change, 2018. ↩