Summarize with AI:
Everyone’s running experiments. Most are just making changes. The difference is one ingredient—and every serious framework already has it.
We’re all running one big experiment right now.
Organizations are placing bets they can’t fully calculate. And somewhere in between, developers and designers are shipping agents, redesigning workflows and implementing systems at a speed that would have been unthinkable five years ago.
We’re all wearing lab coats.
The problem is most of us aren’t actually running experiments. We’re changing things. Replacing processes. Spinning up agents. We’re just playing around. Which feels like we did a thing.
Feels like progress.
But change is not the same as experimentation. And in the AI era, that distinction has never mattered more.
You may have seen this on supermarket shelves. A product you’ve bought for years as a loyal customer suddenly has a badge on the packaging: New Formula. Sounds like a good thing. So you hope for better ingredients and improved results.
But when you try it, it tastes the same. Maybe even worse.
What actually changed was the profit margin. The company found a way to alter the formula so it benefits the bottom line. Cheaper sourcing, different ingredients—same packaging, just with a new badge. They optimized internally and called it innovation. But they didn’t test whether the outcome improved for you. Most of the time, they tested whether you’d notice.
The gesture of value replaced the value itself.
This happens in software constantly. A team adopts AI tooling, automates a handful of tasks, watches output increase and declares the experiment a success. But often without a definition of success before starting. No baseline set. No asking whether the delta—the measurable difference between before and after—justified the investment.
The feeling of improvement replaced the proof of it.
The formula for running a real experiment isn’t new. It’s not waiting to be invented. Every serious framework that exists already contains it.
PDCA: Plan, Do, Check, Act. The scientific method: hypothesize, test, measure, conclude. Lean Startup: Build, Measure, Learn. Strategyzer’s experiment approach from Testing Business Ideas.
Every single one has the same ingredient baked in with different names—the same thing: a solid test.
Strategyzer builds it into what they call the Test Card. A simple tool that forces four things explicit before any experiment runs: what you believe, how you’ll test it, what you’ll measure and, most importantly: “We are right if ...”
That last field is the one that gets skipped.
“We are right if ...” is written before the experiment starts. It defines the delta upfront—the specific, measurable difference that would tell you whether the change actually worked. Not whether something changed. Whether it improved. For whom. By how much.
A simple but powerful example: before you deploy a new AI agent into your support workflow, the Test Card asks you to complete the sentence. We are right if response time drops by 20% without increasing error rate. We are right if customer satisfaction scores hold or improve. We are right if the cost per resolved ticket decreases.
That’s the ingredient. Without it, you’re not running an experiment. You’re making a change and hoping it becomes progress.
Science can afford to be wrong. A negative result is still a result. It advances knowledge. The experiment that disproves the hypothesis is just as valuable as the one that confirms it. Being wrong is not the same as failure.
Business doesn’t work that way. Being wrong has a direct consequence. Your reputation or career affected. A budget or a project cut. There is a pressure to show results and when the cost of being wrong is high enough, people will always find a shortcut. Teams stop running experiments and start building the appearance of them. They find the metric that confirms it worked, declare success because usage went up or retrofit the proof.
It’s not dishonesty. It’s success obsession without enough depth.
In science, the evidence has to support the hypothesis. In business, success is often measured by the bottom line result. If the result looks good enough, the quality of the evidence doesn’t get questioned. That is where experiments become changes with a good story.
The “We are right if ...” field on the Test Card is designed to close that gap. It forces success criteria to be set before anyone knows the result. It makes retrofitting impossible. And it requires something business culture rarely demands: the genuine willingness to find out you were wrong.
In the AI era, that willingness is more expensive than ever. You’ve shipped the agent. The team is invested. Leadership is running around with slide decks about the initiative. The last thing anyone wants is an experiment that comes back negative.
But that’s exactly why you need one.
Law firms are a useful case. In 2024, legal professionals predicted AI would significantly improve efficiency, raise quality and transform how they bill. Technology spending rose 9.7% in 2025—the fastest growth the industry had ever seen.
But nearly 60% of in-house counsel reported they had seen no noticeable savings from outside counsel using AI.
Usage went up. Client improvement did not.
Kinda like a new formula in the supermarket.
Here’s what actually happened. Firms optimized for efficiency inside a billing model that punishes efficiency. They built tools that compress work from hours to minutes—then stayed trapped in billable hour structures that reward time consumption, not outcomes. They measured what changed internally. They never asked we are right if from the client’s perspective.
Legal AI vendors moved from flat-seat pricing to consumption-based models, a shift that also requires attention to data privacy and compliance obligations alongside cost implications. The costs that had been subsidized to get firms hooked started landing on the balance sheet. And firms had no evidence base to fall back on—because they had never built one.
They had tested whether clients would notice. Not whether quality improved. Just like the supermarket. The moment costs rose, the margin argument collapsed. And without a quality story, there was nothing left to defend.
According to RAND Corporation research, more than 80% of AI projects fail to reach meaningful production—roughly twice the failure rate of traditional IT projects. The most common reasons: the wrong problem being solved, poor data foundations and a tech-first focus that chases tools instead of outcomes. In other words, experiments without a clear hypothesis and measurement.
By the time you realize you needed that baseline, you’ve been spending for months with nothing to compare against.
Before you deploy the next agent, redesign the next workflow or ship the next feature—stop and complete one sentence.
We are right if…
If you can’t finish it, you’re not ready to run the experiment. You’re ready to make a change. And that change might look like progress, feel like progress—but get you nothing.
In a world where building has never been faster, measuring has never mattered more.
The formula is not new. It’s just the ingredient most teams skip.
“We are right if ...”
And I’m ready to be wrong.
Teon Beijl is a business designer with over a decade of experience in enterprise software for the oil and gas industry.
Formerly Global Design Lead for reservoir modeling, remote operations and optimization software at Baker Hughes, he now helps people who feel stuck through his own business, Unpuzzler. Teon works with leaders on business design and with professionals on career design, leveraging his experience as both designer and leader to help people create clarity and live on purpose—by design. Connect with Teon on LinkedIn or Substack.