You're shipping more than ever. Is any of it moving the needle?
You can ship several times the amount of work you could two years ago. Look at SWE-bench, a key benchmark for what agentic systems can do. Two and a half years ago, they could complete 2% of their tasks autonomously. Today it's close to 80%.
That's massive progress. If you've succeeded with agentic coding, I'm sure you're proud of it. I'm sure you've got a slide for your board showing shipping velocity going up.
So here's my question. Out of everything you shipped last year, how much actually moved the number you funded it to move?
Yes, we're shipping faster. But are we impacting business metrics? That's what matters in the end.
Right now, most of us can't answer that, and it's not because we're bad at our jobs. The systems we have in place weren't designed for it. You've got a slide from January, or whenever your planning cycle is, saying this new product line will bring in this much revenue through these strategic initiatives. What you actually shipped, and how it contributed, is often anyone's best guess.
Shipping faster doesn't make you better at product. It means you reach the end of the wrong decision faster. If it's the right decision, great. If it's the wrong one, you have more code to unwind, more usability issues, and everything else that comes with it
The reason is that shipping is one step in a bigger system. Product development is a loop: you decide where to invest, you build, and you learn whether it worked. That loop only moves as fast as its slowest step. Coding agents have sped up the build step dramatically, and the rest of the loop hasn't caught up. So the work for product leaders now is making the whole loop faster, so that every cycle connects what you funded to what you shipped to what it did for the business.
Most of the industry conversation is about coding agents. I want to look at the whole system, where coding agents are one part.
Three phases, six jobs
If you think about the overall product development process, it's really three main phases:
- Planning: Where are we? Where are we strong and weak? How do we allocate resources?
- Building: Specify, build, and deliver.
- Learning: Measure and evaluate. Did it work or not?
This is slightly different from Lean Startup. Build-measure-learn is for discovering a business model. This loop is for running a business model you already have, with a constrained budget and resources.
Funny timing, too. Steve Blank, the father of modern entrepreneurship and the Lean Startup movement, just wrote a post saying the minimum viable product is dead, because now you can ship the product right away. All the learning still has to happen. You just don't need the "minimal" part.
Across those three phases, there are six jobs:
- Planning: (1) diagnose and (2) allocate
- Building: (3) specify and (4) ship
- Learning: (5) measure and(6) listen
Listening feeds diagnosis, the loop closes, and you go again. It's simplified, and life isn't linear, but it works for this purpose.

I spoke to the CPO of a very large company recently. He told me things are great: they've made these agentic advancements and now have an agentic shipping factory. "But now I need to feed the beast. That's my biggest problem."
So what does it take to run the whole loop with agents? Let's go job by job.
1. Diagnose
Where is your product portfolio strong, and where is it weak? Which capabilities deliver real value to customers, and for which segments? Where are you losing deals to competitors, and why?
This is where the product picture and the business picture meet. A capability matters to the business only when you can tie it to revenue, or to whatever your business runs on, like assets under management or market size. Once you can do that, you, your CFO, and your revenue leaders are working from one shared diagnosis instead of competing opinions.
So you want agents that give you a portfolio read organized by the capabilities customers get from you, which is different from your engineering architecture or your epics. That read should show needs by segment, why you lose deals, and the business impact of each. To get there, the harness needs:
- A unified data model. You can't just tell an agent to connect to the MCP for the CRM and the MCP for the support system. You need a structured repository with one consistent interpretation of the data, so you get the same answer every time you ask.
- A way to classify the customer life cycle, because a signal from a prospect means something different from a signal from a customer you lost to churn.
- An engine that rolls up revenue by segment or vertical, and a scoring method that shows what matters most in each.
- An interface where you can slice the data and dig into what's happening.
I looked for purpose-built agentic tools for this job and couldn't find any. Most teams still do it with BI, manual data work, and a lot of spreadsheets.
Once you know where you stand, the next question is where to put your money.
2. Allocate
Which initiatives do you fund? How much capital and capacity goes to each one, and what do you expect each one to change in the business?
Allocation is the decision you'll be asked to defend later. If you tie every initiative to a problem and set its goals before work starts, you can answer both "why did we fund this?" and "did it work?" Most teams can't answer those today. The plan promises revenue from a set of initiatives, and by the end of the year, how the shipped work contributed to that number is anyone's best guess.
Agents can help here by showing capital to impact, tying each initiative to a problem, and sizing resourcing to it. They should also keep a clear log of what you said yes to, what you said no to, and what's out of scope.
If you don't write scope down, the execution process will decide it for you. The harness needs an objective model covering initiatives, objectives, and goals, along with prioritization scoring, resource allocation, and an interface to interpret all of it.
This is a more established market, with tools like Jira Align, Planview, and ServiceNow SPM, and the plumbing underneath is complex.
A funded initiative is still only an intent. The next job turns it into something a team or an agent can build.
3. Specify
What exactly are you building, for which segment, and what does done look like? What does the customer evidence say, and what in the codebase can you reuse?
The spec is the contract between your intent and whoever builds it, person or agent. Agents build what the spec says. If the spec isn't clear, you get the wrong feature, and with agents you get it faster. The better the spec, the better your odds of building the right thing.
That's why this harness carries more requirements than most:
- Assembled context: the customer signal distilled for this specific feature, and ideally a searchable codebase.
- Business context: your pricing, packaging, and product principles, plus guardrails like regulatory compliance.
- A layer that translates the codebase into terms PMs use. Agents are now very good at mapping a codebase, but their reading is often still technical.
- A use case writer, with examples and evals so the output stays consistent. Use cases should be complete user journeys, so you deliver the full value instead of pieces of a solution.
- An acceptance criteria generator, so agents or people can QA the result against the requirements.
- A human review gate that shows what someone actually approved.
That last one matters more than it sounds. People generate a spec, don't take the time to read it, and say, "Okay, sounds good, off you go." Nobody reads page five. You need a process that separates what a person approved from what's just AI slop. If you work in a regulated industry, you already know this. Some of you can't build anything until five or ten people, sometimes including agencies, have reviewed it.
New tools are showing up here, including GitHub Spec Kit, Anthropic's Outcomes beta, and AWS Kiro. The acceptance criteria you write at this stage are what you'll check against after you ship.
4. Ship
Is the code working, tested, and observable at launch? And do the people selling and supporting it know what shipped, who it's for, and why it matters?
You probably won't build a coding harness yourself. The market exploded here, with Claude Code, Cursor, GitHub Copilot, Codex, Devin, Antigravity, and Amp, plus Linear and Jira managing the work.
Where your harness needs work is everything around the code. Count a feature as shipped only when it's release-ready: working, tested code with observability at launch, plus documentation, release notes, enablement for the field, and the product marketing machinery behind it. Plan that hand-off as part of shipping, not after it.
And enablement has truly become a bottleneck. When shipping velocity is this high, the field can't keep up.
At Productboard, we shipped 47 features in the past 30 days, and our go-to-market team's reaction was, "Oh my god." Two of my former coworkers, one now at OpenAI and one at Anthropic, say enablement of the field is a huge issue for them too.
If sales, support, and customers don't know what you shipped, who it's for, and why it's valuable, the feature won’t deliver value no matter how fast you built it because it’ll be hard to explain it to customers.
And usage is a key part of this loop.
5. Measure
Is it being used? What does engagement look like, and did it hit the goal you set when you funded it?
Usage is often the easiest proxy you have for value delivered, and it connects what you shipped back to what you funded. It also changes behavior before launch. Once PMs know their feature's usage will be visible, they start asking early whether it will drive engagement and whether they can defend it.
At Productboard, we built a leaderboard, like a sales leaderboard. It lists every feature, the PM who owns it, adoption through internal testing, beta, and GA, and engagement. Every week we put it up and publicly praise the features people use and call out the ones they don't. The point isn't to shame anyone, but it's a great forcing function.
Features aren't apples-to-apples, though. A settings feature you use once can't be compared with one meant for daily use, so group features into categories before comparing them.
Tools like Amplitude and Mixpanel already have built-in agents. The gap your harness has to close is the level of detail. These tools hold event-level data, and you need to report at a feature or capability level everyone understands, possibly by calling these tools' MCPs.
The numbers tell you what happened. They don't tell you why.
6. Listen
How are customers reacting to what you shipped? Which product gaps are blocking deals, and how much revenue would closing them unblock? What do healthy customers want more of, and what's pushing unhappy ones toward contraction or churn? What are competitors and early upstarts shipping?
Measurement is the quantitative side. Listening is the qualitative side that explains it. When you tie feedback to customers and revenue, it also tells diagnosis how much each problem is worth.
That only works if the feedback is consistent. If each PM drops feedback into an agent ad hoc, the agent interprets it differently every time, and you can't track how a complaint or request trends over time. The harness needs:
- Connectors to every feedback source, using direct integrations or APIs where MCPs aren't mature yet.
- A feedback pipeline that sorts feedback into consistent topics or themes, with evals.
- A data model that ties feedback to customers and revenue.
Because copying is so easy now, everyone is watching competitors more closely, so invest in competitive intelligence.
Side note: New agents are popping up here, like ListenLabs, doing synthetic persona and synthetic user research. That gives you signal from a synthetic representation of the voice of the market, on top of your existing customer conversations. People have historically frowned on this category. In the last six months or so, it seems to have really exploded.
Listening feeds diagnosis, and the loop starts again. The faster you get around it, the faster you learn.
The hub: Orchestrating the loop
The six jobs are the spokes of a wheel. The last piece is the hub that orchestrates them.
You need an agent you can ask: What's in what status? What's working? Who's working on what, across the entire loop?
The outcome is a connected system that holds work, features, feedback, and revenue together and is accessible to everyone. Anyone can ask why something is being worked on, or why something is blocked, and get an answer. And as agents become more autonomous, you need gates where humans approve.

General-purpose agents like Claude, OpenAI Codex, and Glean live here, indexing and interpreting data. They're not specialized for this job, but they provide the interaction layer and can connect to a lot of tools through MCP.
In the Q&A, someone raised a related problem: everything is becoming a markdown file. There's a spec for the spec for the spec, all sitting on someone's local machine where no one else can see it.
Honestly, I don't have a full answer. You need to be very thoughtful about what you save in structured systems. For the markdown files, you need the same discipline you apply to your codebase: a repository, access controls, permissions, and clear ownership of who put each file there and whether it's current.
Dashboards have the same problem. It used to be hard to get a visualization. Now people share 100 different visualizations of the same data with me, and 90 of them are junk. They're not connected, and the people making them don't understand information design. It's not as easy as every PM building their own skills and agents. You need the system.
Three ways to get there
It's a lot to build. I know it's overwhelming, and that's kind of my point. As an industry, we have a lot of work ahead of us. For the whole loop to move faster, all six jobs have to work in unison. Otherwise, there will always be a bottleneck.
A lot of larger companies plan once a year. That's a very long cycle when startups are running on a daily or monthly basis and learning super fast. Find a way to shorten your planning cycle, at least on the product side, if not across the whole business.
There are three ways to get there.
Build it yourself. I'm talking to companies building the entire infrastructure themselves, including CRM and HR systems. "Hey, screw it, we're going to build the entire agentic future ourselves." I'm seeing it especially in private equity portfolios, where firms want to build one universal agentic infrastructure for every company they own. It's a very big investment, but it might make sense.
Assemble it. If you go this route, make sure you have a unifying system where the data is connected. That means a central, structured data store that attributes revenue to customers, feedback, and work items.
Or, a little pitch: work with us. This is what we're working on with our agentic product system, Productboard Spark. It's not fully built out yet, but that's the mission. Our ambition isn't to rebuild Amplitude and all the other systems. We want to build the optimized harness that knows what each source provides, helps you make sense of it, and orchestrates the whole wheel. It's a journey for us too, because the industry is moving so fast.
Loop speed is the advantage. See you out there. Or your agents.
