> ## Content Index
> Fetch the complete content index at: https://www.productledalliance.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# You can't prove AI value with metrics that don't  ask a user: Q&A with Keith Zubchevich, CEO of Conviva
- URL: https://www.productledalliance.com/you-cant-prove-ai-value-with-metrics-that-dont-ask-a-user-q-a-with-keith-zubchevich-ceo-of-conviva/
- Published: 2026-10-01T09:31:42.000Z
- Updated: 2026-10-01T09:31:41.000Z
- Description: Discover why AI agents can pass every check and still lose customers, and how product leaders should measure agent quality.
- Author: Keith Zubchevich
- Tags: AI & Machine Learning, AI, Analytics & Metrics, Articles, #conviva

You're trying to get a refund on a $1,000 couch. You explain the problem to a support agent. Then you explain it again. And again. After seven exchanges, you finally get your refund.

On paper, that's a win. The ticket's resolved, and the agent did its job, but you're probably not coming back.

[Keith Zubchevich](https://www.linkedin.com/in/keithzubchevich/), President and CEO of [Conviva](https://conviva.ai/), calls this "the grind that still finishes." In his view, it's one of the hardest agent failures to spot, because every metric you're tracking says it went well.

Following his keynote at the [Virtual AI-Native Product Summit](https://virtual.productledalliance.com/location/ainative/), product leaders asked Keith how to measure what AI agents actually cost users, and how to catch problems before customers quietly stop showing up.

Here’s what Keith had to say:

## 1\. How do you measure the quality of agentic AI solutions?

Right now, I’ve found most teams are meeting only the bare minimum with basic quality-of-service and quality-of-response metrics. You can’t truly understand whether an [AI agent](https://www.productledalliance.com/building-ai-agents-where-and-how-to-start/) worked well for the user if all you know is (1) the system responded, and (2) it was friendly and factual.

Instead, measure the agent by what it costs the person to get through it. Did they have to repeat themselves? Backtrack? Did they settle for less than what they actually wanted? That's one layer of the problem. 

The second layer is what they did after. Did they come back? Did they go around your agent and just do it themselves? Or did they quietly never return? 

None of these outcomes show up on a system dashboard. 

Whatever's measuring performance can't be another model grading the output. It has to be computed; the same input, the same answer, every time. Otherwise you're not measuring your agent; you're measuring your judge's mood that day.

## 2\. How can personal agents be tracked and addressed as individual customers?

It’s my view that every conversation should be its own thread from start to finish, not chopped into turns and averaged. That's the only way you catch everything that happened to one person. 

Threads highlight if: 

- A user reformulated a request
- The agent looped on the same tool call
- The user gave up halfway through

None of these signals exist if you're scoring turns in isolation. 

Once you have this data, you need to tag the thread with everything around it to understand: what kind of request it was, what surface it happened on, and where it broke down. 

This process allows you to pull that one person's bad experience out of the pile, instead of letting it get smoothed into an average. And you have to do it for everyone, not a sample. The failure that matters is narrow, and sampling is how you miss it.

## 3\. What happens when an AI agent gives the "right" answer, but the user's behavior shows that it didn't actually solve their problem?

That's the whole ballgame.

Nobody wants to talk about that case. When the agent passed every check you had: it was correct, on brand, and fast. But your customer left anyway. 

Reaching the outcome and losing the customer aren’t opposites; they happen together all the time. So if your only signal is "did it complete," you will never see this failure. 

Instead, you'll see a resolved ticket and a lost customer, and your dashboard will tell you everything's fine.

[AI agents: 5 lessons for getting it rightAI agents are here, but 40% of projects fail. Learn five lessons for PMs to get it right. Discover how to launch successful AI agent projects and lead the next wave of automation.![](https://storage.ghost.io/c/06/40/064053a7-301f-45ab-974f-9f6baeaffcc5/content/images/icon/android-chrome-192x192--1--3c813a8f-3376-41af-8d85-c05fbae651bc.png)Product-Led Alliance | Product-Led GrowthFarah Ayadi![](https://storage.ghost.io/c/06/40/064053a7-301f-45ab-974f-9f6baeaffcc5/content/images/thumbnail/social---AI-agents-5-lessons-for-getting-it-right-dfdb5aee-50e1-4a2b-8bb0-e40dab6c0725.png)](https://www.productledalliance.com/ai-agents-lessons-for-getting-it-right/)

## 4\. What’s the definition of success against which the agent or LLM is measuring itself? Can this be improved with a simple Like/Don't Like option the user can use to indicate whether the agent is providing a quality result?

Right now, most agents are grading their own homework. Success means "I produced a response that clears the bar I was given." Not: "I solved this person's actual problem." 

A thumbs-up button doesn't fix that. It's one click, on one turn, from a fraction of your users, and it tells you nothing about the five exchanges that led up to it or what your user did the moment they closed the chat. 

Instead of relying on a user having to manually tell you if they’re happy or not via a thumbs up/down, consider looking at implicit signals like:

- How often did the user have to repeat their question?
- Did the user check the agent’s work?
- Did the user get what they came for?

## 5\. What is the hardest type of agent failure to detect through user behavior, especially when the user ultimately achieves their goal?

I call this the grind that still finishes. 

If somebody is already committed (like in our refund example), they might push through several exchanges of restating and backtracking until they get to their desired outcome. 

That likely reads as a win everywhere you'd look, because ultimately they got what they came for. But for the user, it isn't even close to a win. 

This failure just spent goodwill and trust you won't get back. And worse, you won't find out until they stop showing up. 

That's the case that hides this dilemma best, because the scoreboard and the truth are pointing in opposite directions.

## 6\. Is there a tool specialized in agentic experience metrics?

Yes, that's what we’re solving at Conviva. We spent twenty years doing this for video, moving the industry away from "did the stream start" and towards "did the viewer have a good experience." We did that at full census, in real time, using computed metrics, not samples or approximations. 

Agents require that same discipline. Conviva reads what the user came for, measures both the effort it took and what they did afterward, across every conversation. 

Your agent can pass every evaluation you throw at it. Conviva is how you discover whether it worked for the person on the other end.

## Summary 

For Keith, the biggest mistake teams make with AI agents is measuring the wrong thing. 

Response times, accuracy, and tone all matter, but they only tell you whether the system performed. They don't tell you what the experience cost the person on the other end, or what that person did next.

Getting that picture means changing how you measure. Keith argues for computed metrics that give the same answer for the same input every time, rather than another model grading the output. 

It also means following every conversation from start to finish, not scoring turns in isolation or relying on a sample. That's where signals like repeated questions, looping tool calls, and abandoned conversations show up.

Because, as the couch refund shows, the most damaging failures often look like wins. If you're only tracking whether your agent finished the job, or waiting for a thumbs-down, you might not discover the issue until your users stop showing up.

Learn more about how [Conviva](http://conviva.ai) can help.