Skip to content
Back to Blog
July 25, 2026

Measuring User Satisfaction in AI Products

NPS misses the point for AI agents. Task completion, return usage, and conversation-length tell you whether your AI travel agent actually works.

Measuring User Satisfaction in AI Products
M

Someone on my team once asked whether we should add an NPS survey to our AI travel agent. My answer was no, and the reason reveals something important about how AI products should be measured.

NPS (Net Promoter Score) asks "how likely are you to recommend this product?" It is a fine metric for products where satisfaction is the primary concern. But for an AI agent that handles real tasks with real outcomes, the right question is not "did you enjoy this?" but "did this work?"

A user can have a pleasant conversation with an AI agent that completely fails to find them a reasonable flight. They might rate the experience highly because the agent was friendly and responsive. Meanwhile, their actual need remains unmet. Conversely, a user might have a terse, efficient interaction that books the perfect trip in three minutes and rate it neutrally because there was nothing remarkable about the experience. But they will come back for every trip after that.

Satisfaction metrics for AI products need to measure outcomes, not feelings.

Task completion as the primary signal

Illustration for this section

The north star metric for our AI travel agent is task completion rate: the percentage of conversations where the user accomplished what they came to do.

This is harder to measure than it sounds. Not every conversation has a clear goal. A user might open the app to browse destinations with no intention of booking. Another might start planning a trip they will not book for months. A third might ask a quick question about their existing booking.

We classify conversations by inferred intent and measure completion within each category:

  • Active booking intent: user is trying to book something. Completion = booking confirmed.
  • Research intent: user is exploring options. Completion = user received useful information and did not express frustration.
  • Management intent: user is checking or modifying an existing booking. Completion = information retrieved or modification executed.
  • Exploratory intent: user is browsing without clear intent. Completion = conversation lasted more than two turns (engagement threshold).

Each category has its own completion benchmark. Measuring them separately prevents the high completion rate of simple management tasks from masking poor performance on complex booking tasks.

Conversation length as efficiency proxy

In most cases, shorter conversations are better conversations. A user who books a flight in four turns had a more efficient experience than one who needed twelve turns for the same outcome.

But length alone is deceptive. A four-turn conversation that ended because the user gave up is not efficient; it is failed. A twelve-turn conversation where the user explored five destinations before falling in love with one is not inefficient; it is discovery.

We measure conversation length relative to task complexity. A simple one-way flight booking should complete in 3-5 turns. A multi-city trip for a group might legitimately take 15-20 turns. If a simple booking consistently takes 10 turns, our clarification or search quality needs work. If a complex trip takes 5 turns, the user might have felt rushed.

The metric we actually watch is "excess turns": turns beyond the expected minimum for a given task type. High excess turns on simple tasks indicate friction in the agent's conversation design.

Return usage: the ultimate signal

Supporting diagram

Here is the metric I care about most, even though it has the longest feedback loop: does the user come back for their next trip?

People do not book travel frequently. The average leisure traveler books 2-4 trips per year. That means the return usage signal takes months to appear. But when it does, it is the most reliable indicator of genuine satisfaction.

A user who rebooks within 90 days is a satisfied user. Full stop. They had other options. They chose to come back. That is a stronger signal than any survey rating.

We track 90-day return rate as our long-term quality metric. It correlates strongly with task completion rate, which makes sense: users who successfully booked a trip are the ones who come back. The correlation coefficient between task completion rate and 90-day return rate is approximately 0.82 in our data. That is high enough to use task completion as a leading indicator for the lagging return metric.

Explicit feedback within conversations

We do collect explicit feedback, but we integrate it into the conversation flow rather than popping up a survey modal.

After a booking is confirmed, the agent asks: "How did that go? Anything I could do better next time?" This feels natural in a conversation. It does not interrupt the experience. And the responses are far more useful than a 1-5 star rating because users provide specific, contextual feedback.

"The hotel options were great but I wish you had shown me ones closer to the train station." "The booking was smooth but I was confused when you asked me to confirm twice." "Perfect, nothing to improve."

These qualitative responses are gold for product improvement. They identify specific issues that quantitative metrics miss. A drop in recommendation acceptance rate tells you something is wrong. A user comment about train station proximity tells you exactly what.

Triangulating satisfaction

No single metric tells the full story. We triangulate across three data types:

Behavioral signals: task completion rate, conversation length, recommendation acceptance rate, booking rate. These are the most reliable because they measure what users do, not what they say.

Explicit signals: in-conversation feedback, post-trip ratings, support ticket sentiment. These add qualitative depth and specific issue identification.

Outcome signals: return usage rate, booking value trends, referral behavior. These are the ultimate measure of product-market fit but have long feedback loops.

When all three signal types agree (high task completion, positive feedback, strong return usage), we know the product is working. When they diverge, the divergence itself is informative. High task completion with low return usage might indicate that the booking worked but the trip itself was disappointing (recommendation quality issue). Positive feedback with low booking rate might indicate the conversation is engaging but the booking flow has friction.

The key insight from two years of measuring AI travel agent quality: satisfaction in AI products is an outcome, not a feeling. Measure outcomes first. Feelings follow.


Nowah is an AI travel agent that searches and books real flights and hotels through conversation — no filters, no thirty open tabs. Plan your next trip.

Share this article

Ready to Plan with Nowah?

Bring the idea. Nowah will help turn it into a trip.

Try Nowah