The Sound of AI: Designing Audio Feedback for Agent Interactions
Booking confirmation should sound celebratory. Typing indicators need subtle presence. Voice output timing matters. Sound design shapes AI personality.

Close your eyes and imagine booking a flight. What does it sound like?
If you are on a traditional travel website, the answer is: nothing. Maybe a click sound from your mouse. Maybe the hum of your laptop. The actual booking experience is silent. You fill out forms, click buttons, and see a confirmation screen. No audio. No ceremony. A $3,000 purchase and a $3 coffee have the same sonic experience: silence.
This feels like a missed opportunity. Sound shapes how we feel about digital interactions. The macOS trash can crumple. The iPhone lock click. The Slack "knock brush" notification. These sounds are so embedded in our experience that imagining those products without them feels wrong. They carry emotional information that pixels alone cannot.
AI interactions, where conversations happen and transactions complete and agents work on your behalf, deserve the same level of audio design attention. We have been thinking about this at Nowah, and the territory is more interesting than I expected.
Sound creates personality

Every AI product has a personality, whether the team designed one or not. The speed of responses, the word choices, the level of formality, all of these create a character in the user's mind. Sound adds another dimension to that personality.
Consider the moment when a booking confirms. Visually, you see a success screen with a checkmark and booking details. That is informative. Now add a sound. A clean, upward-resolving chime, maybe two notes ascending with a warm harmonic quality. Suddenly the moment has emotion. It feels celebratory. The sound says "something good just happened" in a way that a green checkmark alone does not.
The peak-end rule from psychology tells us that people remember experiences primarily by their emotional peak and their ending. A booking confirmation is often both the peak and the end. If that moment has a satisfying sound, the entire booking experience retroactively feels better. Kahneman would love it.
Apple understood this decades ago. The Mac startup chime was not functional. It was emotional. It said "welcome, something is beginning." The iPhone keyboard clicks were not necessary (the phone has no mechanical keys). They provided tactile feedback that made typing feel responsive and real.
Duolingo is the modern master of sound design in a learning app. The lesson completion sound is a small burst of joy. The streak sound creates a Pavlovian motivation loop. The error sound is gentle enough to avoid discouragement while clearly signaling a mistake. Every sound in Duolingo serves an emotional purpose beyond information delivery.
For an AI travel agent, the sound palette needs to convey a specific personality: competent, warm, attentive, and celebratory at the right moments. Not robotic. Not overly playful. The sound of a trusted agent who takes your travel seriously but also wants you to feel excited about your trip.
The five moments that deserve audio
Not every interaction needs a sound. Sound overuse creates noise, and noise creates annoyance. We identified five specific interaction moments where audio adds genuine value.
Thinking. When the AI is processing a complex request, a subtle ambient sound signals activity. Not a spinner's silence, not a loud animation, just a quiet indication that work is happening. Think of it like hearing someone typing on the other end of a phone call. You know they are doing something, and the absence of silence feels reassuring. This should be soft, low-frequency, almost felt more than heard. A gentle pulse.
Searching. When the AI triggers a flight or hotel search, the sound shifts slightly. Still subtle, but with a rhythmic quality that suggests progress. Like a sonar ping scanning outward. This differs from the thinking sound because the agent is now actively querying external systems, and the user should sense that something more active is happening.
Presenting. When flight or hotel cards appear in the conversation, a crisp, clean sound accompanies the reveal. This is the "here are your options" moment. The sound should be clear and attention-getting without being startling. A single clean tone that says "look at this."
Confirming. The booking confirmation sound is the most important. It needs to feel earned and celebratory. Two or three ascending tones with warmth. Not a fanfare (too much). Not a ding (too little). Something that says "you just made a great decision, and now something exciting is going to happen." This is the sound the user should associate with Nowah in their memory.
Alerting. Notifications about trip changes, flight updates, or price drops need a distinct sound. It should convey urgency without anxiety. Informative without alarming. A firm but warm attention-getter. The user should hear it and think "I should check this" not "something bad happened."
These five sounds, used consistently, create an audio identity for the AI agent. Over time, users recognize them the way they recognize the Slack notification sound or the iMessage sent swoosh. The sounds become part of the product's personality in the user's mind.
Voice output: when the AI should speak vs. type
Nowah supports voice input. The question of voice output is trickier and more nuanced.
The default assumption might be that if the user speaks to the AI, the AI should speak back. But our research and testing suggests this is not always right.
Voice output works well for short, confirming responses. "I found three flights to Paris. Here are your options." The user hears a brief verbal summary while looking at the flight cards. Voice adds a layer of engagement without replacing the visual information.
Voice output works poorly for detailed, data-heavy responses. Nobody wants an AI to read aloud the departure time, arrival time, airline, number of stops, layover duration, and price for three flights. That information is much faster to process visually. Hearing it spoken is slow and hard to compare across options.
The right framework is context-dependent:
- Summarize verbally, detail visually. The agent says "I found three options ranging from $450 to $620" while displaying the full cards for visual comparison.
- Speak when hands are busy. If the user is packing and asked about their flight time, a spoken answer is more useful than a displayed one.
- Be quiet when the user is reading. If the user is reviewing flight cards in detail, additional voice output is distracting.
- Match the user's modality. If the user typed their question, respond in text. If they spoke, respond with voice plus text.
We have not perfected this yet. The timing of voice output relative to card display, the right balance of verbal summary and visual detail, the detection of when voice is helpful versus intrusive, these are active design problems. But getting them right is part of creating an AI agent that feels natural rather than mechanical.
Accessibility: sound as complement, never sole channel
Any discussion of sound design has to address accessibility. Approximately 1.5 billion people worldwide have some degree of hearing loss. Sound design that carries important information exclusively through audio fails these users.
Our principle is that sound is always complementary. Every piece of information conveyed by sound is also conveyed visually. The booking confirmation has a visual success state regardless of whether sound is enabled. The thinking indicator has a visual animation. Notifications have visual badges.
Sound adds emotional texture and peripheral awareness. It should never be the only channel for information. A user with sound disabled should have the exact same functional experience, just without the emotional audio layer.
This also means all sounds must be controllable. Global mute. Per-category sound controls. Vibration alternatives for haptic feedback on mobile. The user should have full control over what they hear, when they hear it, and whether they hear anything at all.
Building an audio identity
The collection of sounds in a product forms an audio brand. Apple has one. Duolingo has one. Slack has one. Most travel apps do not.
We think there is an opportunity to build a distinctive audio identity for Nowah that users associate with the experience of AI-assisted travel booking. A sound that, when heard, evokes the feeling of having a competent agent handle your travel. Like a jingle but functional. Like a brand sound but earned through repeated positive associations rather than marketing repetition.
The booking confirmation sound is the anchor. It is the sound users hear at the most emotionally positive moment in the product. If that sound is distinctive and satisfying, it becomes the audio signature of the entire experience. Everything else, the thinking pulse, the search rhythm, the notification tone, orbits around that central moment.
We are still early in this work. Sound design for AI interactions is a young discipline with few established best practices. But we believe the products that invest in it will create deeper emotional connections with users, and in a category like travel where decisions are emotional and memorable, that connection matters.
The best AI travel agent should not just feel smart. It should sound right.
Nowah is an AI travel agent that searches and books real flights and hotels through conversation — no filters, no thirty open tabs. Plan your next trip.