UXR LEAD · META WEARABLES AI · JANUARY–AUGUST 2025

Bringing live translation capabilities to wearables

A 2-phase mixed methods approach with surveys, interviews, and log analyses

Mark Zuckerberg and mixed martial artist Brandon Moreno, both wearing Ray-Ban Meta glasses, laugh on stage while a screen behind them shows their conversation translated between English and Spanish.
Live translation demoed on stage with mixed martial artist Brandon Moreno at Meta Connect, September 25, 2024. Frame from Meta’s keynote video (29:07).
01
CONTEXT

Live Translation shipped in two steps.

What Live Translation is. Live Translation capabilities on Ray-Ban Meta glasses enable users to hold live conversations with people in languages other than their primary language. At launch it supported English (EN), French (FR), Italian (IT), and Spanish (ES) bi-directional language pairs; voice invocation (“Hey Meta, start live translation”); and visual transcripts within the companion app, which transcribes the conversation in real-time, using both the audio from the conversing partner and the glasses wearer, while the glasses wearer hears the audio translation on their glasses.

The two-stage rollout. Live Translation shipped in two steps: a Beta release to an external Early Access Program in late December 2024, and a general-audience launch in late April 2025. I ran a survey + follow-up interview study at each stage, with the same core questions, so the second study could be read against the first.

Stage 1 — Early Access ProgramStage 2 — Post-GA
FieldedJanuary – February 2025June – July 2025 (readout August 2025)
AudienceExternal Ray-Ban Meta Glasses users in the Early Access Program: US & Canada based, mostly male, mostly iPhone, mostly affluentGeneral public in all launched markets; ~90%+ North America
Surveyn = 44n = 377 completes (in-app, Meta AI app Device Tab)
Interviewsn = 10, 60-minuten = 10, 60-minute, remote over Zoom
Core questionWhat are the risks and opportunities with launching this feature for April 2025? We’ve done studies with live translation in controlled indoor settings (e.g., Oct. 2024), but what happens in the wild?Past UXR involved only early adopters. We’d love to understand how the general audience uses this feature.

Both studies were conducted to: understand live translation use cases in the wild (both for 2-way conversations with other people and listening to a 1-way monologue or video); understand the qualitative reasons driving satisfaction ratings; and identify issues and areas for improvement.

A Live Translation onboarding screen in the Meta AI app: a woman in Ray-Ban Meta glasses smiles at a conversation partner, above the text 'View translations in the app — in a conversation, your response is translated for your partner to read in the app. All translation happens on your device, so your conversation stays private.' and a Continue button.
Live Translation onboarding in the Meta AI app: the wearer hears the translation on the glasses; the partner reads theirs in the app.
02
STAGE 1 · EARLY ACCESS PROGRAM · JAN–APR 2025

Method for Early Access Program

Phase 1 – Early Access Program survey (n = 44): Respondents can choose any of the following activities: having a conversation with a friend / family member / ChatGPT, or watching any chosen video in Spanish, French, or Italian. Afterwards, they fill in a 15-minute survey on perceived accuracy, latency, overall satisfaction, and more.

Phase 2 – Qualitative interviews with Early Access Program survey respondents (n = 10): 60-minute qualitative interviews were conducted to follow up on their use cases, what worked well, and what didn’t. Mix of satisfaction levels represented; real world usage prioritized when applicable.

03
STAGE 1 · EARLY ACCESS PROGRAM · JAN–APR 2025

Use cases in the wild

While traveling in faraway places: One Early Access Program participant connects with shopkeepers, Airbnb hosts, and locals in Guatemala & Mexico while traveling. His conversations while traveling tend to be around ~5 minutes. He also used Live Translation to listen to an entire 30-minute bat cave tour in Spanish. He does not speak Spanish. As a side perk, his friend can also enjoy the visual translations available on the mobile phone while the audio is playing on the glasses.

While working: Another participant is a family physician in Canada who oversees patients who are not fluent in English. He uses live translation to understand French-speaking patients, and to confirm the accuracy of state-provided human translators – for patient interviews over the course of 1-2 hours, or until the battery runs out.

While connecting with neighbors: A third participant lists the languages he hears from fellow residents in shared accessibility housing – some supported by glasses, many not (yet). He’s legally blind, and these capabilities give him a new way to connect with others.

Early live translation usage is often colloquial & done in low-stakes environments with friends & family. Testing with slang and “naughty words” created a bonding experience and created trust in the accuracy of the live translation tool.

Where it falters. 1-way translation falters when the audio source speaks very fast (and you don’t have a way of adjusting it), or if you’re listening to outdoor announcement systems where audio quality can be fuzzy. Participating or listening in group discussions with mixed language abilities can be trickier for the live translation feature to follow, but the desire to use it for group scenarios remains. And no philosophy talks or usage in high-stakes legal discussions yet to report: when that same participant got stopped by police in a foreign country, he pulled out Google Translate instead of using the glasses. He did not want to bother with explaining the smart glasses feature to a (potentially uncooperative) conversation partner experiencing it for the first time.

Data science partners show that most sessions were over within several minutes. These concentrated bursts of activity match experiences shared within the qualitative interviews.

04
STAGE 1 · EARLY ACCESS PROGRAM · JAN–APR 2025

Live Translation at a glance

Early Access Program users are satisfied overall with live translation. Accuracy and companion app experiences are strong points. “Pauses” at the start and middle of conversational turns make live translation challenging for early users.

MetricCSAT (n=44)% Dissatisfied
Overall SatisfactionX.XX.X%
Accuracy of TranslationX.XX.X%
Satisfaction with Visual TranscriptsX.XX.X%
Product Education within Meta AI appX.XX.X%
Naturalness of TranslationX.XXX.X%
Wait time required for hearing audio translationX.XXX.X%

“I was in a group setting speaking to one person demonstrating the capability of the glasses. The speaker was the father of my daughter in law. He is a native Cuban and I have always struggled to understand him but this made a HUGE difference.” – P17

But live translation comes with a steep learning curve. The wait time required for hearing audio translations is a common stumbling point, with early users handling both 1) this new tech and 2) a language they’re not fluent in.

“The pauses take me off guard. It would be nice as the speaker is speaking, to automatically continuously keep hearing the translation. Many times it paused and I felt lost, had to refer to the text translation.” – P31

Takeaway: How we handle the “pauses” at the start of conversation turns matters more than incremental gains in accuracy at this point – at least for the language pairs we currently have available {ES, IT, FR}.

05
STAGE 1 · EARLY ACCESS PROGRAM · JAN–APR 2025

Recommendations (Stage 1)

  1. Provide product education (or at least advance notice) that natural pauses should be expected in live translation. Conversing in a language you don’t understand will not be frictionless for a long while, but we can prepare people for a soft landing when using live translation.
  2. Empower conversation partners with product education to facilitate smoother interactions, such as speaking slowly, requesting to repeat a phrase, or explain things a different way. If they’re a complete stranger to the glasses wearer (e.g., travel use case), it’s likely their first time with live translation on glasses, too.
  3. Focusing on improvements to latency (particularly the “pauses” before the starts of new conversational turns) will likely have the greatest impact on user satisfaction, especially when perceived accuracy holds a relatively strong position (X.X for perceived accuracy compared to X.X for audio translation wait times). Latency improvements scale – they help with all language pairs. Dialect and accent recognition improvements can help for each individual language pairing, but the “pauses” are affecting every single language pair.
06
STAGE 2 · POST-GA · APR–AUG 2025

Method for Stage Two, once released out in public

Between the two studies, the team shipped a ~40% improvement in latency and the feature went to the general public in late April 2025.

Phase 1: Live Translation post-GA Survey (n = 377 completes). Recruitment criteria: each person has successfully started live speech translation at least once via voice invocation or via the app itself. Survey is within the Meta AI app (Device Tab), and features about 10 questions on use cases, perceived accuracy, latency, overall satisfaction, and more. Survey displayed in all currently launched markets, as long as phone language settings were set to English (en_US, en_GB).

Phase 2: Follow-Up Interviews (n = 10). Mix of satisfaction levels represented; real world usage prioritized (especially in cases regarding work, travel, conversations with friends, family, and strangers); remote, 60-minute interviews over Zoom.

07
STAGE 2 · POST-GA · APR–AUG 2025

TL;DR (in five points)

#1: Top use cases post-GA are diverse, beyond tourism. Work use cases were mentioned twice as much as travel in the open ended responses. The top categories were work, family/friends/connections, and travel. Many want to be able to use Live Translation beyond 1:1 interactions, especially group conversations and livelier outdoor environments.

Takeaway: To support these desired use cases, we’ll need 1) better noise suppression, and 2) better support for speaker differentiation.

#2: Overall Satisfaction is lower for gen pop compared to early adopters within the Early Access Program. Why the discrepancy? Early Access Program users focused on 1:1 chats, utility use cases (e.g., ordering food, shopkeepers), listening tours. They forgive frictions more than gen pop, even if they recognize the same challenges. Post-GA users have greater experimentation with group conversations & live usage in busy venues.

Implication: Live speaker diarization (“who spoke when?”) will be necessary to support the group conversation use case.

#3: Beyond latency, challenges with picking up background noise is the most common critique, followed by missing words in translations. Missing words can also be exacerbated by background noise, listening to multiple people speaking at once in a group conversation, or listening to an especially fast speaker.

Top Reported Challenges (select all that apply)%
1Latency: Delays in audio translation made conversation awkward, difficult to keep upXX.X%
2Picking up background noise unrelated to the conversationXX.X%
3Translations were missing wordsXX.X%
4Translations were not accurate enoughXX.X%

#4: Accuracy is a greater driver for overall satisfaction / CSAT (~0.X coefficient) compared to latency (~0.X coefficient). For every 1 point gained in accuracy sentiment, we can expect a ~0.X increase in CSAT. For every 1 point gained in latency sentiment, we can expect a ~0.X increase in CSAT.

DriverCoefficient95% CI
Accuracy0.XX0.XX – 0.XX
Naturalness0.XX0.XX – 0.XX
Latency0.XX0.XX – 0.XX

Takeaway for server-side language expansion: we can take the latency hit when switching from on-device to server-side processing. Improvements in language expansion & accuracy are likely better for capturing utility and supporting overall sentiment.

#5: Top language requests: Chinese (Mandarin) XX.X%, Japanese XX.X%, German XX.X%, Chinese (Cantonese) XX.X%, Arabic XX.X%, Korean XX.X%. Only X.X% of our total respondents noted that they were satisfied with the current offerings (FR, IT, ES, EN).

Takeaway for server-side model decision: Unlocking language expansion from shifting to the server would address longstanding feedback on “more languages please” from gen pop, dogfooders, and early adopters over the past year.

08
STAGE 2 · POST-GA · APR–AUG 2025

Use cases in the wild: what's new with the general public compared with the Early Access Program

In the Early Access Program, we saw people go to restaurants to help order items off the menu. In post-GA, we see restaurant workers using live speech translation to understand custom orders. P3 knows simple words like “lechuga” for lettuce in Spanish, but wouldn’t be able to take an entire order in Spanish on his own. He uses live speech translation on his glasses to complete the order – a “full-sized Italian sub with all the meats, oil vinegar”, and cut in half instead of in thirds.

In the Early Access Program, we saw punchy utility conversations among strangers. In post-GA, some even share their glasses during translation sessions. P9 had a 10 minute conversation for home renovations with someone she hired. For more seamless conversation, P9 lends the Ray-Ban Meta Glasses to her conversation partner so he’s able to get the translated audio, and keeps the phone for looking at the live visual transcripts.

Early successes with 1:1 conversations gives gen pop the confidence to increase the stakes and experiment with group conversations. P2 married into a Puerto Rican family, and does not speak Spanish. 1:1 with his brother-in-law, the accuracy was “spot on”. Confidence in strong 1:1 live speech translation performance encouraged him to test out the feature further with a group conversation of 8-9 people at the bar. He tried to address the subsequent performance drop by standing much closer to his conversation partners.

Workplace scenario: even for fluent bilingual speakers, there can be a use case for live speech translation – especially for formal technical and legal language. P8 works as a program coordinator for a substance use disorder group and speaks fluent Spanish. But he may not know health services and regulatory compliance jargon in Spanish (e.g., telehealth agreements, HIPAA). He’ll use live speech translation to help.

Opportunity: Specialized knowledge graphs could give live translation utility for even fluent speakers, especially with work use cases becoming more common and prevalent than travel use cases. From the open-ends, it looks like healthcare, social services, and hospitality are early adopters.

Latency in turn-taking can feel disruptive; keeping live speech translation focused on purposeful & paced conversations (where pauses are natural) helps minimize disruption. P1 is a monolingual English speaker who is newly married to a monolingual Spanish speaker. The pauses from latency nudge him to use Ray-Ban Meta Glasses (and translation apps at large) in a very purposeful way: where to raise the kids? which bills need to be paid? what should we do for date night?

1-way use cases: at work, eavesdropping, ceremonies & rituals requiring complete presence: media (TV, radio, YouTube) XX%, listening in on conversations XX%, live speech XX%. Workplace settings were very common for live speech translation usage, but also celebrations and ceremonies where active phone usage may be less appropriate – baptisms, funerals, mass.

Higher stakes conversations are starting to be seen as well: whether that’s discussing milestone purchases, livelihoods, or literal lives.

“I work in a Spanish speaking kitchen and use it to discuss bigger topics like pay rates and vacations with my employees. Very helpful” – en_US, US

“I used live translation at a car dealership to communicate with Spanish-speaking customers about financing, monthly payments, and vehicle protection plans. The environment had moderate background noise from phones and customer traffic.” – en_US, US

“I work for [a children’s hospital] and I encounter people from all over the world and most can’t speak English so I need to be able to communicate with multiple languages, this would make my job easier” – en_US, US

09
STAGE 2 · POST-GA · APR–AUG 2025

Deep dives, briefly

People are excited to use live speech translation for social settings beyond 1:1 conversation. Because of this, background noise pops up more post-GA than the Early Access Program as the diversity of use cases and settings expand. Live speech translation’s current state is genuinely valuable, but difficulties with handling background noise & multiple speakers can limit use cases in public. 10 out of 10 of the qualitative follow-up interviews mentioned attempted usage in group or non 1:1 conversations, suggesting that speaker differentiation is a critical feature for live speech translation’s continued success.

“When I am around multiple people it tends to pick up speech from multiple sources instead of just the person I am speaking to / looking at as advertised” – en_US, US

Accuracy challenges are often caused by background noise or competing dialogs and conversations. In the past with the Early Access Program, we saw that most complaints were around handling language dialects. For post-GA audiences, we still see some critiques around dialects and too-formal word choice, but missing words entirely is far more of an issue with the general population cohort as they take it to new environments.

“At a restaurant the glasses were picking up 3 or 4 conversations at once and I didn’t know which one was my server… If the glasses could somehow isolate the voice of the person your speaking to it would be revolutionary.” – en_US, US

Latency / wait time for the audio translations remains a top noticeable challenge. Latency was the lowest rated dimension within the live translation survey (X.XX, XX.X% dissatisfied). Which is interesting, because latency has improved by over ~40% since the Early Access Program. Post-GA users notice how background noises affect latency negatively, much like how it affects accuracy negatively. Once overwhelmed with all sorts of information, many people revert to old habits and patterns of behavior.

Takeaway: Latency may be easier to comment on for participants, but improvements in perceived accuracy will likely have better bang for the buck in driving overall sentiment for live speech translation. Let’s consider this when making potential tradeoffs in the on-device vs. server-side processing discussion, where there are tradeoffs in expected accuracy and latency compared to on-device models (i.e., an additional 2-3 seconds, or 2x as long).

4 top drivers for language expansion desires: connecting with family, living & working in a multilingual environment, competitor expectations, travel. “Please unlock German so I can talk to my mother in law!” – en_US, US

10
ACROSS BOTH STAGES

What held, and what changed

Early Access Program (early adopters)Post-GA (general population)
Where people used it1:1 chats, utility use cases (ordering food, shopkeepers), listening tours. Less complex environments.Much more outdoor & group convo usage, and diversity of listening use cases. More complex environments.
Top use caseTravel & short bursts with concrete goalsWork, mentioned 2x as much as travel
Top frictionThe “pauses” at the start of conversation turnsStill latency (XX.X%), but now background noise (XX.X%) and missing words (XX.X%) close behind
Accuracy complaintsDialect & accent recognition, word order (French)Missing words entirely, from noise and multiple speakers
RecommendationFocus on latency / the “pauses”; it scales across all language pairsAccuracy is the greater CSAT driver compared to latency; take the latency hit for server-side language expansion
11
THE IMPACT

Shipped on time, with investment pointed at the right problems.

The two studies gave the team the evidence to launch on schedule, and then told it where to invest next.

  • Drove the GA go/no-go decision and the language-specific launches that followed, with a mixed-methods evidence base of intercept surveys, interviews, and log analyses.
  • Widened the product’s goals to cover one-way use, originally de-prioritized, alongside two-way conversation, after the field data showed people listening to tours, media, and ceremonies as much as talking.
  • Established translation accuracy as the greater CSAT driver compared to latency, which gave the product team confidence to take the latency hit in exchange for rapid language expansion via server-side processing.
  • Turned two recurring failure patterns in field data into committed 2026 roadmap items: live speaker diarization and background-noise handling, directed through engineering and product prioritization after the post-GA study showed in-the-wild use going beyond the original 1:1 intention.

The server-side decision this research supported is its own story: see CS·04, De-risking the AI model decision.