MODEL STRATEGY · 2025 · META WEARABLES AI

De-risking the AI model decision for multi-language capabilities on glasses

A five-locale survey that cleared the UX-regression concern behind a model migration — a decision that cut operating cost by about half and moved 19+ language launches up by six-plus months.

Ray-Ban Meta glasses on a tan surface in window light, with five small flag tiles in a row beneath them — France, Italy, Spain, Mexico, Germany.
Ray-Ban Meta glasses and the five survey locales.
01
THE PROBLEM

On-device models, or cloud-hosted?

Wearables AI had to decide whether to keep investing in on-device models or migrate to cloud-hosted models for French, Italian, Spanish (Spain and Latin America), and German — and for the broader language expansion behind them.

Note: I’m happy to discuss further details live (without divulging confidential information), but will share some illustrative examples below as a conversation starter.

02
THE APPROACH

A five-locale survey, aimed at the regression concern.

I designed and drove a survey with Wearables AI users across five locales — fr_FR, it_IT, es_ES, es_LA, and de_DE — testing whether cloud-hosted models could unlock rapid language expansion without experience regressions.

Note: The data below is just a placeholder, and is meant to be illustrative.

fr_FRit_ITes_ESes_LAde_DE
Model AModel BModel AModel BModel AModel BModel AModel BModel AModel B
n = 126n = 147n = 139n = 129n = 142n = 122n = 136n = 147n = 90n = 124
Overall satisfaction (CSAT)3.4±0.353.6±0.263.3±0.153.5±0.323.3±0.153.5±0.343.5±0.253.7±0.233.4±0.233.5±0.37
Accuracy3.3±0.373.7±0.343.5±0.153.7±0.393.5±0.313.4±0.373.5±0.263.6±0.353.7±0.273.9±0.21
Relevance3.5±0.323.6±0.363.7±0.303.8±0.343.2±0.323.4±0.273.4±0.313.6±0.163.4±0.273.4±0.36
Naturalness3.3±0.333.7±0.323.5±0.213.8±0.303.4±0.273.6±0.183.3±0.163.6±0.293.4±0.403.5±0.32
Latency3.4±0.263.5±0.343.6±0.183.7±0.323.4±0.163.5±0.353.2±0.353.1±0.223.4±0.203.3±0.15

Illustrative scores (not actual) on a 5-point unipolar scale with 95% confidence intervals — the real values are withheld. Model A = on-device, Model B = cloud-hosted. Highlighted cells mark the higher score where the gap clears the margin of error; everywhere else the two models are statistically level. Here, no regression is the success condition.

Sample sizes chosen based on: availability of Wearables AI users across each language (while being mindful of not burning the entire survey pool), and getting a minimum sample size to notice larger effects.

03
THE FINDING

The regression wasn't there, both in the CSAT scores and the open-ended responses.

The survey showed neutral-to-moderate gains for the cloud-hosted models across all five locales — clearing the UX-regression concern that had kept the decision stuck.

04
THE IMPACT

Half the operating cost, and 19+ languages six months sooner.

  • De-risked the migration. The five-locale survey cleared the UX-regression concerns that stood between the team and a Go decision on cloud-hosted models.
  • Cut OPEX ~50%. The decision the evidence unlocked roughly halved the projected operating cost of multi-language support ($XX million in training additional on-device models, and additional engineering headcount that can be allocated to other workstreams).
  • Accelerated 19+ language launches by 6+ months. Language expansion moved from a per-language build to a migration the roadmap could absorb at once.