De-risking the AI model decision for multi-language capabilities on glasses
A five-locale survey that cleared the UX-regression concern behind a model migration — a decision that cut operating cost by about half and moved 19+ language launches up by six-plus months.
On-device models, or cloud-hosted?
Wearables AI had to decide whether to keep investing in on-device models or migrate to cloud-hosted models for French, Italian, Spanish (Spain and Latin America), and German — and for the broader language expansion behind them.
Note: I’m happy to discuss further details live (without divulging confidential information), but will share some illustrative examples below as a conversation starter.
A five-locale survey, aimed at the regression concern.
I designed and drove a survey with Wearables AI users across five locales — fr_FR, it_IT, es_ES, es_LA, and de_DE — testing whether cloud-hosted models could unlock rapid language expansion without experience regressions.
Note: The data below is just a placeholder, and is meant to be illustrative.
| fr_FR | it_IT | es_ES | es_LA | de_DE | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model A | Model B | Model A | Model B | Model A | Model B | Model A | Model B | Model A | Model B | |
| n = 126 | n = 147 | n = 139 | n = 129 | n = 142 | n = 122 | n = 136 | n = 147 | n = 90 | n = 124 | |
| Overall satisfaction (CSAT) | 3.4±0.35 | 3.6±0.26 | 3.3±0.15 | 3.5±0.32 | 3.3±0.15 | 3.5±0.34 | 3.5±0.25 | 3.7±0.23 | 3.4±0.23 | 3.5±0.37 |
| Accuracy | 3.3±0.37 | 3.7±0.34 | 3.5±0.15 | 3.7±0.39 | 3.5±0.31 | 3.4±0.37 | 3.5±0.26 | 3.6±0.35 | 3.7±0.27 | 3.9±0.21 |
| Relevance | 3.5±0.32 | 3.6±0.36 | 3.7±0.30 | 3.8±0.34 | 3.2±0.32 | 3.4±0.27 | 3.4±0.31 | 3.6±0.16 | 3.4±0.27 | 3.4±0.36 |
| Naturalness | 3.3±0.33 | 3.7±0.32 | 3.5±0.21 | 3.8±0.30 | 3.4±0.27 | 3.6±0.18 | 3.3±0.16 | 3.6±0.29 | 3.4±0.40 | 3.5±0.32 |
| Latency | 3.4±0.26 | 3.5±0.34 | 3.6±0.18 | 3.7±0.32 | 3.4±0.16 | 3.5±0.35 | 3.2±0.35 | 3.1±0.22 | 3.4±0.20 | 3.3±0.15 |
Illustrative scores (not actual) on a 5-point unipolar scale with 95% confidence intervals — the real values are withheld. Model A = on-device, Model B = cloud-hosted. Highlighted cells mark the higher score where the gap clears the margin of error; everywhere else the two models are statistically level. Here, no regression is the success condition.
Sample sizes chosen based on: availability of Wearables AI users across each language (while being mindful of not burning the entire survey pool), and getting a minimum sample size to notice larger effects.
The regression wasn't there, both in the CSAT scores and the open-ended responses.
The survey showed neutral-to-moderate gains for the cloud-hosted models across all five locales — clearing the UX-regression concern that had kept the decision stuck.
Half the operating cost, and 19+ languages six months sooner.
- De-risked the migration. The five-locale survey cleared the UX-regression concerns that stood between the team and a Go decision on cloud-hosted models.
- Cut OPEX ~50%. The decision the evidence unlocked roughly halved the projected operating cost of multi-language support ($XX million in training additional on-device models, and additional engineering headcount that can be allocated to other workstreams).
- Accelerated 19+ language launches by 6+ months. Language expansion moved from a per-language build to a migration the roadmap could absorb at once.