GOOGLE · ASSISTANT · 2019

Bringing conversational experiences to everyone

A four-city international baseline — Tokyo, New Delhi, New York, Mountain View — to make sure the Assistant works everywhere, for everyone, and to measure progress over time.

Assistant results for “restaurants nearby” shown side by side in San Francisco, Tokyo, and New Delhi — on iPhone, Android, and the Jio feature phone.
From left to right: Assistant search results for the query “restaurants nearby” in San Francisco, Tokyo, and New Delhi. The interfaces are production results from iPhone, Android, and the Jio feature phone (available in India only).
01
THE PROBLEM

How can we ensure the Google Assistant works everywhere and for everyone?

The Geo Assistant aims to be an effective personal companion for the real world – whether that’s providing useful information for the morning commute, suggesting local places to explore, or answering questions people have about their surroundings.

Given this, how would we make sure we’re actually providing value and solving genuine problems – not just for people in Silicon Valley, but also for people and communities worldwide? And how might we measure our progress in doing so over time?

To do this, we conducted a UX baseline lab study with external participants in 4 cities worldwide: Tokyo 🇯🇵, New Delhi 🇮🇳, New York City 🇺🇸, and Mountain View 🇺🇸. We focused on how production versions of the Geo Assistant perform for our critical user journeys in Commute, Transit, Local Answers & Actions. The goal of this study is to understand Geo Assistant performance and establish a baseline to track UX progress (e.g., HEART metrics) every six months.

We feature a combination of participant self-reported metrics (through an administered survey after each task) and researcher-reported metrics (e.g., task success, discoverability of relevant features).

Happy to explain further details in person. Confidential information will not be provided.

02
BASELINE STUDY APPROACH

Understand how production versions of the Geo Assistant perform worldwide, and establish a baseline to track UX progress every six months.

Overhead view of a lab session: a participant fills in a paper survey next to their own phone showing Assistant results, with a task card reading “Red Rooster.”
In the lab in New York City — each participant brought their own phone.

For studies within the United States, we brought people into labs on campus and had them evaluate experiences related to daily commutes, public transit, and local place discovery. Each person brought their own phone for this study.

I set up and moderated the first study in New York City. My colleague Jenny Spencer (UXR) set up and moderated the second study in Mountain View, California.

A localized task card reading “CVS / Matsumoto Kiyoshi nearby” beside Assistant results listing three Matsumoto Kiyoshi drugstores in Shinjuku, Tokyo.
A localized task from the Tokyo study — the drugstore chain Matsumoto Kiyoshi stands in for CVS.

For the international studies, we worked together with local research partners and vendors to make sure we had representative participants, the proper localized tasks, and translated study scripts. I partnered with the research team in Google India, while Ariel Liu (Senior UXR) partnered with a research vendor in Tokyo, Japan.

Slide titled “What kinds of tasks did we test?” pairing three spoken queries — commute time, gas stations along a route, restaurants near me — with the Assistant screens they produce.
A sample of the tasks / critical user journeys we evaluated on the mobile Google Assistant surface. There are more available.
03
CHALLENGES

Challenges we addressed along the way

Two spoken tasks — “Where am I?” and “CVS / Matsumoto Kiyoshi nearby” — shown with their Assistant results in both Tokyo and New York.
The same tasks, localized: “Where am I?” and “CVS nearby” as they rendered in Tokyo and New York.

Localization. A one-size-fits-all script wouldn’t be relevant in all places around the world. International communities have their own tastes and needs. So we start with a skeleton or base script, and then tailor the particular destinations from there. We also work with translators and international teammates to make sure we meet the users where they are. That adds more time necessary for preparation work.

Logistics. Coordinating with research partners and third-party vendors who may have the localization expertise, but may not be as familiar with this particular digital product or experience, or the context and history. There’s an onboarding process to bring partners up to speed.

Managing (enthusiastic) stakeholder expectations. The great news is: nearly all of our stakeholders were excited about this research, and wanted to get involved. But when we’re juggling asks from different roles (e.g., engineering asks and product manager asks) and different teams/product areas (and each team has a multitude of user journeys they want to test), the list of requests and potential study questions ballooned in scope. We only have 1 hour scheduled with each participant, and there won’t be enough time to account for all of the different tasks out there. By involving the diverse stakeholders early in the script writing process, we narrowed down the different asks to ten scenarios/tasks, along with some additional bonus tasks if there was extra time.

04
DELIVERABLES & IMPACT

This was the first international baseline study for the Geo Assistant team.

Informed 2020 product strategy on themes related to personalization and a more helpful, proactive Assistant. Teams from four different product areas (Transit, Driving, Local Answers + Actions, Next Billion Users) were present for the presentation and follow-up discussion. Stakeholders included product managers, interaction designers, engineering, data scientists & analysts.

  • Captured 10 bugs in production versions of the Geo Assistant for the engineering team.
  • Team research impact: First international baseline study for the Geo Assistant team. Serves as a template for future studies, especially since the team plans on repeating this study every six months to see how far the Geo Assistant has improved with every product release cycle.
  • Organizational research impact: Influenced other Google Assistant teams and organizations who were looking for ways of establishing their own baseline to measure their product performance over time.
  • Research communication: UX baseline scores were integrated into a dashboard with the engineering team, where the UX success metrics now live along with other key performance metrics.
05
REFLECTIONS & LEARNINGS

Evaluating a conversational experience over time is tricky.

Usability metrics don’t work too well for a conversational experience, especially over a longer period of time. What does the difference between a 4.2 and a 4.5 on a self-reported metric like Ease of Use actually tell you? The real value in these kinds of research projects is the qualitative feedback; the scores provide some preliminary measures.

Scaling the traditional lab study to work internationally and not just locally is tricky. But it’s so worth it. It also takes a while. Being able to unite all the separate studies into a single narrative can get incredibly tricky. Or at least time-consuming.

Evaluations across cultures are not always equal. Some participants on average are more forgiving than others.

How often should we re-evaluate and try the baseline again? The original timeline we proposed is once every 6 months. But are all teams shipping at a rapid rate of 6 months, or pushing enough changes to warrant another study? Maintaining these kinds of measurements takes a lot of extra effort.

Google Assistant directory page for Google Maps listing example voice commands such as “Are there traffic jams?” and “How long will it take me to work?”, with a 4.6-star rating below.
Screenshot from public Google Assistant + Maps documentation. Not all of these suggestion chips were scenarios we tested, but it's along similar lines. You could try out some of the commands yourself.