MARIANA OKA

04 / LIBRESTREAM

2021 to 2022

Librestream

Turning recorded field-service calls into something a manager can act on

FOR
Operations managers and directors at industrial companies, overseeing a workforce of field technicians.
ROLE
Product designer and QA tester.
TEAM
Two machine learning researchers, one full-stack developer, one product owner, and me.
TIMELINE
Eight weeks.
TOOLS
Figma, Adobe XD, Pendo, Hotjar, Mailchimp, Zoom, Notion.
OUTCOME
70% of customers approached for the beta opted in. 65% said they would pay more for it.
The Call Insights overview. A What changed this month panel with trending, needs-attention and regional readings, above a ranked list of top topics and a sentiment-over-time chart.
FIG 01The overview. It carries the whole argument, so it opens the page rather than sitting in a strip where it can be scrolled past.

Impact

BETA PROGRAMME

Customers found real patterns in their own data before the beta was over.

THE NUMBER I WOULD DEFEND

12%

spotted trends around specific parts worth investigating, during the beta

Opt-in measures curiosity and willingness to pay measures intent. A customer finding a real pattern in their own data is the product doing its actual job, unprompted.

BETA COHORT

 

70%

of customers approached opted in to the beta

 

65%

said they would pay more for it

And that it would significantly improve their workforce operations.

Introduction

Onsight is remote assistance for industrial work. An expert joins a technician's video call and helps them fix the thing in front of them. Call Insights reads the recordings of those calls and turns them into something a manager can act on: sentiment, recurring topics, summaries, and the named entities that keep coming up.

Problem statement

Operations managers had thousands of recorded field-service calls and no way to see a pattern above a single call, so a recurring fault kept arriving as a string of isolated incidents.

The organisation already had the data and no way to look at it above the level of a single call. A manager could watch one recording. Nobody could answer whether the same part had failed eleven times this quarter.

Every call was handled as a one-off. Each was a record of what broke, what was tried and whether it worked, but nothing connected one call to the next. So a recurring fault arrived as a string of isolated incidents, and the pattern that could have prevented the next one stayed buried in the recordings.

THE PROBLEM

439 calls about one cause, handled as 439 separate incidents.

AS IT WAS4,812 calls, each one handled aloneWHAT WAS IN THEMOne recurring cause, 439 callsINCIDENTINCIDENTINCIDENTREAD ACROSS CALLSCHIPSET SHORTAGE. 834 MENTIONSNOTHING WAS ADDED BETWEEN THE TWO. THE PATTERN WAS IN THE RECORDINGS ALL ALONG.
FIG 02The same 4,812 calls, twice. Each dot stands for about twelve. On the left, the way they arrived: alike, and handled alone. On the right, the 439 about the chipset shortage, joined. Nothing was added between the two panels; only the reading across them changed.

What I looked at

I started with a competitive analysis of how existing call analysis platforms visualised their data. There were plenty of them. None covered what our customers needed end to end, and more usefully, most of them were close to unusable. The analysis was capable and the interfaces were so dense that managers spent more time navigating the tool than learning anything from it.

The gap in that market was not capability. It was that nobody had made the capability legible to the person who had to act on it.

That set the priority for the whole project. I ran ten interviews over Zoom with managers and directors running operations, alongside fifty survey responses through Mailchimp, and synthesised them into a persona and an empathy map. The pattern was consistent: they were not short of data, they were short of a way to see the shape of it. They needed to spot a trend, not read a transcript.

What I was trying to do

Make a machine's reading of a workforce useful to a manager without making it authoritative. Every derived claim had to stay close enough to its evidence that a person could check it, and the interface had to be readable by someone who had never used an analytics tool and did not want to learn one.

I wrote the bet down as a hypothesis before any screen existed: applying natural language processing to support-call transcripts would give operations managers insight they could act on, because sentiment and recurring topics turn thousands of calls into something a person can read. User stories built on that hypothesis set the scope and the order features were built in.

The decisions

01I learned enough machine learning to argue with the researchers

I took Duke's Introduction to Machine Learning through Coursera partway through the project. Not to build anything, but so I could tell the difference between a feature the models could support and one I was inventing.

The platform was being designed while the capability underneath it was still being established. There were long stretches of investigation to find out whether something was even feasible. A designer who cannot participate in that conversation ends up drawing interfaces for capabilities that do not exist, and then negotiating them away one at a time. Being able to read what the researchers were doing meant proposals arrived already scoped to what was real. Diarization, telling one speaker from another in a recording, was the clearest case: a capability we were still learning while I designed the transcript on top of it.

On a team building something that has not been built before, the designer's job includes learning enough of the adjacent discipline to be wrong in useful ways rather than expensive ones.

02Three levels, because one view cannot serve both questions

The product resolves into three screens and the split is the design. A dashboard for the pattern across all calls. A topic deep dive for one recurring theme. A transcript for the single call where you need the actual words.

Managers arrive with two incompatible questions: what is going on across my operation, and what exactly happened on Tuesday. A single view answering both would have served neither. Each level drills into the one below it, so a conclusion is always one click from the material it came from.

FIG 03  2 IMAGES, SWIPE OR USE THE ARROWS

01 / 02

  • The Call Insights overview, with top topics ranked by mentions and change against the previous period.

    FIG 03ALevel one. The pattern across every call, ranked and compared to the period before it.

  • The Chipset shortage topic page. Mentions, unresolved rate and top region as metric cards, above a filterable list of the 439 calls the topic was derived from.

    FIG 03BLevel two. One topic, with the calls it came from listed underneath rather than summarised away.

The overview repeats here on purpose. A drill-down cannot be shown with only its destination, and the pair makes the relationship legible in one figure: a topic named in the ranked list on the left becomes a page with its own numbers, its own comparison window, and the 439 calls it was derived from.

03The transcript is the receipt

Sentiment analysis on your own workforce is a serious thing to put in a product. A summary saying morale is declining on the night shift is a conclusion about people, generated by a model, handed to someone with authority over them.

So the transcript view carries everything the derived claims were built from: the sentiment reading, both the extractive and abstractive summaries, the custom entities, who was on the call and when. The summary is a convenience. The transcript is the evidence, and the design keeps reminding you which is which.

A single call. An audio scrubber marks the four points where the chipset shortage was mentioned. Below it, the transcript with speakers, roles and timestamps, tabs for summary and entities, and a call details panel holding category, location, linked topic and resolution.
FIG 04Level three, and the reason the other two are allowed to make claims. Full-bleed rather than behind a swipe, because the evidence should not be the thing a reader has to go looking for.

04Instrumented before launch, not after

We set KPIs at the start and built the measurement in: analytics for engagement, Hotjar for behaviour, Pendo for in-product surveys, interviews alongside, and a cost-benefit analysis against call centre operations.

That is unglamorous and it is the reason there are numbers below rather than impressions. A product that ships without instrumentation can only be defended with anecdote.

Also at Librestream: LibreHack 2022, first place

A 48-hour remote hackathon, with developers I had never worked with, around one question: what could Azure Spatial Anchors do for an industrial workforce? We landed on two uses of the same capability. Navigation, placing arrows through a large, complex facility so a worker reaches the right equipment without asking anyone. And training, anchoring a hologram of a machine in the room so its parts can be walked around and explained.

Forty-eight hours rules out building the real thing, so I prototyped the experience instead: 3D models placed in real rooms with Adobe Aero, composited into the redesigned Onsight mobile interface I had been working on, with a text-to-speech narrator as the guide. The helper character was designed to replace a loading screen with something that explains while you wait. The demos ran on a phone; the interaction was designed to carry to tablets and HoloLens.

FIG 05The navigation concept: a helper that animates while the route loads, then arrows anchored to the floor, room to room, to the right switch. The hackathon was remote, so my apartment stood in for the facility and a breaker panel for the equipment. Sound on to hear the guide.
FIG 06The training concept: a BMW 801 aircraft engine anchored in the room and exploded into its parts, narrated by a generated voice (ElevenLabs). Sound on to hear it.

We won. The development team also left with a basic working version of Azure Spatial Anchors on HoloLens, and both concepts were taken forward as avenues for the business. One finding shaped the pitch: clients wouldn't need 3D artists to author training, because a photogrammetry app like Polycam can scan real equipment into models.

With more time I would have added switchable captions, since a voice-only guide fails first on a loud factory floor, and layered instructions over the interface rather than beside it.

What I would do differently

Sentiment about a workforce is a serious thing to show a manager. I'd add confidence indicators and a clear "why the model thinks this" view from the start, rather than relying on the transcript alone as the receipt.

Reflection

This was AI product design before the phrase meant anything. Summarisation, sentiment and topic clustering, in an industrial setting, with a model whose limits were real and close. The problems were the ones I still work on now: what a system is allowed to assert, how a person checks it, and what happens to trust when the machine is confidently wrong.

Keeping the evidence one click from the conclusion is the same instinct as an agent that drafts and waits. Neither is about restricting what the machine can do. Both are about refusing to let the person lose their grip on what it actually did.