Skip to main content

What is a UIWorker?

When you put a voice agent in front of an app, talking isn’t enough. The agent needs to see what the user sees and act on the screen: read the page, point at things, fill in fields, click buttons. A UIWorker is the server-side agent that makes this possible. It owns the screen, so the rest of your app doesn’t have to. The connection is two-way, over the RTVI UI channel:
  • Client → server. The client streams the screen to the UIWorker as accessibility snapshots, and forwards the user’s UI interactions as events.
  • Server → client. The UIWorker drives the page back (scrolling, highlighting, selecting text, filling inputs, clicking, or running app-defined commands) and surfaces long-running work as progress cards.
A UIWorker is the screen half of a voice/UI split:
  • The voice agent holds the conversation and does all the talking. It never sees the page.
  • The UIWorker owns the screen. When the voice agent needs something from it, it sends the UIWorker a job, and the UIWorker answers with short data: a label, a yes or no, a list of items. Never the whole page.
The voice LLM keeps a small context focused on the conversation, and it stays the one voice the user hears.
The two directions map to RTVI UI messages: the client sends ui-snapshot and ui-event; the UIWorker sends ui-command and ui-job-group. You rarely touch these directly. PipelineWorker wires the channel up automatically when RTVI is enabled (the default). See The RTVI Standard for the wire protocol and the UIWorker API reference for the class.

The two-way interface

What the UIWorker sees (client → server)

The screen, as a snapshot. The client sends an accessibility snapshot of the page whenever it changes, and the UIWorker keeps the latest one. Each element carries a stable ref the UIWorker uses to act on it. Rendered for an LLM, it looks like this:
Your code reads the raw snapshot through the snapshot property, and the text the user has selected, if any, through selection. When the UIWorker runs its own LLM, it injects the latest <ui_state> into that LLM’s context automatically. User interactions, as events. The client dispatches app-defined events (a button click, a custom gesture) with sendUIEvent(name, payload). Route them to handlers with @ui_event(name); each runs in its own task:

What the UIWorker does (server → client)

Drives the page. The UIWorker acts on the screen by sending UI commands. The built-in helpers cover the common actions, and send_command(name, payload) sends any app-defined command: The standard client handlers ship in @pipecat-ai/client-react; apps can override them or define their own command names. Surfaces long work. Every job group the UIWorker dispatches shows up on the client as a progress card, with a line per peer agent. See Long-running work below.

Hello world

The smallest UIWorker is the class itself, with an LLM and a system prompt. It answers the built-in respond job: it runs one LLM turn with the latest <ui_state> in context, and the reply its LLM writes is the answer.
The voice agent gets a tool that sends the respond job and hands the answer back to its LLM, which speaks it:
Register both with the runner. The UIWorker comes online to receive snapshots and jobs as soon as it starts:
Here’s the full round trip for one question:
1

Snapshot

The client streams the current screen as a ui-snapshot. PipelineWorker broadcasts it on the bus; the UIWorker keeps the latest one.
2

Route

The user asks “what does the second story say?”. The voice LLM can’t see the page, so it calls ask_page, which sends a respond job to the UIWorker.
3

Ground

The UIWorker runs one LLM turn with the latest <ui_state> injected, so its answer is grounded in what’s on screen.
4

Speak

The job returns {"answer": ...} to ask_page. The voice LLM phrases it for the user, and its TTS speaks it.
By default a UIWorker is stateless: it clears its context at the start of each respond job, so every turn sees only the current <ui_state> and question. Set keep_history=True to accumulate history across turns, useful for follow-ups like “and the one after that?”, at the cost of more tokens.

Asking about the screen without an LLM turn

A respond job runs a full LLM turn over the whole page. Many screen questions are smaller than that: which field is “the email field”? Is anything still unchecked? What has the user selected? The UIWorker answers these with its built-in screen job, using a classifier instead of an LLM turn. screen_tools() gives the voice LLM a single screen(action, target, value) tool that sends this job:
The tool’s description teaches the voice LLM each action: Every answer is short data the voice LLM can use directly, such as {"done": true, "label": "Email"}. Describe in the voice LLM’s prompt when to use each action. For example, a voice-guided form might say:
And a reading assistant that resolves “this” and “that”:

Choosing the classifier

Without a classifier argument, the UIWorker answers screen questions with its own LLM through an LLMClassifier. That works, but costs an LLM call per question. Pass a JevClassifier for answers in about a tenth of a second, with calibrated probabilities:

Custom jobs

The built-in jobs are generic. When your app has its own vocabulary, such as a shopping list where the user adds, checks off and removes items, give your UIWorker subclass its own @job handlers. A handler reads the screen with plain code, asks the classifier what it needs to, sends commands, and answers with a short result:
On the voice side, a tool with the app’s own parameters sends the job, and the voice LLM speaks from what it returns:
The snapshot is the source of truth. A read-only job, such as a summary that lists what’s on the list from self.snapshot, lets the voice agent answer “what’s left?” from what’s actually on screen, including items the user checked off by hand. The UIWorker also exposes the classifier questions the screen job uses, for your own handlers: which_element(), select_elements(), check_screen(), and act(), which finds an element and acts on it in one call:

Long-running work

When a job kicks off work that takes a while, fan it out to peer agents with job_group(). Every group a UIWorker dispatches appears on the client as a cancellable progress card, with each agent’s progress streaming in as it arrives. Give the group a label to title the card:
The voice agent’s tool waits on the job while the user watches the cards. To let the user know work has started, the tool can speak a short acknowledgement before it sends the job:

Choosing an approach

They combine freely. One UIWorker can answer respond and screen jobs and your own, and the voice LLM can hold screen_tools() alongside your app’s tools.
  • ReplyToolMixin is deprecated. Give the voice LLM screen_tools() instead, or a custom job for app-specific actions, and let it say the answer.
  • respond_to_job(text, tts_speak=True) is deprecated. A UIWorker should not speak: answer with data, or let the respond job answer with the LLM’s reply, and the voice LLM says it. - BaseUIWorker is deprecated. Dispatch job groups from a @job handler on your UIWorker, which reports them to the client itself.

What’s next

You’ve built agents that converse, call tools, and drive the screen. Next, learn how to transfer control between them.

Agent Handoff

Activation, deactivation, and seamless control transfer

UIWorker API Reference

Full reference for UIWorker, its built-in jobs, screen_tools, and UI commands.