> ## Documentation Index
> Fetch the complete documentation index at: https://daily-ms-ws-body-url-encode.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Controlling the UI

> Let a voice agent read and act on a client GUI with UIWorker: respond and screen jobs, screen_tools, classifiers, and job groups over RTVI.

## What is a UIWorker?

When you put a voice agent in front of an app, talking isn't enough. The agent needs to *see what the user sees* and *act on the screen*: read the page, point at things, fill in fields, click buttons. A `UIWorker` is the server-side agent that makes this possible. It owns the screen, so the rest of your app doesn't have to.

The connection is **two-way**, over the RTVI UI channel:

* **Client → server.** The client streams the screen to the `UIWorker` as accessibility snapshots, and forwards the user's UI interactions as events.
* **Server → client.** The `UIWorker` drives the page back (scrolling, highlighting, selecting text, filling inputs, clicking, or running app-defined commands) and surfaces long-running work as progress cards.

A `UIWorker` is the screen half of a voice/UI split:

* The **voice agent** holds the conversation and does all the talking. It never sees the page.
* The **`UIWorker`** owns the screen. When the voice agent needs something from it, it sends the `UIWorker` a job, and the `UIWorker` answers with short data: a label, a yes or no, a list of items. Never the whole page.

The voice LLM keeps a small context focused on the conversation, and it stays the one voice the user hears.

<Note>
  The two directions map to RTVI UI messages: the client sends `ui-snapshot` and
  `ui-event`; the `UIWorker` sends `ui-command` and `ui-job-group`. You rarely
  touch these directly. `PipelineWorker` wires the channel up automatically when
  RTVI is enabled (the default). See [The RTVI
  Standard](/client/rtvi-standard#user-interface) for the wire protocol and the
  [UIWorker API reference](/api-reference/server/workers/ui-worker) for the
  class.
</Note>

## The two-way interface

### What the UIWorker sees (client → server)

**The screen, as a snapshot.** The client sends an accessibility snapshot of the page whenever it changes, and the `UIWorker` keeps the latest one. Each element carries a stable `ref` the `UIWorker` uses to act on it. Rendered for an LLM, it looks like this:

```
<ui_state>
- heading "Shopping list" [level=1] [ref=e3]
- list:
  - checkbox "milk" [checked] [ref=e5]
  - checkbox "eggs" [ref=e6]
</ui_state>
```

Your code reads the raw snapshot through the `snapshot` property, and the text the user has selected, if any, through `selection`. When the `UIWorker` runs its own LLM, it injects the latest `<ui_state>` into that LLM's context automatically.

**User interactions, as events.** The client dispatches app-defined events (a button click, a custom gesture) with `sendUIEvent(name, payload)`. Route them to handlers with `@ui_event(name)`; each runs in its own task:

```python theme={null}
from pipecat.workers.ui import UIWorker, ui_event

class MyUIWorker(UIWorker):
    @ui_event("note_click")
    async def on_note_click(self, message):
        ref = (message.payload or {}).get("ref")
        await self.scroll_to(ref)
        await self.select_text(ref)
```

### What the UIWorker does (server → client)

**Drives the page.** The `UIWorker` acts on the screen by sending UI commands. The built-in helpers cover the common actions, and `send_command(name, payload)` sends any app-defined command:

| Helper | Effect |
| - | - |
| `scroll_to(ref)` | Bring an element into view |
| `highlight(ref)` | Briefly flash an element |
| `select_text(ref)` | Select an element's text, to point at it |
| `click(ref)` | Click a checkbox, radio, or button |
| `set_input_value(ref, value)` | Fill a text input or textarea |
| `send_command(name, payload)` | Any app-defined command (e.g. `"add_item"`) |

The standard client handlers ship in `@pipecat-ai/client-react`; apps can override them or define their own command names.

**Surfaces long work.** Every job group the `UIWorker` dispatches shows up on the client as a progress card, with a line per peer agent. See [Long-running work](#long-running-work) below.

## Hello world

The smallest `UIWorker` is the class itself, with an LLM and a system prompt. It answers the built-in **`respond`** job: it runs one LLM turn with the latest `<ui_state>` in context, and the reply its LLM writes is the answer.

```python theme={null}
from pipecat.workers.ui import UIWorker

UI_PROMPT = """\
You answer questions about the page the user is looking at, in one or \
two plain sentences. Don't tell the user what you can't see; answer, \
or say you don't know."""

ui_worker = UIWorker(
    "ui",
    llm=OpenAILLMService(
        api_key=os.environ["OPENAI_API_KEY"],
        settings=OpenAILLMService.Settings(system_instruction=UI_PROMPT),
    ),
)
```

The voice agent gets a tool that sends the `respond` job and hands the answer back to its LLM, which speaks it:

```python theme={null}
from pipecat.adapters.schemas.direct_function import tool_options
from pipecat.pipeline.job_context import JobError, JobParams


@tool_options(cancel_on_interruption=False, timeout_secs=60)
async def ask_page(params: FunctionCallParams, question: str):
    """Ask about the page the user is looking at.

    Call it for any question that could be about the page; nothing else
    can see it. Returns the answer.
    """
    try:
        async with params.pipeline_worker.job(
            "ui", params=JobParams(name="respond", payload={"query": question}, timeout=30)
        ) as t:
            pass
    except JobError as e:
        await params.result_callback({"error": str(e)})
        return
    await params.result_callback(t.response)


context = LLMContext(tools=[ask_page])
```

Register both with the runner. The `UIWorker` comes online to receive snapshots and jobs as soon as it starts:

```python theme={null}
await runner.add_workers(ui_worker, worker)
```

Here's the full round trip for one question:

<Steps>
  <Step title="Snapshot">
    The client streams the current screen as a `ui-snapshot`. `PipelineWorker`
    broadcasts it on the bus; the `UIWorker` keeps the latest one.
  </Step>

  <Step title="Route">
    The user asks "what does the second story say?". The voice LLM can't see the
    page, so it calls `ask_page`, which sends a `respond` job to the `UIWorker`.
  </Step>

  <Step title="Ground">
    The `UIWorker` runs one LLM turn with the latest `<ui_state>` injected, so
    its answer is grounded in what's on screen.
  </Step>

  <Step title="Speak">
    The job returns `{"answer": ...}` to `ask_page`. The voice LLM phrases it
    for the user, and its TTS speaks it.
  </Step>
</Steps>

<Note>
  By default a `UIWorker` is stateless: it clears its context at the start of
  each `respond` job, so every turn sees only the current `<ui_state>` and
  question. Set `keep_history=True` to accumulate history across turns, useful
  for follow-ups like "and the one after that?", at the cost of more tokens.
</Note>

## Asking about the screen without an LLM turn

A `respond` job runs a full LLM turn over the whole page. Many screen questions are smaller than that: which field is "the email field"? Is anything still unchecked? What has the user selected? The `UIWorker` answers these with its built-in **`screen`** job, using a [classifier](/pipecat/learn/classifiers) instead of an LLM turn.

`screen_tools()` gives the voice LLM a single `screen(action, target, value)` tool that sends this job:

```python theme={null}
from pipecat.workers.ui import UIWorker, screen_tools

context = LLMContext(tools=screen_tools("ui"))

ui_worker = UIWorker("ui", llm=ui_llm)
```

The tool's description teaches the voice LLM each action:

| Action | What it answers or does |
| - | - |
| `find` | Which element a description means, such as "the checkout button" |
| `check` | Whether something is true of the screen, with a probability |
| `select` | Which elements match a description, such as "dairy products" |
| `list` | The named elements on screen with their state and values, optionally of one role |
| `selection` | The text the user has selected |
| `click`, `scroll_to`, `highlight`, `select_text`, `fill` | Do that to the element a description means; `fill` writes a `value` |

Every answer is short data the voice LLM can use directly, such as `{"done": true, "label": "Email"}`. Describe in the voice LLM's prompt when to use each action. For example, a voice-guided form might say:

```text theme={null}
- screen(action="list", target="textbox"): the form's inputs with their
  current values. Call it at the start of every turn to see which fields
  are filled and steer toward the next empty one.
- screen(action="fill", target=..., value=...): write one value into the
  field the target names, such as "the email field".
- screen(action="click", target="the submit button"): submit, at the end.
```

And a reading assistant that resolves "this" and "that":

```text theme={null}
You cannot know what the user has selected; only your tools can. Whenever
the user says "this", "that" or "this paragraph", call screen("selection")
first and answer from the text it returns.
```

### Choosing the classifier

Without a `classifier` argument, the `UIWorker` answers screen questions with its own LLM through an `LLMClassifier`. That works, but costs an LLM call per question. Pass a `JevClassifier` for answers in about a tenth of a second, with calibrated probabilities:

```python theme={null}
from pipecat.classifiers.jev.classifier import JevClassifier

ui_worker = UIWorker(
    "ui",
    llm=ui_llm,
    classifier=JevClassifier(api_key=os.getenv("TYPESAFE_API_KEY")),
)
```

## Custom jobs

The built-in jobs are generic. When your app has its own vocabulary, such as a shopping list where the user adds, checks off and removes items, give your `UIWorker` subclass its own `@job` handlers. A handler reads the screen with plain code, asks the classifier what it needs to, sends commands, and answers with a short result:

```python theme={null}
from pipecat.bus.messages import BusJobRequestMessage
from pipecat.classifiers.base_classifier import ChoiceQuestion
from pipecat.pipeline.job_decorator import job
from pipecat.workers.ui import UIWorker


class ListWorker(UIWorker):
    @job(name="update")
    async def _update(self, message: BusJobRequestMessage) -> None:
        payload = message.payload or {}

        # Adding needs no classifier: the voice LLM already carries the text.
        for text in payload.get("add", []):
            await self.send_command("add_item", {"text": text})

        # For items named in the user's words, ask the classifier which
        # checkbox on screen each one means, all in one call.
        items = {ref: name for ref, name in self._checkboxes()}
        to_check = payload.get("check", [])
        questions = {
            text: ChoiceQuestion(instructions=f"the list item {text!r} refers to", options=items)
            for text in to_check
        }
        not_found = []
        if items and questions:
            results = await self.classifier.choice({"items": to_check}, questions)
            for text, result in results.items():
                if result.confidence >= 0.5:
                    await self.send_command("set_checked", {"ref": result.choice, "checked": True})
                else:
                    not_found.append(text)

        await self.send_job_response(message.job_id, {"not_found": not_found})

    def _checkboxes(self) -> list[tuple[str, str]]:
        """The (ref, name) of each checkbox in the latest snapshot."""
        ...
```

On the voice side, a tool with the app's own parameters sends the job, and the voice LLM speaks from what it returns:

```python theme={null}
@tool_options(cancel_on_interruption=False, timeout_secs=15)
async def update_list(
    params: FunctionCallParams,
    add: list[str] | None = None,
    check: list[str] | None = None,
):
    """Change the shopping list.

    Args:
        add: Items to add, named as they should appear on the list.
        check: Items to check off, as on the list or as the user said them.
    """
    async with params.pipeline_worker.job(
        "ui", params=JobParams(name="update", payload={"add": add or [], "check": check or []})
    ) as t:
        pass
    await params.result_callback(t.response)
```

The snapshot is the source of truth. A read-only job, such as a `summary` that lists what's on the list from `self.snapshot`, lets the voice agent answer "what's left?" from what's actually on screen, including items the user checked off by hand.

The `UIWorker` also exposes the classifier questions the `screen` job uses, for your own handlers: `which_element()`, `select_elements()`, `check_screen()`, and `act()`, which finds an element and acts on it in one call:

```python theme={null}
class ReviewWorker(UIWorker):
    @job(name="add_note")
    async def _add_note(self, message: BusJobRequestMessage) -> None:
        text = (message.payload or {}).get("text", "")
        textarea = await self.which_element("the notes textarea")
        save = await self.which_element("the Save button")
        if not text or not textarea or not save:
            await self.send_job_response(message.job_id, {"done": False})
            return
        await self.set_input_value(textarea, text)
        await self.click(save)
        await self.send_job_response(message.job_id, {"done": True})
```

## Long-running work

When a job kicks off work that takes a while, fan it out to peer agents with `job_group()`. Every group a `UIWorker` dispatches appears on the client as a cancellable progress card, with each agent's progress streaming in as it arrives. Give the group a `label` to title the card:

```python theme={null}
from pipecat.pipeline.job_context import JobGroupError, JobGroupParams, JobStatus


class ResearchWorker(UIWorker):
    @job(name="research")
    async def _research(self, message: BusJobRequestMessage) -> None:
        query = (message.payload or {}).get("query", "")
        try:
            async with self.job_group(
                "wikipedia", "news", "scholar",
                params=JobGroupParams(payload={"query": query}, label=f"Research: {query}"),
            ) as group:
                pass
        except JobGroupError as e:
            await self.send_job_response(message.job_id, {"error": str(e)}, status=JobStatus.ERROR)
            return
        summaries = {name: r.get("summary") for name, r in group.responses.items()}
        await self.send_job_response(message.job_id, {"results": summaries})
```

The voice agent's tool waits on the job while the user watches the cards. To let the user know work has started, the tool can speak a short acknowledgement before it sends the job:

```python theme={null}
@tool_options(cancel_on_interruption=False, timeout_secs=60)
async def research(params: FunctionCallParams, query: str):
    """Research a topic across three sources and return their summaries."""
    await params.llm.push_frame(TTSSpeakFrame(f"Researching {query} now."))
    async with params.pipeline_worker.job(
        "ui", params=JobParams(name="research", payload={"query": query}, timeout=60)
    ) as t:
        pass
    await params.result_callback(t.response)
```

## Choosing an approach

| | `respond` job | `screen` tool | Custom `@job` |
| - | - | - | - |
| **What answers** | The `UIWorker`'s LLM, with the page in context | The `UIWorker`'s classifier | Your code, with the snapshot and the classifier |
| **Cost** | One LLM turn | One classifier call, fast with `JevClassifier` | Whatever your handler does |
| **Best for** | Open questions about the page's content | Finding, checking, pointing, filling, clicking | App-specific actions and long-running work |
| **Setup** | A system prompt | `screen_tools()` in the voice LLM's context | A `@job` handler and a voice tool that sends it |

They combine freely. One `UIWorker` can answer `respond` and `screen` jobs and your own, and the voice LLM can hold `screen_tools()` alongside your app's tools.

<Accordion title="Migrating from earlier versions">
  * **`ReplyToolMixin`** is deprecated. Give the voice LLM `screen_tools()`
    instead, or a custom job for app-specific actions, and let it say the answer.
  * **`respond_to_job(text, tts_speak=True)`** is deprecated. A `UIWorker`
    should not speak: answer with data, or let the `respond` job answer with the
    LLM's reply, and the voice LLM says it. - **`BaseUIWorker`** is deprecated.
    Dispatch job groups from a `@job` handler on your `UIWorker`, which reports
    them to the client itself.
</Accordion>

## What's next

You've built agents that converse, call tools, and drive the screen. Next, learn how to transfer control between them.

<CardGroup cols={2}>
  <Card title="Agent Handoff" icon="arrow-right" href="/pipecat/learn/agent-handoff">
    Activation, deactivation, and seamless control transfer
  </Card>

  <Card title="UIWorker API Reference" icon="book" href="/api-reference/server/workers/ui-worker">
    Full reference for `UIWorker`, its built-in jobs, `screen_tools`, and UI
    commands.
  </Card>
</CardGroup>
