Skip to content
Kevo Rojas

My AI assistant on Meta Quest 3: how it works and a prompt to build yours

KR
Kevo Rojas
Oct 9, 2026 · 11 min read
Tutorial video

Edison is an AI assistant that lives in my Meta Quest 3: I talk to it, it answers with voice and panels in my room. How it works, plus a prompt to build yours.

In short: Edison is an AI assistant on Meta Quest 3: a web page I open in the headset's browser, connected to a server on my computer. I talk to it, it answers with voice, and panels float around my room. Voice is processed locally, simple requests are answered by a small model running on Ollama, and the heavy lifting goes to the cloud. At the end there's a prompt so your coding agent can build you a first version.

Someone on TikTok asked me how I made it. This article is the answer: what Edison is, how it works inside and a prompt to build your own.

What Edison is

Edison is a voice assistant I built for my Meta Quest 3, in mixed reality. I put on the headset, open a page in the browser, enter immersive mode and still see my room, with panels floating on top. I talk to it and it answers out loud.

The name comes from One Piece: Edison is one of Vegapunk's satellites, the one who thinks.

It's a personal project. It isn't a product or an app you can download, and some things still fail. I built it in April 2026, and what I describe here is how it was at that point.

The videos

The video above is part 1: the idea, how I built it and the first test with the headset on. It's in Spanish.

Part 2 shows what it already does and what still fails: tasks, calendar and weather by voice, and a request that didn't go as planned. It's here: https://youtu.be/eAniLy-YehY

How it works

There are two pieces: the headset and my computer, on the same wifi network.

Meta Quest 3 (browser, WebXR)
        │  secure WebSocket (wss), local network
        ▼
Server on my computer (FastAPI)
  1. whisper.cpp ── your voice to text       (local)
  2. router ─┬── chat and simple requests ── Ollama, small model (local)
             └── reasoning, planning, code ── cloud model
  3. Kokoro ── the answer to speech          (local)
        │
        ▼
Audio back to the headset + floating panels
  • The headset has no app installed. It's a page built with Three.js and WebXR that opens in the Quest browser. For debugging, the same page works in the computer's browser.
  • Everything runs over HTTPS on the local network, with a certificate made with mkcert. That's not a detail: without HTTPS, the Quest won't give the page access to WebXR or the microphone. Nothing is exposed to the internet.
  • Voice is processed on my computer. whisper.cpp turns the audio into text and Kokoro turns the answer into speech.
  • A router decides which model answers. Chat and simple requests go to a small model running locally on Ollama, which answers in about a second. Anything that needs reasoning, planning or writing code goes to a cloud model; in my case, Claude Code. If the local model fails, the request falls back to the cloud.

That's why I don't call it "100% local": voice is, and so are simple requests, but the heavy work doesn't run on my computer.

Here's the prompt to build it

One note first: this isn't the prompt I used to build Edison. I built it over several sessions, iterating and fixing things. What follows is a prompt that sums up the architecture and what I learned along the way, split into phases, so a coding agent can build you a first working version.

What you need

  • A Meta Quest 3 and a computer on the same wifi network.
  • A computer that can handle whisper.cpp, Kokoro and a small Ollama model. I ran it on an Apple Silicon Mac with 16 GB of RAM; I haven't tried Windows or Linux.
  • A coding agent: Claude Code or any other that can create files and run commands on your computer.
  • Optional: access to a cloud model for complex requests.

How to use it

  1. Fill in the placeholders before pasting it. They're the text between double curly braces:
PlaceholderWhat to putExample
{{ASSISTANT_NAME}}Your assistant's nameEdison
{{COMPUTER_OS}}Your operating system and hardwaremacOS on Apple Silicon
{{TTS_VOICE}}One of Kokoro's voices in your language(see Kokoro's voice list)
{{LANGUAGE}}The language you'll speak to it inEnglish
{{LOCAL_MODEL}}The small Ollama modelgemma3:1b, the one I used
{{CLOUD_LLM}}The model for complex requestsClaude Code, or the API you prefer
  1. Paste it in an empty folder with your coding agent open there.
  2. Go phase by phase. The prompt tells the agent not to move on until each phase meets its criteria, and to tell you what to test when it's done. That part is yours, with the headset on: the agent can't test WebXR on the Quest.
  3. If something breaks on the headset, give it the error from the remote devtools (chrome://inspect with the Quest connected). The last section of the prompt already warns it about the problems I ran into.
(Note: this isn't literally the prompt I used to build Edison; I built it over several sessions, iterating. It's a summary of the architecture and what I learned, meant to help your coding agent build you a first version.)

I want to build a voice AI assistant called {{ASSISTANT_NAME}} that lives in my Meta Quest 3 in mixed reality. I talk to it, it answers out loud, and floating panels appear in my room. Work in phases: don't move on to the next one until the current one meets its acceptance criteria, and when you finish each phase, tell me exactly what I need to test with the headset on (you can't test WebXR on the hardware).

ARCHITECTURE
- Frontend: a web page with Three.js + WebXR ("immersive-ar" session with passthrough) opened from the Quest browser. No native app, no Unity, no XR frameworks on top.
- Backend: FastAPI (Python 3.11+) running on my computer ({{COMPUTER_OS}}), on the same wifi network as the headset.
- STT: whisper.cpp compiled locally, with my hardware's acceleration. Input: 16 kHz mono WAV.
- TTS: Kokoro, local, voice {{TTS_VOICE}} in {{LANGUAGE}}. Output: 24 kHz WAV.
- LLM with a router: {{LOCAL_MODEL}} via Ollama (OpenAI-compatible API) for chat and simple things; {{CLOUD_LLM}} for anything that needs reasoning, planning or writing code.
- Transport: WebSocket (wss) between headset and backend; POST /voice as an alternative.
- Configuration with pydantic-settings and .env. No keys in the repo. Code in English.

NON-NEGOTIABLE REQUIREMENTS
- HTTPS is mandatory: WebXR and the microphone don't work on the Quest without HTTPS. Use mkcert and generate a certificate that includes localhost, 127.0.0.1 and my computer's IP on the local network. Put the IP in a config variable and create a script to regenerate the certificate if the IP changes.
- Nothing exposed to the internet: no public tunnels. CORS limited to the local network.
- A start script that brings up backend and frontend, and a setup script that compiles whisper.cpp, downloads models and creates the virtual environment.
- A desktop mode: the same page must work in the computer's browser with mouse and keyboard, to debug without the headset. Add debug keyboard shortcuts.
- Measure and log the latency of each stage (STT, LLM, TTS) on every turn. Expose it in GET /health along with the status of each component.
- Split the frontend into modules from the start (scene, audio, panels, network); don't put everything in a single index.html.

PHASE 1: voice pipeline on the computer
- Separate stt, llm and tts services. POST /voice: audio in, audio out. GET /health.
- A terminal test script with two modes: text (I type and it answers with audio) and voice (I record with the computer's microphone).
- Short system prompt: answers of 2 or 3 sentences max, plain text with no markdown or emojis (the output is heard, not read).
Criteria: I talk through the terminal and hear the answer; /health shows all three components green; the log shows latency per stage.

PHASE 2 (MVP): talking to it from the Quest
- HTTPS page with an "Enter MR" button that starts the immersive-ar session.
- Microphone capture in the browser (push to talk), sent over WebSocket as binary frames, a text message to mark the end, audio response and status messages. Protocol with simple prefixes (for example STT:, LLM:, DONE:, ERROR:). Automatic reconnection with ping/pong.
- A floating assistant panel with its status (listening, transcribing, thinking, speaking) and the last answer. Panels are planes with a canvas as texture.
- A simple avatar (for example, an icosahedron with a shader) that changes with the status and reacts to the amplitude of the answer's audio using an AnalyserNode.
Criteria: with the headset on I enter MR, see the panel in my room, talk to it, hear the answer, and the panel shows what it understood and what it answered. If the WebSocket drops, it reconnects on its own.

PHASE 3: local/cloud router
- Routing function with configurable modes: off, local-first, local-only and cloud-only. Default: local-first.
- "Complex" heuristic: long query, reasoning words (architecture, refactor, debug, plan) or questions about my personal stuff (projects, tasks, status). Classify by intent, not only by exact name: whisper mangles proper names.
- The local model gets its own system prompt, short and very directive, with mandatory literal phrases for what it doesn't know. 1B models make things up with long prompts.
- A fallback chain that never returns an error to the user: local -> cloud -> second cloud option. Log which backend actually answered and why. Endpoint GET /llm/status.
- If system memory goes over a threshold, skip the local model for that turn. Use a short keep_alive in Ollama to free RAM.
Criteria: simple chat answers locally in about a second with the model already loaded; a complex query goes to the cloud; if I stop Ollama, it keeps answering; /llm/status shows the last decision.

PHASE 4: extra panels and voice tools
- A tools layer that intercepts the transcribed text BEFORE the LLM using patterns: "show me the weather", "my tasks", "close everything". If it matches, it opens the panel and answers with a short confirmation, without going through the LLM. CONTENT: event over WebSocket.
- A generic content panel system: image, markdown, weather (can be mocked at first), task list and conversation log.
- Interaction: drag, resize and close panels; scroll with the controller thumbstick and with pinch; save the layout in localStorage.
Criteria: I say "show me the weather" and the panel appears in under 2 seconds; I move a panel, reload, and it stays where I left it; I can scroll in MR without a mouse.

PHASE 5 (optional): wake word and delegation
- Wake word with voice detection in the browser.
- Delegate long tasks to {{CLOUD_LLM}} or to a coding agent in the background, and have the assistant tell me out loud when it's done.

KNOWN ISSUES ON THE QUEST: handle them from the start
- When entering an immersive session, the browser can suspend the AudioContext and the microphone stops capturing. After starting the session, call resume(), keep a silent buffer looping so it doesn't get suspended again and, if the microphone stream died, call getUserMedia again.
- When resizing a canvas used as a texture, recreate the CanvasTexture.
- Don't use createConicGradient (it isn't available in every Quest browser). Iframes with CSS3DRenderer don't show in immersive mode: plan a fallback.
- To debug on the headset, use the remote devtools (chrome://inspect with the headset connected) and leave prefixed logs, for example [audio].

What to expect

It's a starting point, not a finished Edison. The prompt leaves out things that are specific to how I work, like the panel that shows my agent sessions or Edison's personality, and keeps phase 5 optional. You'll most likely have to fix things along the way, especially the ones you only see with the headset on.

Did you build it?

If you build it, tell me how it went: what you named it, which phase got stuck or what you added. I'd love to see it. You can find me here:

Get new tutorials by email

I email you when I publish a tutorial or a video. No spam; unsubscribe anytime.