Uruvam · Digital humans

A five-minute phone video becomes a face your agent can wear

Uruvam reconstructs a photoreal 3D digital human from ordinary phone footage, then ships it as a 12 MB browser asset that animates on the user's own device — so an AI agent gets a face without a render farm behind it.

How it works

Capture once, animate everywhere

STEP 01

Capture on a phone

Five minutes of ordinary phone video is the whole input. No rig, no studio, no depth sensor.

STEP 02

Build the face

The capture is reconstructed into a photoreal 3D head and compressed into a single 12 MB asset a browser can load.

STEP 03

Ship it once

The asset transfers on first load and is then cached. Every later session reuses the face already on the device.

STEP 04

Animate from speech

The server sends only the agent's speech. The device drives the face from that audio, so nothing has to be rendered upstream.

Why it is different

The server sends speech, never video

That single decision is what removes the render GPU, the encode delay and the per-session cost.

Sub-second responses

Nothing waits on a video encode or a render queue. Audio arrives and the face is already moving on the client.

Zero render GPU

There is no per-session GPU to provision or pay for. Cost per concurrent user is the cost of streaming speech.

Video never leaves the device

Because the server transmits speech rather than frames, no rendered likeness of the user is sent over the wire during a session.

Compared with streamed avatars

Where the work happens changes everything downstream

 UruvamServer-rendered avatar
What the server sends per turnSpeech audioRendered video frames
GPU per concurrent sessionNoneOne render GPU
Response latencySub-secondEncode + network bound
Cost as users scaleBandwidth for audioLinear GPU cost
Capture requirement5 minutes of phone videoStudio session or rig

Put a face on your agent

Send us five minutes of phone video and we will show you the face running in your own browser.