Product · AI video studio
A desktop video studio that turns an idea into a post-ready faceless video, rendered on the user’s own PC.
Introduction
Faceless channels run on volume: a script, a voiceover, footage that matches every line, captions, and an export — for every video, on every channel. Doing that by hand in a general-purpose editor is slow, and cloud tools add watermarks and render queues. ClipMesh is a desktop studio built for that exact job. AI writes the script and the voiceover, matches footage from the creator’s own library to each line of narration, and adds word-timed captions; a multi-track timeline handles the rest. Rendering happens on the user’s own GPU through a native Rust renderer, and the same renderer draws the preview and the export.
Details
The desktop app, rebuilt here in code from its own screens, in the order a video gets made. Pick a step, or click inside the window.
A topic goes in; tone, point of view, type and length shape what comes out. Script Health grades the hook before a word is recorded.
Recreated from the app’s interface. Labels and options are ClipMesh’s own; the projects, footage and numbers are samples.
Pieces of the editor and the studio that carry the most weight, at full size.
Captions
Speech recognition on the device times every word, and a preset decides how the active one pops. These are the app's own previews.
AI keys
A request uses the creator's own key first, then a platform key paid in credits, then a local engine that costs nothing.
Voiceover
On-device engines download their models the first time they are used. Cloud voices are there when one particular voice matters.
AI Match
An AI Director turns each line into a visual query, and the library answers with scored candidates, tiled so the voiceover is covered end to end.
Timeline
Text, image, shape, video, sound, caption and sticker, with linked audio and undo that covers the whole document.
Shortcuts
About two dozen shortcuts, J-K-L shuttle included.
Export
Four aspect ratios or a custom size up to 7680 pixels, five frame rates, and three quality levels.
Renderer
One Rust renderer draws both: compiled to WebAssembly for the preview, natively for the export. It tries the GPU first, and every hardware encoder is tested with a real one-frame encode before it is trusted.
A Rust renderer on wgpu with a CPU fallback, compiled to WASM for preview, so what users see is what they export, with no watermark on paid plans.
Script writer, voiceovers, captions from the clip’s own audio, and "AI Match" that builds a shot list from the voiceover and pulls matching footage from the user’s library.
Local engines for free, the user’s own API keys, or platform keys paid with credits — plus an MCP server so coding agents can assemble montages.
The running parts
The hard parts
The timeline is compiled into a plain numeric description of every frame, with animations baked to keyframes. The same Rust renderer reads it in both places: natively for export, as WebAssembly for preview.
The renderer tries the GPU and falls back to a CPU compositor. Each hardware video encoder is tested with a real one-frame encode before it is trusted, with software encoding as the last resort.
Export happens on the user’s own machine, so the plan cannot simply be trusted there. The API issues a signed grant, valid for five minutes, that carries the watermark settings; the Rust side verifies it and draws the watermark itself.
A request uses the user’s own key, then a platform key paid with credits, then a local engine. Whether the user can afford it is checked before the call and credits are taken after, with failover when a provider has a temporary error.
The video tools and AI models are not bundled. They download the first time they are needed, resume if interrupted, and fall back through mirrors.
Start here
Tell us what is slow, manual or missing. You will hear back within one working day — then get a written scope and a fixed quote, or an honest no.