Skip to content
Kamply

Voice-over in Kamply

Text to speech that lands on real word boundaries

Kamply speaks the script with ElevenLabs, transcribes that take with Whisper, and derives every cut and caption from the transcript.

The longer description

One timed reel build took about an hour, in June 2026.

Free. Keep the files. Akash replies with times.

Finished work, built in the app.

Every piece here was built in Kamply at the size it ships. Sample brands are labelled as such.

NutriScanA plate drawn in one line. Reel · 9:16
NutriScanTuesday’s energy, as a forecast. Reel · 9:16
NutriScanYour year, counted. Year in review · 9:16 reel
NutriScanKnow what dinner costs. Product story · 9:16 reel

Sample brands from the app’s library

Darnfold, Boxkin, Woofstead and Latelift are invented brands, so every figure in their work is fictional.

sample brand
Woofstead · sample brandWhere the dog goes on fireworks night. Reel · 9:16
sample brand
Latelift · sample brandA coach note, rep by rep. Reel · 9:16
sample brand
Boxkin · sample brandA box changes hands. Reel · 9:16
sample brand
Darnfold · sample brandBack again, patched again. Reel · 9:16
  1. 01 Takes before timing

    The voice-over audio exists before anything is timed against it.

  2. 02 Timing from transcript

    Cuts and captions derive from a transcript of that audio, never hand-typed.

  3. 03 You watch the draft

    The whole video plays as a sequence and you sign it off before the master.

  4. 04 The file is checked

    The rendered MP4 itself is inspected, not the plan that produced it.

Hand-typed subtitle timings drift against the read. Deriving them from a transcript of the audio that actually rendered is why the captions land on the word. The last gate reads the finished file, so a video that renders wrong cannot pass by having had a correct plan.

What it actually does

Kamply is a Mac app that makes finished MP4 reels.

What it does, in detail

The voice-over is one step in that build. You approve a written script, Kamply renders the take with ElevenLabs, runs Whisper over the audio, and cuts the video to the timings that come back. Captions sit on the same word boundaries.

What this is not

The voice-over exists to time a video inside a Mac app.

What it leaves to you

There is no web box to paste a paragraph into, and no API.

How it runs, start to finish

Three steps, in the order the app actually does them.

  1. Lock the script

    The crew interviews you, then writes the script into a locked brief. You read the words before anything is spoken. Nothing gets voiced from an unapproved draft.

  2. Render the take

    ElevenLabs speaks the locked script. The audio file exists before the plan page renders, so every timing number downstream comes from a take you can already listen to.

  3. Derive every cut

    Whisper transcribes that take. Cut points and caption spans come from the transcript's word timings, so a line change re-derives the edit instead of drifting out of sync.

The numbers, and where they come from

Every figure on this page comes from Kamply's build log or a named published source. Nothing is modelled.

Voice engine
ElevenLabs
Timing source
Whisper transcript of the rendered take
Build time
About one hour, one timed build
Output
MP4 at 9:16, 1:1 or 4:5
Agency comparison
$300-$1,500 per video
What each number means
Voice engine
Built into the Mac app; the account is provisioned for you
Timing source
Preflight refuses hand-authored voice-over timing
Build time
One timed build in Kamply's build ledger, June 2026
Output
Finished file on disk; you upload it yourself
Agency comparison
Published Vidico and D-MAK production pricing

Questions people actually ask

If yours is not here, ask it in the walkthrough. The form below reaches a human.

Can I paste text and download just the audio?

Kamply is built around finished video, so the voice-over is a step inside a reel build rather than a standalone export screen. If all you want is an MP3 of a paragraph, a dedicated text-to-speech service is the shorter route.

How do captions stay in sync with the voice?

They come from the same source. Whisper transcribes the rendered ElevenLabs take, and both the caption spans and the cut points read their timings from that transcript. Nobody types a start and end time, so there is nothing to drift.

What happens if I change a line of the script?

The take is re-rendered and re-transcribed, and the edit is derived again from the new word timings. You are not nudging keyframes by hand. The gate that checks the render will fail if the timing was authored any other way.

Can I pick the voice?

Yes. Kamply ships with a default narrator and you can set a different ElevenLabs voice for a project. The voice choice is part of the brief you approve, alongside the script, the ratio and the look.

Does Kamply do avatars or lip-sync?

No. There is no talking head, no face swap, no digital presenter. The voice sits over footage, stock imagery, product shots and motion built in the app. If a person on camera is the point, this is the wrong tool.

Will it post the finished reel for me?

No. Kamply has no publishing, no scheduling and no social account connections. It writes a finished MP4 to your Mac at the ratio you picked, and you upload it wherever it is going.

Bring one brief to the walkthrough.

We build the first direction with your brand during the walkthrough.

Akash, who built Kamply, replies with times. The walkthrough is free: he builds your first campaign direction on your brand, and you leave with the files.