How I cloned my voice with AI, and why it only says my words
My narrated shorts use an AI copy of my voice, made with Artlist and now ElevenLabs. It only reads lines I said in my films and guides.
I make long films about hiking alone, mostly for YouTube: about 250 over ten years. This fall the big project is giving them a second life as short vertical videos and as narrated versions of films that never had a voice, with Claude, an AI model made by Anthropic, and KraigOS, the system Claude and I built to run my work. A lot of that needs my voice. So I cloned it.
Why an AI voice at all
Many of my films have no narration, just the trail and music. Short videos need more context than that. When I reviewed the first animated route-map videos on October 5, I asked for two versions of each: one ambient, and one with a "Kraig VO talking with more context."
Recording it myself every time doesn't scale. In July I narrated a Peaks of the Balkans film the old way, over five rounds of takes, and as I said then, "I need to be fresh to record this entire script with inflection and passion in my voice." That works for one film, not for dozens of shorts whose lines change with every cut. With a clone, Claude can re-voice one line after a late edit, and keep each line as its own file so a better voice, or my real one, can be swapped in later.
Three copies of my voice

- My first Artlist clones. Artlist, the music and stock footage library I use, has an AI voice tool, and I'd made a few clones there myself, months ago. The one we used most, "Kraig 5," voiced the first public post with my AI voice, a Kyrgyzstan guide reel on September 27. When we started leaning on it, I told Claude: "I think I made that like months ago. So if there's like a better way to do it, we can do that."
- KraigVO. On September 29 I read part of a script on my Rode shotgun mic, and Claude used a clean 3-minute cut of it to make a fresh Artlist clone, KraigVO. It narrates the route-map videos.
- The ElevenLabs clone. On September 29 I said: "I would like to make it absolutely flawless and like no way that anyone could tell that it's made by AI and generated." That's what ElevenLabs' Professional Voice Clone is for, trained on far more of my audio. On October 5 I made a rule: no new voiceover work until it's set up.
How the ElevenLabs clone was trained
ElevenLabs asks for at least 30 minutes of clean audio for a Professional clone. Instead of reading a long script, Claude went looking for audio I'd already recorded.
- Same mic only. Different mics sound different, so my rule was Rode audio only. Claude measured each recording (room noise, tone, loudness) and kept only what matched.
- What was left. About 31 minutes of narration takes and reads, just over the minimum.
- Cleaned, not polished. Retakes and false starts cut with the help of Whisper transcripts; breaths kept; no noise reduction or effects.
- Uploaded and checked. Claude uploaded it on October 9, and I read ElevenLabs' short live sentence to confirm it's my voice. The clone was ready that afternoon.
The name has a "2" in it because of an idea I had that day: "the voice clones should be trained on the same microphone, right? ... it should be like Kraig voice clone one, DJI mic, Kraig voice clone two, Rode video mic." The DJI one isn't made yet.
The rule: it only says what I said
The AI voice doesn't get to make things up.

- My own words. What I said in a film's narration and what's in that route's hiking guide, itself built only from my films, GPS tracks and notes. Gemini, Google's AI model, has already transcribed every moment I speak in my films.
- A script with sources. Claude trims lines only to join them up: no new claims, no adjectives I didn't use. Every line lists its source, with the film's timestamp and my exact quote.
- The voice reads it, one file per line, with place names respelled so it says them right: on K2, Askole became "Askoli."
- Captions and video. Whisper times every word for the captions, and Claude builds the video.
Here's what that looks like in the K2 Base Camp map video. Every line under the frames is something I said in the film:

Longer pieces follow the same rule. For a narrated remake of my silent Hornstrandir film, Claude pulled lines from my guide video, my guide, a podcast and a vlog. Every feeling in it, worried, super pumped, lonely and bored, is a word I used, and any small inference is marked for me to confirm. And since Claude can't hear audio, every review starts with me listening.
How it's gone so far
- The first long one. I watched the Hornstrandir remake all the way through on a walk. My reaction: "I'm super impressed. It's really good... 90% of this is like great, fantastic." It went up on Facebook on October 5.
- The K2 map video is the narrated one people have seen most. On Instagram it passed 120,000 views in about two days; it's also on YouTube Shorts, with YouTube's AI label, and TikTok, where Acatenango and others followed.
- Not perfect. The Artlist voice reads fast, and unusual place names still need respelling.
- Not for everything. Some sponsors want my real voice, and those get it.
That's also why this article exists. On October 8 I decided I want to be "more open with how we're using AI" here, so this is the whole thing.
What's next
The ElevenLabs clone hasn't read anything yet. Next is a blind test: the same lines from the new clone, KraigVO and a real recording of me, to see whether anyone can pick out the real one. If it passes, the narrated videos can be re-voiced line by line, and I'll decide whether to keep it.
FAQ
Is the narration in your videos AI?
In my narrated short videos and the narrated Hornstrandir remake, yes: it's an AI copy of my voice.
Which tools do you use to clone your voice?
Artlist's AI voice tool, running MiniMax Speech 02 HD, since September 2026, and an ElevenLabs Professional Voice Clone trained on October 9, 2026. Claude writes the scripts and builds the videos.
Does the AI voice say things you never said?
No. Every line comes from my film narration, my guides or my own recordings, trimmed only to join them, and every line in the script lists its source. Any small inference is marked for me to check.
How much audio does a voice clone need?
ElevenLabs asks for at least 30 minutes for a Professional clone and suggests one to two hours. Mine was trained on about 31 minutes of my own narration, all from one microphone.
Can people tell it's AI?
Some can. The goal is a voice no one can tell from mine, which is what the blind test is for.