AI dubbing API for video in 135+ languages

Upload your video, choose a target language, collect the dubbed video. One AI dubbing API request handles transcription, translation, speaker detection, voice cloning and speech synthesis, can keep the voice identity of the original speaker, and returns a dubbed video, separate WAV audio tracks and subtitle files.
Voice cloning
Lip-sync built in
REST + Python SDK
three requests, start to finish
OAuth 2.0
# 1. exchange client credentials for a token (valid 1 hour)
curl -X POST https://rask-prod.auth.us-east-2.amazoncognito.com/oauth2/token \
  -u "$CLIENT_ID:$CLIENT_SECRET" \
  -d "grant_type=client_credentials"# 2. upload by link - YouTube, Drive, S3, Vimeo or direct URL
curl -X POST https://api.rask.ai/api/library/v1/media/link \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"link": "https://youtu.be/...", "name": "Q3 launch"}'# 3. create the dubbing project - this starts the dub
curl -X POST https://api.rask.ai/v2/projects \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "video_id": "b8f7835d-e8c9-4ff4-8e34-5c336d818319",
    "dst_lang": "es-mx",
    "name":     "Q3 launch"
  }'
GET /v2/projects/{project_id}
polling, 10s
# Poll the project until status is merging_done.while :; do
  curl -s https://api.rask.ai/v2/projects/$ID \
    -H "Authorization: Bearer $TOKEN" | jq .status
  sleep 10done
Used by global teams
Capabilities

Everything the dubbing API does, in one call

Automatically transcribe, translate and dub your videos through one API. Fine-tune the results with voice selection, glossaries and script editing. Add optional lip-sync, billed separately.

Dubbing into 135+ languages

The full pipeline behind one request: transcribe, translate, synthesise, mix. A video and a target language are all you send.

Speaker detection and per-speaker voices

Speakers are separated from the source audio automatically and each gets its own voice in the target language. Override any assignment, or pass the count if you know it.

Voice cloning across languages

Keep the voice identity of the original speaker in the dub, so the presenter still sounds like the presenter in Spanish, Hindi, Japanese or Korean.

Lip-sync

Mouth movement matched to the dubbed audio, as a separate task on any dubbed project. Face detection runs first so you only pay where it will work.

Background audio preserved

Music, room tone and effects are separated from speech and mixed back under the dub. The result is a fully localized track with the original music, room tone and effects under the new voice.

Subtitles in .srt and .vtt

Both formats come with every finished project, timed to the dub and ready for your player.

Terminology control with glossaries

Brand names, product names and technical terms translate the same way every time, or stay untranslated on purpose. Versioned, so every finished project keeps the glossary it was dubbed with.

Segment-level script editing

Fetch the transcript, fix a line, adjust a timestamp, add or delete a segment, then regenerate. Your linguists correct the script directly, and the dub regenerates from their edits.

Scale without limits

Send your full catalogue at once and dub into every target language in parallel, at any volume.
See it in action

What the AI dubbing API actually returns

The same clip, dubbed by the same API call you just read. Original voice on the source, the cloned voice in French, German, Portuguese and Spanish, background audio intact.
EN flag
En
FR flag
Fr
DE flag
De
PT flag
Pt
ES flag
Es
Workflow

Three steps from source file to dubbed video

Creating the project starts the dub. Use /generate later to re-dub a project after you edit it.
1

Upload your video

Send a file as multipart form data, or post a link. Supported link sources are YouTube, Google Drive, S3, Vimeo and any direct download URL. Returns a media id and full technical metadata before you spend anything.

POST /api/library/v1/media/link
2

Create the project

Pass the media id as video_id and a target language. Optionally attach your own transcript or a glossary. The dub starts immediately.

POST /v2/projects
3

Collect the output

Read status until it reaches merging_done, then download the dubbed video, the audio tracks and the subtitle files.

GET /v2/projects/{project_id}
Processing time varies by project. Track progress through the API and download your files when dubbing is complete.
Built around you

Not on the list? We build it.

New languages, custom voices, your own delivery method, higher throughput: when a dubbing integration needs something outside the public catalogue, our team scopes it with you and builds it.

Languages and voices

Need a language or voice outside the public catalogue? Talk to our team about your requirements and the available options.

Delivery and integration

Webhooks, push into your own queue, delivery straight to your S3, custom endpoints. Tell us how your pipeline expects to be fed.

Volume and throughput

Batch sizes, target languages per action and processing priority are all set in your contract to match your pipeline.

Security and hosting

Data residency, SSO and SAML, retention rules and security review support are handled as part of enterprise onboarding.
Use cases

Who uses an AI dubbing API

Platforms with video inside

LMS and customer-education platforms, creator and video platforms, webinar and virtual event tools, video hosting and workflow SaaS, digital adoption platforms, live shopping. Educational content, product demos and creator videos are dubbed inside your product, so global audiences hear them in their own language and AI dubbing becomes one of your features.

Localization and production agencies

Connect AI dubbing to the delivery pipeline you already run and keep the human review step. The AI produces the first dubbed version, and your linguists edit the script instead of booking a studio to record it again.

In-house continuous video production

International retail and e-commerce, global franchises, media publishers, corporate universities. Build localization into your publishing workflow, so your team can regenerate language versions whenever the source video changes.
Customer stories

Real teams, measurable results

How companies localize video with Rask, in their own numbers.
arrow
arrow
Code

AI dubbing API code examples

One complete round trip in each language: authenticate, upload, dub, wait, download.
# pip install git+https://github.com/braskai/rask-sdk
import asyncio
from rask_sdk import RaskSDKClient

async def main():
    client = RaskSDKClient(client_id=CLIENT_ID, client_secret=CLIENT_SECRET)

    media   = await client.upload_media_link(link="https://youtu.be/...")
    project = await client.create_project(video_id=media.id, dst_lang="es-mx")

    while project.status != "merging_done":
        await asyncio.sleep(10)
        project = await client.get_project(project.id)

    print(project.translated_video, project.translation_srt_path)

asyncio.run(main())
github.com/braskai/rask-sdk
Python 3.8+, async
import os, time, requests

AUTH = "https://rask-prod.auth.us-east-2.amazoncognito.com/oauth2/token"
API  = "https://api.rask.ai"

token = requests.post(AUTH, auth=(CLIENT_ID, CLIENT_SECRET),
                      data={"grant_type": "client_credentials"}).json()["access_token"]
h = {"Authorization": f"Bearer {token}"}

media = requests.post(f"{API}/api/library/v1/media/link", headers=h,
                      json={"link": "https://youtu.be/..."}).json()

project = requests.post(f"{API}/v2/projects", headers=h,
                        json={"video_id": media["id"], "dst_lang": "es-mx"}).json()

while project["status"] != "merging_done":
    time.sleep(10)
    project = requests.get(f"{API}/v2/projects/{project['id']}", headers=h).json()

print(project["translated_video"])
requests + polling
no SDK required
// no dependencies, Node 18+ global fetch
const API = 'https://api.rask.ai';

const auth = await fetch('https://rask-prod.auth.us-east-2.amazoncognito.com/oauth2/token', {
  method: 'POST',
  headers: { Authorization: 'Basic ' + btoa(`${CLIENT_ID}:${CLIENT_SECRET}`) },
  body: new URLSearchParams({ grant_type: 'client_credentials' }),
}).then(r => r.json());

const h = { Authorization: `Bearer ${auth.access_token}`, 'Content-Type': 'application/json' };

const media = await fetch(`${API}/api/library/v1/media/link`, {
  method: 'POST', headers: h,
  body: JSON.stringify({ link: 'https://youtu.be/...' }),
}).then(r => r.json());

let project = await fetch(`${API}/v2/projects`, {
  method: 'POST', headers: h,
  body: JSON.stringify({ video_id: media.id, dst_lang: 'es-mx' }),
}).then(r => r.json());

while (project.status !== 'merging_done') {
  await new Promise(r => setTimeout(r, 10000));
  project = await fetch(`${API}/v2/projects/${project.id}`, { headers: h }).then(r => r.json());
}

console.log(project.translated_video);
fetch, Node 18+
plain HTTP
# 1. token (valid 1 hour - cache and reuse it)
TOKEN=$(curl -s -X POST https://rask-prod.auth.us-east-2.amazoncognito.com/oauth2/token \
  -u "$CLIENT_ID:$CLIENT_SECRET" \
  -d "grant_type=client_credentials" | jq -r .access_token)

# 2. upload by link
VIDEO_ID=$(curl -s -X POST https://api.rask.ai/api/library/v1/media/link \
  -H "Authorization: Bearer $TOKEN" \
  -d '{"link": "https://youtu.be/..."}' | jq -r .id)

# 3. create the dubbing project
ID=$(curl -s -X POST https://api.rask.ai/v2/projects \
  -H "Authorization: Bearer $TOKEN" \
  -d "{\"video_id\": \"$VIDEO_ID\", \"dst_lang\": \"es-mx\"}" | jq -r .id)

# 4. poll, then download
while [ "$(curl -s https://api.rask.ai/v2/projects/$ID -H "Authorization: Bearer $TOKEN" | jq -r .status)" != "merging_done" ]; do
  sleep 10
done
POST /v2/projects
OAuth 2.0 client credentials
<?php
use GuzzleHttp\Client;

$http = new Client(['base_uri' => 'https://api.rask.ai']);

$auth = json_decode((new Client)->post(
    'https://rask-prod.auth.us-east-2.amazoncognito.com/oauth2/token',
    ['auth' => [$clientId, $clientSecret],
     'form_params' => ['grant_type' => 'client_credentials']]
)->getBody(), true);

$h = ['Authorization' => 'Bearer ' . $auth['access_token']];

$media = json_decode($http->post('/api/library/v1/media/link',
    ['headers' => $h, 'json' => ['link' => 'https://youtu.be/...']])->getBody(), true);

$project = json_decode($http->post('/v2/projects',
    ['headers' => $h, 'json' => ['video_id' => $media['id'], 'dst_lang' => 'es-mx']])->getBody(), true);

while ($project['status'] !== 'merging_done') {
    sleep(10);
    $project = json_decode($http->get("/v2/projects/{$project['id']}",
        ['headers' => $h])->getBody(), true);
}
Guzzle 7
plain HTTP
Python is the official SDK. The Node, PHP and cURL tabs are plain HTTP examples against the same endpoints.
Progress

How a dubbing project reports progress

A project moves through a documented state machine, so you can show your own users a real, staged progress bar. You only need to act on three groups.

Success, stop here

merging_done
Dubbing is complete. Fetch the project once more to retrieve the download links available for your project.

In progress, keep polling

created
uploading
transcription_started
separate_background_started
determine_speakers_started
translation_started
voiceover_started
merging_started
Plus the matching _done states. These map cleanly onto a staged progress UI.

Failure, stop and report

failed
upload_failed
transcription_failed
translation_failed
voiceover_failed
merging_failed
no_audio
no_words
forbidden_link
no_audio, no_words and forbidden_link are specific enough to show directly to your own user.
Typical run, 4 minute source video
GET /v2/projects/{project_id}
   0s  created
  12s  transcription_started
  70s  translation_started
  95s  voiceover_started
 150s  merging_started
 170s  merging_done

Or skip polling entirely

Point us at a webhook endpoint and every state transition arrives as a POST, so your queue reacts to each change the moment it happens. Webhook delivery is part of enterprise onboarding.
Read the docs
preserveAspectRatio="xMidYMid meet" aria-hidden="true" role="img">
Output

Dubbed video, audio and subtitles

Download your dubbed video and translated subtitles in SRT and VTT. On eligible plans, you can also download the translated voice on its own or mixed with the original background audio.
Field
Format
What it is
translated_video
MP4
The dubbed video, ready to publish
voiceover
WAV
Translated speech only, no background, drop it into your own mix
translated_audio
WAV
Translated speech mixed with the original background audio
translation_srt_path
.srt
Subtitles in the target language
translation_vtt_path
.vtt
Same subtitles as WebVTT, for web players
original_video
MP4
Your source file as stored, for reference
Control

Edit the script, then re-dub only what changed

Automatic translation gets you most of the way. When it does not, you can edit the transcript segment by segment and regenerate, and you are only charged for the parts you touched.

Segment-level editing

Segment-level editing
Every project exposes its transcript as addressable segments, so fixing a bad line takes one patch request.
Fetch the transcript and its segments
Patch the source text and we re-translate it
Patch the translated text directly to keep your own wording
Adjust timestamps, add segments, delete segments
Reassign the voice attached to any speaker

Glossaries for terminology

Glossaries for terminology
Attach a glossary to a project and your brand names, product names and technical terms translate the same way every time, or stay untranslated on purpose.

A glossary covers one language pair and applies across all regional variants of that target language, so one Spanish glossary serves es-mx and es-es alike. The project records glossary_version, so later edits never silently change finished work.
Input

Supported file formats

Seven video and four audio formats, plus ingestion straight from a URL. Post an unsupported file and the API rejects it by content type before any minutes are charged.

Video

7 formats
.mp4
.mov
.mkv
.avi
.webm
.m4v
.3gp

Audio

4 formats
.mp3
.wav
.flac
.m4a

Metadata before you spend

The media endpoint returns duration, file size, video codec, frame rate, resolution, audio codec, sample rate and channel layout. Validate a file and estimate cost in advance, because dubbing is billed by the minute.
Security and compliance

Cleared by the people who have to ask

Brask Inc holds SOC 2 Type II and operates under GDPR, and publishes live control status in its Trust Center.

Certifications

SOC 2 Type II
GDPR
CAI and C2PA member

Encryption

TLS 1.2 in transit
AES-256 at rest
SSE-S3 managed keys
Key rotation at least every 12 months

Your content

Closed AI environment
Never used to train general models
Full export, no lock-in

Access

SSO and SAML
Role-based access control
OAuth 2.0 client credentials

Enterprise

Data residency options
Custom SLAs
Dedicated account manager

Voice rights

Synthetic voices with usage rights
Consent-based voice cloning
AI safety controls, human in the loop
Comparison

Rask API vs building it yourself

The same dubbing capability, assembled two ways. One column is a contract; the other is a roadmap.
What you need
Rask API
Assembled from parts
Transcription, translation, synthesis, mixing
One call
Three to four vendors and the glue between them
Speaker detection and per-speaker voices
Included
A diarization vendor plus your own voice-assignment logic
Background music and effects preserved
Included
A source-separation step you build and tune
Voice cloning across languages
Included
A separate vendor and a separate consent flow
Lip-sync
Included
A fifth vendor
Subtitles in .srt and .vtt
Included
Generated from your own timing data
Fixing one bad line
Re-dub that segment, billed for that segment
Re-run the pipeline, pay for the file
Terminology consistency
Versioned glossaries
Pre and post-processing you maintain
Rate limits
None
Whatever your weakest vendor enforces
A language nobody supports yet
We build it
Not available
FAQ

AI dubbing API: common questions

Can't find what you're looking for? Contact sales.

How do I get API access and an API key?

How do I authenticate with the AI dubbing API?

Is the API synchronous or asynchronous?

What are the rate limits?

How many languages can I dub into?

How does voice cloning work with the API?

How long does AI dubbing take?

What file formats are supported?

How does the API handle a video with several speakers?

Can I supply my own script or transcript?

Do I get subtitles as well as a dubbed video?

Is there an SDK?

What does it cost to re-dub after an edit?

Start dubbing in one afternoon

Make your first request today, and talk to us when you need volume terms or a language nobody else has.