Documentation
Smart Editor AI
Add word by word captions, texts, images, B-rolls, background music, sound effects and a voice-over to a video, then export it.
Smart Editor AI edits a video: the speech is transcribed word by word and shown as captions in the style you choose, with optional text overlays (Apple emojis included), images, B-rolls, background music, sound effects and a voice-over, then everything is exported as an MP4. In the API, a Smart Editor AI video is a project (/v1/captions, ids cap_…). The usual flow is to clean an ad with Smart Remover, then create a project from the cleaned video (job_id).
The exported video looks exactly like the preview of the editor: same fonts, sizes, line breaks, line spacing and boxes.
Price
0.5 credit per second of video (30 per minute), with a minimum of 5 credits. A 30 second video costs 15 credits.
- The credits are taken on the first export of a project (
402 insufficient_creditsif the balance is too low). - A failed export gives its credits back.
- Later exports of the same project (after edits) are free, up to 20 per project.
- Videos up to 15 minutes.
- Included from the Growth plan, in the app and through the API (
403 captions_not_allowedotherwise).
Lifecycle
| Status | Meaning |
|---|---|
transcribing |
The speech is being transcribed |
ready |
The transcript is ready; the project is not exported yet, or an export failed (see error). If the speech could not be transcribed, words is empty and error.code is transcription_failed: overlays still work |
rendering |
The video is being exported |
done |
The captioned video is ready in result.url |
failed |
Older projects whose transcription failed |
Create a project
POST /v1/captions
| Field | Type | Description |
|---|---|---|
job_id |
string | A succeeded job: captions are added to the cleaned video |
upload_id |
string | An upload created with POST /v1/uploads |
video_url |
string | A direct file link, Google Drive, Dropbox, TikTok, Instagram or YouTube link |
name |
string | Optional display name |
style |
string or object | A built-in style (classic, tv…), the id (cst_…) or name of one of your styles, or a style object, see below |
overlays |
object[] | Optional texts shown at a given time and place, see below |
images |
object[] | Optional images shown at given times, see below |
music |
object | Optional background music, see below |
sounds |
object[] | Optional sound effects at given times, see below |
brolls |
object[] | Optional B-roll videos shown full frame at given times, see below |
voiceover |
object | Optional voice-over (the captions are transcribed from it), see below |
original_voice |
boolean | false cuts the video's own voice |
auto_export |
boolean | Export as soon as the transcript is ready (default true). Set false to review or edit first |
Give exactly one of job_id, upload_id or video_url. The response (202) is the project, in transcribing status. Follow it with GET /v1/captions/{id} or the caption.succeeded and caption.failed webhooks.
curl https://pixamake.ai/v1/captions \
-H "Authorization: Bearer $PIXAMAKE_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 6f1d0c2a-3b7e-4f7a-9d52-0f3c9a1e8b44" \
-d '{
"job_id": "job_3kT9xQ2vLm8RfZp1Wq7a",
"style": {"base": "highlight", "highlight_color": "#FDE047", "words_per_line": 2},
"overlays": [{"text": "-30 % today 🔥", "start": 0, "end": 3, "y": 18, "style": "Brand yellow"}]
}'Styles
Built-in styles
| Id | Name | Look |
|---|---|---|
classic |
Classic | White text, discreet black outline, shadow |
clean |
Clean | Thin white text, no outline, small shadow, low in the frame |
soft |
Soft | White text with a very thin grey outline |
tv |
Banner | White text on a see-through black banner, like TV subtitles |
white |
White box | Black text on a white banner |
highlight |
Highlight | The spoken word turns yellow |
caps |
Capitals | Bebas Neue in capitals, 3 words per line |
cinema |
Cinema | Small whole sentences low in the frame, like at the cinema |
Built-in styles never change and cannot be edited: to change a setting, save your own style from one of them (base).
Your styles
Each workspace can save up to 50 named styles, in the editor or through the API, then use them wherever a style is expected, by id (cst_…) or by name.
| Method | Path | Purpose |
|---|---|---|
GET |
/v1/caption-styles |
Built-in styles, then saved styles |
POST |
/v1/caption-styles |
Save a style: { "name", "style" } |
GET |
/v1/caption-styles/{id} |
Read a style |
PATCH |
/v1/caption-styles/{id} |
Rename or change a style, immediately |
DELETE |
/v1/caption-styles/{id} |
Delete a style |
curl https://pixamake.ai/v1/caption-styles \
-H "Authorization: Bearer $PIXAMAKE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Brand yellow", "style": {"base": "highlight", "highlight_color": "#FFC300", "font": "poppins"}}'A project keeps a copy of the style it was given: changing a saved style does not change videos already set up. Send the style to the project again (PATCH /v1/captions/{id}) to use the new version. Names are unique in a workspace (409 name_taken otherwise).
Style object
Send a reference (built-in id, cst_… id or name), or an object with base (a reference, the project's current style by default) and only the fields to change. As soon as a field other than position differs, preset is custom.
The vertical position is a setting of each video: {"base": "tv", "position": 60} keeps the built-in tv style (and a saved style stays linked the same way), only placed higher.
| Field | Values |
|---|---|
base |
Style to start from: built-in id, cst_… id or name |
preset |
classic, clean, soft, tv, white, highlight, caps, cinema (in responses, custom for a changed style) |
font |
montserrat, poppins, anton, bebas, archivo |
size |
30 to 220 (pixels for a 1920 pixel tall frame) |
case |
normal, upper, lower |
grouping |
words: words_per_line words at a time; sentence: one sentence at a time (a pause of more than 0.6 s also starts a new caption) |
words_per_line |
1 to 8, with grouping: "words" (the editor offers 3 or 5) |
position |
5 to 95: vertical center, in % of the height |
color |
Text color, like #FFFFFF |
highlight |
How the spoken word stands out: color (the word changes colour) or none |
highlight_color |
Color of the spoken word |
outline, outline_color |
Outline width (0 to 14) and color around the text |
shadow |
Drop shadow (pop and highlight_text are accepted and ignored; highlight: "box" is treated as color) |
background, background_opacity |
Box behind each line of the caption, with lightly rounded corners (a color or null), and its opacity (0 to 100). The spoken word changes color inside it |
background_padding |
0 to 80: space above and below the text in its box (pixels for a 1920 pixel tall frame, default 16); the sides get 1.8 times more |
line_spacing |
0.8 to 2: distance between two lines of one caption, in multiples of the text size (default 1.1) |
Overlays
An overlay is a text shown at a given time and place (price, promo, hook). It takes the look of a style: font, colors, outline, banner and padding, shadow, case. Emojis are drawn as Apple emojis, in the app as in the exported video.
| Field | Description |
|---|---|
text |
Up to 200 characters, emojis included |
start, end |
Seconds |
y |
Vertical center of the text, in % of the height (default 20). Texts are centered across the frame |
size |
20 to 240 (default: the size of the style) |
style |
A style reference (built-in id, cst_… id or name) or a style object, like the project's style. Default: the project's caption style |
In responses, each overlay has style_id (the style it comes from, or custom) and style (the settings it uses).
Images
An image (product photo, logo, sticker) shown at given times, centered across the frame. Transparency is kept.
| Field | Description |
|---|---|
image_url |
Public link of the image (JPEG, PNG, WebP or GIF, 8 MB at most), copied when the project is saved. In responses, a link valid one hour |
id |
In PATCH: the id of an image already on the project, to keep its file without sending image_url again |
start, end |
Seconds |
y |
Vertical center, in % of the height (default 30) |
size |
Width, in % of the frame width (5 to 100, default 40) |
Up to 20 images per project. In responses, each image also has ratio (its height divided by its width).
Background music
A track from the Pixamake library plays under the sound of the video, looped if needed, lowered automatically while someone speaks, and faded out over the last 1.5 seconds. Music is included in the price.
| Field | Description |
|---|---|
auto |
true: Pixamake picks the track that best fits what is said (theme, mood, energy). This is the default when no track_id is given. On PATCH, auto: true asks for a new pick |
track_id |
A track of the library (mus_…), picked by hand |
volume |
0 to 100 (default 30) |
"music": null removes the music. With auto, the track is picked as soon as the transcript is ready; responses show it in music.track (title, artist, category, mood, tags and a preview link).
"music": { "auto": true, "volume": 25 }The library
GET /v1/music lists the tracks, the same for every workspace, newest first. Filters: category (hype, upbeat, chill, inspiring, emotional, dramatic, luxury, funny, corporate, lofi), q (text searched in the title, artist, genre, mood and tags) and limit (100 by default, 500 at most).
{
"object": "music_track",
"id": "mus_4fT8kQ2vLm9RzXp1Wq7a",
"title": "Golden Hour",
"artist": "Studio North",
"duration_seconds": 142.6,
"category": "upbeat",
"mood": "warm, confident",
"genre": "pop",
"energy": 4,
"bpm": 118,
"tags": ["upbeat", "pop", "claps", "instrumental", "beauty", "fashion"],
"description": "Bright pop groove with claps and a warm synth bass. Suits beauty and fashion launches.",
"preview_url": "https://...",
"is_new": true
}B-rolls
Videos shown over the video during a time range, above it and below the captions, texts and images (their own sound is cut, the soundtrack goes on). By default a B-roll fills the frame in its own format, centred, on black. Upload the video with POST /v1/uploads first. Up to 20 per project.
| Field | Description |
|---|---|
upload_id |
A video uploaded with POST /v1/uploads |
id |
In PATCH: the id of a B-roll already on the project, kept without uploading again |
start, end |
When it shows, in seconds (at most the length of the clip) |
trim |
Where the clip starts playing, in seconds into the clip (default 0) |
fit |
contain (default): the whole clip in its own format; cover: it fills the frame, its edges cut |
scale |
Size in % of the fitted size (10 to 400, default 100) |
x, y |
Centre of the clip, in % of the frame (default 50 and 50) |
background |
Black around the clip (default true); false shows the video around it |
In responses, each B-roll has name, length and a preview_url valid one hour.
Voice-over
A voice-over replaces or completes the video's voice, for instance when the video has none. When there is one, the captions are transcribed from it (again, in the background: the project goes back to transcribing).
| Field | Description |
|---|---|
voiceover.audio_url |
Public link of the audio file (MP3, M4A, WAV, AAC or OGG, 100 MB and 10 minutes at most) |
voiceover.start |
When it starts in the video, in seconds (default 0). Moving it moves its captions too |
voiceover.volume |
0 to 200 (default 100) |
original_voice |
false cuts the video's own voice (default true) |
"voiceover": null removes the voice-over; the captions are transcribed again from the video.
"voiceover": { "audio_url": "https://cdn.example.com/voice-over.mp3", "start": 0.5 },
"original_voice": falseSound effects
Sound effects played once at a given time, over the soundtrack, chosen from the Pixamake library. Up to 50 per project, included in the price.
| Field | Description |
|---|---|
sound_effect_id |
A sound effect of the library (mus_…, see below) |
id |
In PATCH: the id of a sound already on the project, kept as it is |
start |
When it starts, in seconds |
duration |
How long it plays (default: the whole effect) |
volume |
0 to 100 (default 80) |
In responses, each sound has name and a preview_url valid three hours.
GET /v1/sound-effects lists the library, the same for every workspace, with the same filters as /v1/music: category (whoosh, pop, click, impact, riser, notification, cash, applause, laugh, magic, transition, ambience), q and limit.
"sounds": [{ "sound_effect_id": "mus_9pQ2xT8kLm3RzWv1Yc6b", "start": 0.8 }, { "sound_effect_id": "mus_2hT7kQ9vLm4RzXp8Wc1a", "start": 8.5, "volume": 60 }]Safe zone
Texts and images never overlap each other or the spoken captions. Around each of them, a margin of 2 % of the frame height stays free while they are shown at the same time. When an element would touch another one, the export moves it to the nearest free height (texts first, in their order, then images); the editor shows the same result, and dragging an element stops at the edge of the others' margins. The y you send is kept as asked: only the export and the preview move the element.
Edit and export again
PATCH /v1/captions/{id} changes words (each { "text", "start", "end" }, plus "line": "start" to start a caption at this word or "line": "join" to keep it in the caption before), style, overlays, images, sounds, brolls (each list replaces the current one), music, voiceover or original_voice. Then POST /v1/captions/{id}/export burns the changes into a new file. With auto_export: false, call the same export endpoint when you are ready.
GET /v1/captions/{id}/srt returns the captions as an SRT file.
The project object
{
"id": "cap_8LmQ2vT9xK3rWp7aZn4c",
"object": "caption_project",
"status": "done",
"name": "spring-ad.mp4",
"source_job_id": "job_3kT9xQ2vLm8RfZp1Wq7a",
"input": { "duration_seconds": 28.4, "width": 1080, "height": 1920 },
"language": "en",
"credits": 29,
"charged": true,
"exports": 1,
"words": [{ "text": "This", "start": 0.12, "end": 0.31 }],
"style": { "preset": "custom", "font": "montserrat", "words_per_line": 2, "highlight": "color", "highlight_color": "#FDE047", "background": null },
"overlays": [{ "id": "ov_x1", "text": "-30 % today 🔥", "start": 0, "end": 3, "y": 18, "size": 80, "style_id": "cst_8LmQ2vT9xK3rWp7aZn4c", "style": { "font": "poppins", "color": "#FFFFFF" } }],
"music": { "auto": true, "track_id": "mus_4fT8kQ2vLm9RzXp1Wq7a", "volume": 30, "track": { "id": "mus_4fT8kQ2vLm9RzXp1Wq7a", "title": "Golden Hour", "artist": "Studio North", "duration_seconds": 142.6, "category": "upbeat", "mood": "warm, confident", "tags": ["upbeat", "pop"], "preview_url": "https://..." } },
"sounds": [{ "id": "snd_x1", "sound_effect_id": "mus_9pQ2xT8kLm3RzWv1Yc6b", "name": "Pop", "start": 0.8, "duration": 0.4, "volume": 80, "preview_url": "https://..." }],
"brolls": [],
"voiceover": null,
"original_voice": true,
"images": [{ "id": "img_x1", "image_url": "https://...", "start": 4, "end": 7, "y": 30, "size": 40, "ratio": 1 }],
"result": { "url": "https://...", "expires_at": "2026-10-06T10:00:00.000Z" },
"srt_url": "https://pixamake.ai/v1/captions/cap_8LmQ2vT9xK3rWp7aZn4c/srt",
"error": null,
"created_at": "2026-10-06T08:58:12.000Z",
"exported_at": "2026-10-06T09:00:41.000Z"
}result.url expires after one hour: download it, or fetch the project again for a fresh link. error.code is one of transcription_failed, insufficient_credits, no_speech or export_failed.