CLI › Reference

doc.json

The editable document: zoom, segments, speed, tilt, audio, overlays and frame.

doc.json is the editable document: every effect the video carries, as plain JSON. Edit it and re-render; nothing re-runs the browser. The full JSON Schema ships in the npm package at schema/doc.schema.json, and vos validate <dir> lints the semantics before you spend a render.

A doc.json is one of two documents, told apart by its anchor. A recording document (a take) carries source and everything on this page. A program document carries program and the shared layers only: overlays, objects, audio, speed, export, plus program.tweenEdits (retimes over the config's recorded tweens, keyed by spec index) and program.duration (the program's own length when the config's is a placeholder). In a program directory doc.json omits program.config, because config.json is the config; a bare program with no layers has no document at all. The reference tables list the program document's own fields under program.

Camera

JSON
"zoom": [{ "id": "z1", "in": 4.0, "out": 8.5, "level": 2.2,
           "cx": 0.62, "cy": 0.4, "source": "manual" }]
  • level is 1 to 5; 1.4 to 2.8 reads well, and anything past 3 is for genuine detail, briefly.
  • A camera span (zoom, speed, tilt) may carry "anchor": { "step": "cta", "offset": -0.2 }, tying it to an actions.json step so vos plan --reuse re-times it exactly onto a re-recording. in/out stay authoritative; rendering never reads it.
  • cx/cy are the focus point as normalized fractions of the video frame, 0 to 1, where 0.5, 0.5 is center. Never pixels; pixel values are the single most common mistake.
  • focusMode: "auto" makes the camera follow the cursor through the span.
  • source is the wand contract: "auto" spans are planner suggestions and regenerate on re-plan; set "manual" on any span you add or edit and it is never touched. Auto ids say their origin: z0 and up are click clusters, k0 and up are typing sessions (a frame held on the field being typed into), d0 and up are dwells.
  • A deleted proposal stays deleted. Removing an "auto" span in the studio writes rejected: [{ id, lane, in, out }] at the top of the document (the lane and the source extent of what you removed, its step anchor with it), and no re-plan proposes that beat again: vos plan, plan --reuse on a re-recording, the studio's own re-plans. Write one by hand the same way when you delete an auto span in doc.json; a manual span never needs it.
  • transition sets the speed of a span's ramps as named steps: instant | fast | smooth | slow, where smooth (the default) is the stock motion and instant is a hard cut. Zoom, tilt and camMotion spans all take it; the same control sits in each span's panel.

zoomStyle names the camera personality: glide (default) | focus | cinema | snappy | cut | none. Every style times the tilt track to its own zoom tempo, so a lean lands with its zoom; how far the card leans is tiltStyle. keynote and drift are retired names that still read as glide plus a medium tilt and cinema plus a subtle one.

Trims and pacing

JSON
"segments": [{ "in": 0.8, "out": 14.2 }],
"speed":    [{ "id": "s1", "in": 6, "out": 9, "rate": 2 }],
"freeze":   [{ "id": "f0", "at": 11.5, "seconds": 2 }]

segments are the kept footage spans; cut dead time here. speed rates run 0.1 to 16; 2x through typing is the classic use. freeze holds the frame just before a source moment for these output seconds, anywhere in the take: a pause on a state while a caption reads, a beat before a cut, the rest a poster ends on. Both are footage-anchored (they follow their frames through trims). The still, the frame that stands for the take (the shelf's cover, the frame a kit's card destinations render), is still in output seconds when you set it (drag the studio's Cover mark on the time ruler, vos plan --still <t>), else where the take's own last freeze begins (a template's freeze never counts), else the hero moment after the card and the opening words have entered. The card is on screen exactly while its clip runs, freezes included, and gone after it: clips that outlast the footage play over the ground alone, so a card under an end card is a freeze of its last frame.

Many media

JSON
"media":    [{ "id": "m1", "videoKey": "media/second.webm", "cursor": [], "meta": { "durationMs": 6000, "width": 1600, "height": 900 } }],
"segments": [{ "in": 0.8, "out": 14.2 }, { "in": 0, "out": 6, "media": "m1" }],
"zoom":     [{ "id": "u1", "in": 1, "out": 3, "level": 1.6, "cx": 0.5, "cy": 0.5, "media": "m1" }]

A take may hold several recordings or uploads: media lists the others, each the source shape with an id, and a segment on one plays it wherever it sits in the sequence. A source-anchored span names the media its seconds belong to; absent means the primary, source. A recording is media plus the facts only a capture has, so a file with no cursor track gets no dot and no auto-zoom. A media wears its own card: media[].frame carries the card-owned fields (inset for its placement and size, browserBar, radius, shadow, shadowContact, shadowColor, border, borderWidth, borderColor, fit, focus, focusFollow) over the take's frame while it plays. Absent fields fall to the take's; the bar names the media's recorded page, and an upload with no page wears none. The frame-wide fields (the aspect, the padding, the ground, the backdrop, the card's anim) stay the take's. Many cards: an image or video overlay may show a document media by reference ("key": "media:m1", media: alone the primary), with its cursor track and recorded page, and wear a card (frame: browserBar, a lean in degrees, shadow, shadowContact, shadowColor, cursor), drawn on its own plane above the primary card with the media's cursor and clicks inside it; a layer without frame is the flat picture. The primary card stays primary. The primary's sound plays at its own moments; another media's audio is not spliced into the cut yet.

A footage clip meets the clip beside it by a TRANSITION: segments[].anim carries exit (how it leaves at the boundary after it) and enter (how it arrives at the one before), by slide, fade, scale or none; a slide names its side (left | right | up | down: where an enter comes from, where an exit goes to; in from the right and out to the left by default), a step its seconds (0.15 to 2, never longer than half the shorter clip). An exit and the next clip's enter at one boundary run together as a push: the incoming clip plays live while the outgoing, frozen on its last frame, moves away; the camera rests through the window and one cursor dot crosses from the outgoing card's last point to the incoming's. The first clip's enter and the last clip's exit are the card's own (frame.anim). A page change inside one recording is the same thing: two clips of one media with the load between them cut, which vos plan proposes at every step the recorder marked navigated.

Tilt

JSON
"tilt": [{ "id": "u0", "in": 4, "out": 8, "rx": 6, "ry": -9, "source": "manual" }]

Source-anchored, non-overlapping regions where the 3D card leans to a pose and returns to flat rest. rx/ry are degrees (±45 hard, ±5 to 18 reads premium): positive rx brings the top edge toward the camera, positive ry the left edge, so leaning toward a right-side focus takes negative ry. Tilt is punctuation: at most one pose change per 5 second beat, paired with zoom moments. tiltStyle (off | subtle | medium | strong) records the Dynamic-tilt wand intensity that derives auto spans from the zoom plan.

Webcam

JSON
"cam": { "visible": true, "position": "bottom-left", "size": 0.25,
         "shape": "circle", "mirror": true },
"camMotion": [{ "id": "m1", "in": 3, "out": 8,
                "x": 0.5, "y": 0.55, "size": 0.45 }]

cam is the bubble's rest pose (it only applies when the take has a cam track): x/y place the bubble center as frame fractions and win over the position corner anchor, size is the diameter as a fraction of frame height, and window is a source-time visibility span.

camMotion animates the layout: source-anchored, non-overlapping spans where the bubble holds a pose and morphs there over about 0.65 seconds, settled at the span start. Outside spans it returns to the rest pose; spans close together morph pose-to-pose without returning to rest. Absent pose fields inherit the rest pose, so a span can move without resizing. The signature move is a bubble parked small in a corner that comes front-and-center (x: 0.5, y: 0.55, size: 0.45) while you talk to camera, then returns on its own.

Audio

JSON
"audio": [{ "id": "a1", "key": "https://assets.vos.so/music/ember-glow.mp3",
            "start": 0, "gain": 0.5, "duck": true, "loop": true }]

Output-anchored music and SFX, mixed and muxed on full renders. key is a file dropped into the take dir ("/music.mp3") or a hosted catalog track; the CC0 catalog with slugs, moods and durations is at GET vos.so/api/music. A music bed: start 0, gain around 0.5, duck: true under a voiced recording, loop to fill. micGain and systemGain (0 to 1) are the voice and system-audio master gains on recordings that carry them.

Frame and background

JSON
"frame": { "padding": 96, "radius": 12, "shadow": 0.5,
           "browserBar": { "kind": "mac" },
           "backgroundMedia": { "kind": "video",
             "key": "https://assets.vos.so/backgrounds/ember-drift-1080p.webm",
             "duration": 10, "dim": 0.2 } }

The card chrome: padding, corner radius, shadow, border, aspect ratio and the browser-bar mock. backgroundMedia puts a vos animation or a still behind the card; a video needs duration (its loop length; time is output-anchored modulo it, so trims never retime the ambience), and dim is a black scrim for legibility. The hand-picked set is GET vos.so/api/backdrops: write key from urls["1080p"], duration from the row, and frame.background from its ground so the frame before the first decoded frame is the loop's own colour. A loop's length is a fact of the asset, never a number you choose. Depth dials: frame.parallax (0 to 1) makes the media counter-pan with the zoom, and backgroundMedia.blur is design px. The card stroke is three fields: frame.border is the switch and its alpha (0 turns it off), frame.borderWidth its width in design px (absent = 1.5) and frame.borderColor any CSS colour (absent = white). The stroke grows outward from the card's edge, so it never covers the recording.

Overlays

JSON
"overlays": [{ "id": "t0", "kind": "text", "start": 1, "duration": 3,
               "text": "Ship it", "preset": "title",
               "transform": { "x": 0.5, "y": 0.82, "scale": 1, "rotation": 0 },
               "fx": { "fx": "typewriter", "unit": "char" } }]

Screen-space clips above the card, outside the zoom, output-anchored. Text presets: title (Lexend 600, 64px), caption (Lexend 400, 32px), label (JetBrains Mono, 22px), overridable with size, color, a hosted catalog family (GET vos.so/api/fonts), weight, align, letterSpacing, lineHeight, stroke, a box pill, and maxWidth as a wrap budget. transform.x/y are fractions of the frame, the zoom convention: 0.5/0.5 is center at any aspect ratio and a caption's lower third sits at y around 0.82. A layer about something on the page names it with pin ({ step | press | rect, side?, gap?, mark?, leader?, color? }) and is placed beside that thing, following it as the camera moves; transform is then the fallback for a pin that cannot resolve. The HTML layers guide has the fields. anim is how the clip enters and leaves, pure f(t), the one vocabulary every visual thing shares (the card's frame.anim, a prop's anim.idle): enter and exit take a kind (fade | rise | none; words also pop | blur | typewriter) or a step { kind, seconds, unit: block | line | word | char, direction, stagger }. The card's own frame.anim.enter also takes slide, and a footage clip's segments[].anim takes slide | fade | scale at its boundaries (see the transitions above).

Media kinds (image | video) ride the same lane: key (take-dir file or URL), width as a fraction of the frame width, radius, opacity, loop. Overlay video is clip-local in time and muted by design; soundtracks belong to audio.

Any overlay can move: motion is a list of pose keyframes over clip-local time, [{ "at": 0 }, { "at": 2, "x": 0.8, "scale": 1.4 }] slides and grows the element over its first two seconds. Values interpolate between poses with a smooth in-out ease, absent fields inherit the base transform, a hold is two identical poses, and opacity is a 0 to 1 multiplier on the clip's own alpha. In the studio, poses render as diamonds on the clip and dragging the element on the paused preview sets the pose at the playhead.

3D objects

JSON
"objects": [{ "id": "p0",
              "asset": { "kind": "primitive", "shape": "knot", "color": "#ffb03a" },
              "span": { "start": 1, "duration": 3 },
              "transform3d": { "x": 0.8, "y": 0.28, "z": 0.5,
                               "rx": 0, "ry": 40, "rz": 0, "scale": 0.18 },
              "animation": "spin" }]

World-space props between the card and the overlays. Assets are curated primitives (cube | sphere | torus | knot), a real GLB ({ "kind": "gltf", "key": "/model.glb" }, bbox-normalized so scale means the same thing), or extruded 3D text (kind: "text3d" with a catalog typeface and a fleet-audited material: standard | metal | glass | neon). x/y are frame fractions, z is world units in front of the card plane, scale is a fraction of frame height, animations (spin | float) are deterministic, and span gates visibility with soft fades. Props take the same motion pose keyframes as overlays, over the full 3D transform (x, y, z, rx, ry, rz, scale); the spin and float presets compose on top, so a prop can fly in on poses while it spins.

Cursor and export

cursor controls the drawn dot and click effects: visible, hideWhenIdle (fades out after about a second parked, always back at full opacity on a click), smoothing, size and click-effect styling. visible: false hides the dot only; the track still drives cursor-follow zoom.

export declares the intent: resolution (720p | 1080p | 2k | 4k), fps (30 or 60), format. Never set resolution above the footage; vos validate warns which preset matches the recording.

Field reference

Every field, generated from schema/doc.schema.json (the schema vos validate enforces), so this table cannot drift from the CLI. A * marks a required field.

fieldtypedescription
source*objectThe recording this doc edits. Written by record/plan — treat as read-only except videoKey/camKey/micKey when relinking assets. micKey (extension takes recorded after the audio split) is a separately-recorded mic sidecar: the voice, gained by micGain; the recording's own audio track is then system/tab audio, gained by systemGain.
segments*arrayKept SOURCE-time footage spans; the output timeline is their concatenation. Trim dead time by shrinking spans. Empty array = untrimmed. Never carry `rate` here (only lowered/derived segments do).
segments[].in*number 0..SOURCE seconds
segments[].out*numberSOURCE seconds, > in
segments[].holdnumber 0..30LEGACY spelling of a freeze at this segment's end, read into `freeze` on every path; write `freeze` instead
segments[].animobjectTHE animation vocabulary, one for every visual primitive. The card enters by tilt-in | pull-out | rise | fade | slide and leaves by recede | fade; words enter by none | fade | rise | pop | blur | typewriter and leave by none | fade | rise; an image or video clip enters and leaves by none | fade | rise; a prop idles by spin | float; a FOOTAGE clip meets the clip beside it by slide | fade | scale (a transition at the boundary: the incoming clip plays live while the outgoing, frozen on its last frame, moves away; a slide names its side). Absent = the house default; none = an explicit nothing
segments[].anim.entervariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
segments[].anim.exitvariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
segments[].anim.idle"spin" | "float" | nullprops only: what it does while it stays
segments[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
speedarraySpeed-change spans (SOURCE time, footage-anchored, non-overlapping). Absent = all 1×. The planner proposes source:'auto' spans for typing passages, scroll runs and idle gaps (ids s0…); set source:'manual' on spans you add or edit so they survive a re-plan (absent source also counts as manual).
speed[].id*string
speed[].in*number 0..
speed[].out*number
speed[].rate*number 0.1..16playback rate: 2 = twice as fast
speed[].source"auto" | "manual"'auto' = planner suggestion (typing/scroll/idle), replaced by re-plan; 'manual' (or absent) = user/agent work, always preserved
speed[].anchorobjectRe-record tie: `vos plan --reuse` re-times this span onto a NEW recording of the same script by resolving the step in the new meta.steps. Metadata only — `in`/`out` stay authoritative and lowering never reads it, so it costs nothing at render time
speed[].anchor.step*string | integerthe actions.json step: its `id` when it has one, else its record-time index (meta.steps)
speed[].anchor.at"start" | "end"which edge of the step the span's `in` is measured from (default start)
speed[].anchor.offsetnumberseconds from that edge to the span's `in` (negative = before it)
speed[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
freezearrayFreezes (SOURCE moments, footage-anchored): the footage frozen on the frame just before `at` for `seconds` of output, anywhere in the take (a pause on a state while a caption reads, a beat before a cut, the rest a poster ends on). The retime primitive beside `speed`: a speed span is a source range with a rate, a freeze a source point with a length. A freeze whose frame is trimmed away has no effect and comes back when the trim does. The still every still-taking surface takes (the thumbnail, a poster's render, the stills export's first moment) is where the LAST freeze begins. Absent = none. The output ending past the footage needs no freeze: the card holds its last frame under the clips that outlast it.
freeze[].id*stringstable identity (f0, f1, …)
freeze[].at*number 0..SOURCE seconds: the frame just before this moment freezes
freeze[].seconds*number 0.25..30OUTPUT seconds the frame holds
freeze[].anchorobjectre-record tie to a step (vos plan --reuse resolves it against the new take's steps and it wins over the step map); `at` stays the truth
freeze[].anchor.step*string | integerthe actions.json step: its `id` when it has one, else its record-time index (meta.steps)
freeze[].anchor.at"start" | "end"which edge of the step the freeze's `at` is measured from (default start)
freeze[].anchor.offsetnumberseconds from that edge to the freeze's `at` (negative = before it)
freeze[].fromstringthe template (a vos id, or 'endcard') that placed it, the clips' `from`: a re-apply replaces it, a loop drops it
freeze[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
zoom*arrayZoom regions (SOURCE time, footage-anchored, non-overlapping). The camera ramps in around `in`, holds until `out`, ramps out — or pans straight to the next span when the gap is short. Planner-made spans have source:'auto' (replaced on re-plan); set source:'manual' on any span you add or edit so it survives.
zoom[].id*stringstable identity (`z{n}` planner, `u{n}` user)
zoom[].in*number 0..SOURCE seconds
zoom[].out*numberSOURCE seconds, > in (≥ 0.3s span reads well)
zoom[].level*number 1..5zoom level; 1.4–2.8 reads well
zoom[].cx*number 0..1focus X as a NORMALIZED [0..1] fraction of the video frame (0.5 = center). NOT pixels.
zoom[].cy*number 0..1focus Y as a NORMALIZED [0..1] fraction of the video frame (0.5 = center). NOT pixels.
zoom[].easestringarrival ease (@vosjs/timeline EASINGS name). Absent = the style's default ramp.
zoom[].transition"instant" | "fast" | "smooth" | "slow"transition SPEED for this span's ramps (in, out, and a chained pan arriving here). Absent = 'smooth' (the camera style's stock motion); 'instant' = hard cut; 'fast' ≈ half, 'slow' ≈ 1.6×.
zoom[].focusMode"manual" | "auto"'auto' = camera follows the cursor through the span (needs a cursor track); absent/'manual' = fixed cx/cy
zoom[].screenobjectwhere the target lands ON SCREEN at the apex, as FRACTIONS of the frame (0.35, 0.5 = a third of the way across, centred), under the stage camera (frame.camera: 'stage'). Absent = the centre. A placed span is the author's composition, ground beside the card included, so the stage camera's cover band does not clamp it. Ignored on a span that follows the cursor (focusMode 'auto')
zoom[].screen.x*number 0..1
zoom[].screen.y*number 0..1
zoom[].source"auto" | "manual"'auto' = planner suggestion, replaced by re-plan; 'manual' = user/agent work, always preserved
zoom[].anchorobjectRe-record tie: `vos plan --reuse` re-times this span onto a NEW recording of the same script by resolving the step in the new meta.steps. Metadata only — `in`/`out` stay authoritative and lowering never reads it, so it costs nothing at render time
zoom[].anchor.step*string | integerthe actions.json step: its `id` when it has one, else its record-time index (meta.steps)
zoom[].anchor.at"start" | "end"which edge of the step the span's `in` is measured from (default start)
zoom[].anchor.offsetnumberseconds from that edge to the span's `in` (negative = before it)
zoom[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
zoomStyle"glide" | "focus" | "cinema" | "snappy" | "cut" | "none" | "keynote" | "drift"Camera style preset driving planner + camera motion: glide (default), focus, cinema, snappy, cut, none. Absent = 'glide'. 'none' disables auto-zoom planning (manual spans move like glide). Every style also sets the tempo of the tilt track (a lean lands with its zoom); the lean's intensity is tiltStyle. 'keynote' and 'drift' are RETIRED names still accepted for old documents: they read as glide + tiltStyle medium and cinema + tiltStyle subtle (docSchemaVersion 5 rewrites them on read).
zoomParamsobjectPer-doc overrides layered on the named style (the Custom seam). Any named-style pick clears them.
speedParamsobjectAuto-speed planner overrides: idleMin/idleRate/typingMin/typingRate/scrollMin/scrollRate. Absent = defaults (idle ≥5s→4×, typing ≥3s→3×, scroll ≥2.5s→2×).
tiltarrayCard tilt regions (SOURCE time, footage-anchored, non-overlapping). While a span is active the 3D card leans to its pose; between spans it returns to FLAT (there is no static card pose: a lean is a moment on the timeline). Ramps ~0.9s in / ~0.8s out; spans ≤ ~1.35s apart swing pose-to-pose. Direction: +rx brings the TOP edge toward the camera, +ry the LEFT edge — to lean toward a right-side focus use NEGATIVE ry. Wand-made spans have source:'auto' (replaced when Dynamic tilt re-runs); set source:'manual' on spans you add or edit.
tilt[].id*stringstable identity (`t-{zoomId}` wand, `u{n}` user)
tilt[].in*number 0..SOURCE seconds
tilt[].out*numberSOURCE seconds, > in (≥ 0.8s so the pose can settle)
tilt[].rx*number -45..45pose DEGREES about the horizontal axis (+ = top edge toward camera). NOT radians, NOT fractions; ±5..18° reads premium.
tilt[].ry*number -45..45pose DEGREES about the vertical axis (+ = left edge toward camera; lean toward a right-side focus = negative). NOT radians, NOT fractions.
tilt[].easestringarrival ease (@vosjs/timeline EASINGS name). Absent = the house tilt ease.
tilt[].transition"instant" | "fast" | "smooth" | "slow"transition SPEED for this span's ramps. Absent = 'smooth' (~0.9s in / ~0.8s out); 'instant' snaps the card to the pose.
tilt[].source"auto" | "manual"'auto' = Dynamic-tilt suggestion, replaced by re-plan; 'manual' = user/agent work, always preserved
tilt[].anchorobjectRe-record tie: `vos plan --reuse` re-times this span onto a NEW recording of the same script by resolving the step in the new meta.steps. Metadata only — `in`/`out` stay authoritative and lowering never reads it, so it costs nothing at render time
tilt[].anchor.step*string | integerthe actions.json step: its `id` when it has one, else its record-time index (meta.steps)
tilt[].anchor.at"start" | "end"which edge of the step the span's `in` is measured from (default start)
tilt[].anchor.offsetnumberseconds from that edge to the span's `in` (negative = before it)
tilt[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
tiltStyle"off" | "subtle" | "medium" | "strong"Dynamic-tilt wand intensity: derives source:'auto' tilt spans FROM the zoom spans, leaning toward each zoom's focus (max 5/9/14° per axis). 'off'/absent = no auto tilt; manual spans work either way.
rejectedarrayDeleted planner proposals, kept so no re-plan proposes them again ("not this one"). Each entry is the lane and the SOURCE extent of an auto span you deleted; a fresh proposal on that lane overlapping the extent by half of the shorter span is dropped by vos plan, by plan --reuse (which re-times these like manual spans) and by the studio's re-plans. Manual spans never need this: a re-plan keeps them by contract. The renderer never reads it.
rejected[].id*stringstable identity (`r{n}`)
rejected[].lane*"zoom" | "tilt" | "speed"which planner's proposals this rejects
rejected[].in*number 0..SOURCE seconds, the deleted span's start
rejected[].out*numberSOURCE seconds, > in
rejected[].anchorobjectthe deleted span's step anchor, when it had one, so a re-record re-times the rejection the same way
rejected[].anchor.step*string | integerthe actions.json step: its `id` when it has one, else its record-time index (meta.steps)
rejected[].anchor.at"start" | "end"which edge of the step the span's `in` is measured from (default start)
rejected[].anchor.offsetnumberseconds from that edge to `in` (negative = before it)
rejected[].notestringwhy, in the deleter's words
rejected[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
audio*arrayMusic/SFX clips placed on the OUTPUT timeline (they do NOT follow footage through trims; mic audio does). The CLI render mixes + muxes these (Opus for webm, AAC/Opus for mp4) on full single-flight renders; --range renders stay silent and --parallel is forced to 1 when audio is present. `key` may be a file inside the take dir (e.g. "/music.mp3").
audio[].id*string
audio[].key*stringaudio file URL (blob or asset URL)
audio[].fromstringthe template that placed this clip (see the overlay clip's from)
audio[].name*string
audio[].start*number 0..OUTPUT seconds
audio[].in*number 0..kept span start within the source file, seconds
audio[].out*numberkept span end, > in
audio[].duration*numberfull source-file length, seconds
audio[].gain*number 0..1
audio[].fadeIn*number 0..
audio[].fadeOut*number 0..
audio[].loopboolean
audio[].loopLennumber
audio[].duckbooleanduck under the mic while speech is detected
audio[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
micGainnumber 0..1master gain for the VOICE: the mic sidecar when source.micKey exists, else the recording's own audio. Absent = 1. 0 = muted.
systemGainnumber 0..1master gain for the recording's own system/tab audio track — split takes (source.micKey) only; legacy takes have one track on micGain. Absent = 1. 0 = muted.
cursor*objectCursor rendering (visible, hideWhenIdle, smoothing, size, style, click effects). See CursorStyle in @vosjs/studio-core. `style: "arrow"` (the default for new takes) draws an OS-style pointer with its tip on the recorded point; `"default"` (every earlier take) draws the white dot. `visible: false` hides the drawn pointer only — the track still drives cursor-follow zoom, and click effects keep drawing (silence those with clickFx.style: "none"). `hideWhenIdle` (default true) fades the dot out after ~1s parked and back in when it moves; scrolling is not movement, and it is always back at full opacity on a click.
cam*objectWebcam bubble style + SOURCE-time window — the bubble's REST pose. Only applies when the take has a cam track. Placement: x/y (bubble CENTER as frame fractions, the zoom cx/cy convention) when present, else the `position` corner anchor; size = diameter as a fraction of frame height. Look: shape circle|rounded; radius (rounded corner, design px, default 18); border { width design px (default 3; 0 = none), color } (default white at 0.9); shadow none|soft|strong (default soft). Animate the pose over time with camMotion spans.
camMotionarrayAnimated cam layout (SOURCE time, footage-anchored, non-overlapping): while a span is active the webcam bubble holds its pose; outside spans it rests at doc.cam; spans ≤ ~1.2s apart in output time morph pose-to-pose. Ramps ~0.65s, settled at span start (a cam move frames what follows). Absent pose fields inherit the rest pose, so a span may move without resizing. Only renders when the take has a cam track (source.camKey).
camMotion[].id*stringstable identity (`m{n}` user-created)
camMotion[].in*number 0..SOURCE seconds
camMotion[].out*numberSOURCE seconds, > in (≥ 0.5s so the move can settle)
camMotion[].xnumber 0..1bubble CENTER as a fraction of the frame width [0..1] (the zoom cx/cy convention). NOT pixels. Absent = the rest pose's x.
camMotion[].ynumber 0..1bubble CENTER as a fraction of the frame height [0..1]. Absent = the rest pose's y.
camMotion[].sizenumber 0.08..0.6bubble diameter as a fraction of the frame height. Absent = the rest pose's size.
camMotion[].easestringarrival ease (@vosjs/timeline EASINGS name, e.g. 'power2.out'). Absent = the house settle ease.
camMotion[].transition"instant" | "fast" | "smooth" | "slow"transition SPEED for this span's morphs. Absent = 'smooth' (~0.65s); 'instant' jump-cuts the bubble to its pose (the layout-cut).
camMotion[].source"auto" | "manual"reserved for a future planner; spans you add or edit are 'manual'
camMotion[].mediastringthe media this belongs to: a `media[].id`; absent = the primary, `source`
frame*objectCard framing: background (CSS), padding, radius, shadow 0..1, border 0..1, aspectRatio, browserBar.
frame.backgroundstringCSS background (gradient/color) — the underlay/fallback, always painted under backgroundMedia
frame.backgroundMediaobject | nullOptional media layer drawn OVER the CSS background and UNDER the card: a vos animation baked to a seamless loop, or a still image. Video time is OUTPUT-anchored modulo the loop (bgT = t % duration) — trims/speed never retime it. Draws outside the zoom transform. `key` may be a file inside the take dir (e.g. "/bg.webm"), a public https://assets.vos.so/... loop, or an /api/... URL. Fail-open: if it can't load, the CSS background still shows.
frame.paddingnumber
frame.radiusnumber
frame.shadownumber 0..1the card shadow's strength 0..1, shared by three layered shadows whose blur and offset grow (3/1, 14/6, 48/22 design px) while each stays at a low alpha, so the card reads as lifted rather than sitting in a dark pool
frame.shadowContactnumber 0..1contact shadow strength: a second tight layer (blur 10, offset 3 design px) over the ambient one, what makes a light card sit on a light ground. Absent = 0
frame.shadowColorstringshadow colour as #rrggbb; the strengths are its alpha. Absent = black
frame.insetobjectper-side card placement as FRACTIONS of the frame (left/right of its width, top/bottom of its height), each overriding `padding` on its side; a NEGATIVE side bleeds the card past the frame edge (headroom above, the window running off the bottom). Absent sides keep padding. Under contain the card centres inside the inset area; under cover the inset area is the card
frame.inset.topnumber -2..0.9
frame.inset.rightnumber -2..0.9
frame.inset.bottomnumber -2..0.9
frame.inset.leftnumber -2..0.9
frame.bordernumber 0..1stroke around the card: the switch AND its alpha (0 = off)
frame.borderWidthnumber 0..24border stroke width in design px (scales with the canvas, like radius), drawn OUTWARD from the card's edge so it never covers the recording. Absent = 1.5
frame.borderColorstringborder stroke colour, any CSS colour string. Absent = #ffffff; `border` is the alpha it is drawn at
frame.aspectRatiostring
frame.parallaxnumber 0..1V2: background media counter-pans subtly with the zoom camera (depth cue). 0/absent = static; 0.6 reads well
frame.fit"contain" | "cover"How footage meets an off-ratio frame. contain (default) letterboxes the card onto the background; cover makes the padded area the card and fills it with footage, cropped around `focus` — what a 440x280 store tile or a 2.5:1 marquee demands
frame.camera"card" | "stage"The zoom camera's model. card (absent; every older take) is the magnifier: the card scales about the focus, which stays where it is on screen, and the focus is clamped so the zoomed card always covers the canvas. stage is a camera: the card scales and slides so the focus lands at the frame's centre, the focus is never clamped, and past the card's edge the frame shows the ground. New takes open on stage
frame.focusobjectcover-crop anchor, normalized video-frame fractions (the zoom cx/cy convention). Absent = center; ignored under contain
frame.focus.cxnumber 0..1
frame.focus.cynumber 0..1
frame.browserBarobject
frame.browserBar.kindstring'mac' | 'windows' | 'minimal' | 'none' | 'original' (window takes with a chrome crop)
frame.animobjectTHE animation vocabulary, one for every visual primitive. The card enters by tilt-in | pull-out | rise | fade | slide and leaves by recede | fade; words enter by none | fade | rise | pop | blur | typewriter and leave by none | fade | rise; an image or video clip enters and leaves by none | fade | rise; a prop idles by spin | float; a FOOTAGE clip meets the clip beside it by slide | fade | scale (a transition at the boundary: the incoming clip plays live while the outgoing, frozen on its last frame, moves away; a slide names its side). Absent = the house default; none = an explicit nothing
frame.anim.entervariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
frame.anim.exitvariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
frame.anim.idle"spin" | "float" | nullprops only: what it does while it stays
frame.entranceobjectDEPRECATED: the card's anim.enter (migrated on read). How the card ENTERS at t = 0: tilt-in swings in from a perspective pose while the card settles up and in; pull-out opens zoomed in and pulls out; rise settles up and in flat
frame.entrance.kind*"tilt-in" | "pull-out" | "rise" | "none"
frame.entrance.secondsnumber 0.2..3
frame.focusFollow"camera"under fit: cover, the crop follows the zoom track's focus (a 9:16 cut of a 16:9 take keeps the affordance in frame). Absent = focus, or the centre
overlaysarrayScreen-space layers drawn ABOVE the card, outside the zoom transform. OUTPUT-anchored (start is final-cut seconds — trims/speed never retime a title). A layer about something on the page names it with `pin` and is placed beside it, following it as the camera moves; a layer with no referent is a caption and sits in the margin. Optional; absent = none.
overlays[].id*string
overlays[].kind*"text" | "image" | "video" | "html"
overlays[].pinobjectThe layer's REFERENT: the thing on the page it is about. Exactly one of step, press or rect names it. The lowering places the layer a gap off one side of the referent and carries it with the referent through the camera (the layer keeps its screen size), so the statement and its subject stay bound; transform.x/y stay the fallback when the pin cannot resolve (validate says why). Not the re-record `anchor` on spans.
overlays[].pin.stepstring | integera recorder step (actions.json `id`, else its index in meta.steps) whose element rect is the referent — hover, click, type and drag steps carry one; an older take resolves it from the presses inside the step's window
overlays[].pin.pressnumberSOURCE seconds: the press nearest this moment (within 0.5 s) names the referent — for a human recording, which has no steps
overlays[].pin.rectobjectthe referent itself, in FRACTIONS of the video frame [0..1] (the zoom cx/cy convention; NOT pixels)
overlays[].pin.rect.x*number
overlays[].pin.rect.y*number
overlays[].pin.rect.w*number
overlays[].pin.rect.h*number
overlays[].pin.side"auto" | "right" | "left" | "below" | "above"which side of the referent the layer sits on; auto (the default) takes the first of right, left, below, above whose box fits inside the frame through the layer's life, else the roomiest
overlays[].pin.gapnumber 0..design px between the referent and the layer's visible box (default 24)
overlays[].pin.mark"none" | "ring" | "underline"a standing mark on the referent for the layer's life: a rounded ring around it, or a bar under it (default none)
overlays[].pin.leaderbooleana hairline from the layer's near edge to the referent, with a dot at the referent
overlays[].pin.colorstringthe mark's and the leader's ink (CSS colour; default white over a dark rim, legible on any ground)
overlays[].htmlstringhtml kind: the layer's MARKUP, well-formed XML (foreignObject is XML, not HTML: <br />, &amp;, quoted attributes). A UI component authored as DOM, laid out and painted by the browser and composited over the footage sharp at any export size. Required for html. No script runs. An image it names by URL is inlined for you when the URL is on https://assets.vos.so/ (the platform's own catalogs); any other address paints blank, a file of your own on vos.so included, so inline it as a data: URI. Faces are found from the CSS
overlays[].cssstringhtml kind: the layer's CSS, scoped by the layer's own wrapper (plain selectors are fine). Its box-shadow, border and radius ARE the layer's chrome (the compositor adds none). A font-family it names is resolved against the hosted catalog (GET https://vos.so/api/fonts) at every font-weight it names. An animation or transition inside a STILL layer runs on the wall clock, never the timeline, and validate warns; on a live layer (live: true) its @keyframes are scrubbed to the clip's time
overlays[].livebooleanhtml kind: a LIVE layer, the picture a function of clip-local time. CSS @keyframes inside it are scrubbed to t (paused, delayed by -t) and {{t}} / {{data.<name>}} placeholders in the markup and CSS fill per frame, so a counter, a progress bar or a typed line come out on the timeline's clock in the preview and in every export chunk alike. One rasterize per frame while on screen; an authored animation-delay is overridden (fold it into the keyframes). Absent = a still
overlays[].dataobjecthtml kind: values for a live layer's {{data.<name>}} placeholders, escaped as text (never markup)
overlays[].boxobjectbackground pill behind the text block (text only; absent = none). Paddings/radius are EMs of the resolved font size
overlays[].box.color*stringCSS pill color
overlays[].box.opacitynumber 0..1
overlays[].box.paddingXnumber 0..4EMs, default 0.6
overlays[].box.paddingYnumber 0..4EMs, default 0.35
overlays[].box.radiusnumber 0..2EMs, default 0.25
overlays[].bleednumber 0..html kind: design px of room around the box for what paints OUTSIDE it (a shadow, a glow), or the SVG viewport clips it square. Absent = derived from the CSS's shadows; stated, it wins
overlays[].fontsarrayhtml kind: faces the CSS names that the catalog does NOT host, by URL (the page fetches and inlines them). A hosted family needs no entry: name it in the CSS
overlays[].fonts[].family*string
overlays[].fonts[].url*stringhttps URL of a woff2 face
overlays[].fonts[].weightnumber 100..900
overlays[].fonts[].style"normal" | "italic"
overlays[].start*number 0..OUTPUT seconds
overlays[].duration*number
overlays[].textstringcontent; \n breaks lines. With `emphasis` set (OPT-IN; {} is enough), *words between asterisks* are set in the emphasis weight (a two-weight caption: light words with bold ones), a word-by-word reveal keeps each word in its own weight, \* is a literal asterisk and a lone * stays as typed. Without `emphasis`, asterisks are text as typed
overlays[].emphasisobjecttext kind: turns *markers* in text ON ({} is enough) and says how marked words are set. Absent = emphasis off: asterisks are text as typed
overlays[].emphasis.weightnumber 100..900the weight of *marked* words, snapped to what the family hosts (absent = the family's bold step, 700 or the nearest)
overlays[].emphasis.colorstringthe colour of *marked* words (absent = the clip's colour)
overlays[].preset"title" | "caption" | "label"house style: title = Lexend 600 64px, caption = Lexend 400 32px, label = JetBrains Mono 22px
overlays[].sizenumber 12..200font-size override, DESIGN px (H=1080 space)
overlays[].colorstringCSS color override
overlays[].familystringfont family override — a catalog family name (GET https://vos.so/api/fonts). Unknown names fail open to the preset stack on the render fleet
overlays[].weightnumber 100..900weight override — snapped to the nearest weight the catalog hosts for the family (weights are files, not synthesis)
overlays[].italicbooleansynthesized oblique (no italic files are hosted)
overlays[].align"left" | "center" | "right"multi-line alignment within the block (default center)
overlays[].letterSpacingnumber -10..60letter spacing, design px at the resolved size (default 0)
overlays[].lineHeightnumber 0.8..3line-height multiplier (default 1.25)
overlays[].strokeobjecttext outline drawn under the fill
overlays[].stroke.color*string
overlays[].stroke.width*number 0.5..40design px at the resolved size
overlays[].transform*object
overlays[].transform.x*number -0.5..1.5anchor CENTER as a FRACTION of the frame width [0..1] (0.5 = center at ANY aspect — the zoom cx/cy convention; NOT pixels)
overlays[].transform.y*number -0.5..1.5anchor CENTER as a FRACTION of the frame height [0..1] (lower-third ≈ 0.82; NOT pixels)
overlays[].transform.scalenumberuniform multiplier on the preset size (default 1)
overlays[].transform.rotationnumberdegrees, screen-space (default 0)
overlays[].maxWidthnumber 0.1..1text kind: wrap budget as a FRACTION of the frame width (greedy word wrap at measured widths; absent = lines break only on \n; no intra-word breaks)
overlays[].shadowvariantmedia kinds: the card shadow, 'none' | 'soft' | 'strong' (absent = 'soft', the house picture-in-picture look). Text kind: the legibility shadow STRENGTH 0..1 (absent = the preset's; 0 = none, which words on a plate ground want)
overlays[].borderobjectmedia kinds: an outline stroke over the clipped media edge. Absent = none
overlays[].border.color*stringCSS color
overlays[].border.width*number 0..40design px (scales with the canvas, like radius)
overlays[].animobjectTHE animation vocabulary, one for every visual primitive. The card enters by tilt-in | pull-out | rise | fade | slide and leaves by recede | fade; words enter by none | fade | rise | pop | blur | typewriter and leave by none | fade | rise; an image or video clip enters and leaves by none | fade | rise; a prop idles by spin | float; a FOOTAGE clip meets the clip beside it by slide | fade | scale (a transition at the boundary: the incoming clip plays live while the outgoing, frozen on its last frame, moves away; a slide names its side). Absent = the house default; none = an explicit nothing
overlays[].anim.entervariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
overlays[].anim.exitvariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
overlays[].anim.idle"spin" | "float" | nullprops only: what it does while it stays
overlays[].fromstringthe template that placed this clip (its vos id): provenance the lowering never reads. A loop drops a template's clips, a re-apply replaces one template's and leaves the maker's
overlays[].enter"none" | "fade" | "rise"DEPRECATED: anim.enter (migrated on read)
overlays[].exit"none" | "fade" | "rise"DEPRECATED: anim.exit (migrated on read)
overlays[].motionarrayPose keyframes (MO): the clip's transform animated over CLIP-LOCAL time. Values interpolate across the gap between poses (~symmetric in-out ease; a hold is two identical poses); the base transform is the value before the first pose, and absent pose fields inherit it. Pure f(t) — scrub, export and chunked server renders agree.
overlays[].motion[].at*number 0..CLIP-LOCAL seconds (0 = the clip's start; poses ride along when the clip moves)
overlays[].motion[].xnumber -0.5..1.5anchor CENTER as a fraction of the frame width (the transform.x convention). NOT pixels. Absent = the base transform's x.
overlays[].motion[].ynumber -0.5..1.5anchor CENTER as a fraction of the frame height. Absent = the base transform's y.
overlays[].motion[].scalenumberuniform scale multiplier. Absent = the base transform's scale.
overlays[].motion[].rotationnumberdegrees, screen-space. Absent = the base transform's rotation.
overlays[].motion[].opacitynumber 0..1opacity MULTIPLIER on the clip's own alpha (default 1)
overlays[].motion[].easestringarrival ease (@vosjs/timeline EASINGS name). Absent = 'power2.inOut'.
overlays[].fxobjecttext kind only: entrance animation evaluated per unit — pure f(t), segmentation baked at lowering (scrub/seek/server chunks agree by construction)
overlays[].fx.fx*"fade" | "rise" | "pop" | "blur" | "typewriter"the entrance; typewriter is a step reveal by unit count
overlays[].fx.unit"block" | "line" | "word" | "char"what animates as one thing (default block — the whole text; char is grapheme-safe)
overlays[].fx.direction"forward" | "reverse" | "center"unit start order (default forward; center ripples outward)
overlays[].fx.staggernumber 0..2seconds between unit starts (defaults: typewriter 0.05, others 0.06 when unit ≠ block; clamped so the entrance fits the clip)
overlays[].fx.durationnumber 0.05..2per-unit seconds (default 0.35); typewriter ignores it
overlays[].keystringimage/video kinds: media URL or take-dir file (e.g. "/logo.png", "/clip.webm") — required for media. A media layer may show a DOCUMENT media by reference: media:<id> (media: alone is the primary), which brings the media's facts (its cursor track, its recorded page) and follows its kind.
overlays[].widthnumber ..1media and html kinds: base width as a FRACTION of the frame width (media default 0.35; html default = the design size, the box plus its bleed over 1920); height follows the picture's aspect; × transform.scale
overlays[].radiusnumber 0..media kinds: corner radius, design px (default 12)
overlays[].opacitynumber 0..1media and html kinds: opacity (default 1)
overlays[].loopbooleanvideo kind: loop while active (default: hold the last frame)
overlays[].frameobjectA media layer as a CARD, drawn by the card painter on its own plane above the primary card: a browser bar (allowed on any card, opens off; its address defaults to the media's recorded page for a media:<id> layer), a lean in degrees (the tilt convention), the layered shadow, and for a document media its cursor dot and click rings. Absent = the flat picture.
overlays[].frame.browserBar
overlays[].frame.leanobject
overlays[].frame.lean.rxnumber
overlays[].frame.lean.rynumber
overlays[].frame.shadownumber 0..1
overlays[].frame.shadowContactnumber 0..1
overlays[].frame.shadowColorstring
overlays[].frame.cursorboolean
objectsarrayCompositor v2 (V3): world-space 3D props between the card and the overlays — the drafted V4 engine spec. Curated primitives, GLB models by key, and extruded 3D text (TX7).
objects[].id*string
objects[].asset*object
objects[].asset.kind*"primitive" | "gltf" | "text3d"
objects[].asset.shape"cube" | "sphere" | "torus" | "knot"primitive kind
objects[].asset.colorstringCSS color (default #e4e4e7); text3d: the material's base/emissive color
objects[].asset.keystringgltf kind: GLB URL or take-dir file (e.g. "/model.glb") — bbox-normalized so scale means the same as primitives
objects[].asset.textstringtext3d kind: the extruded string (REQUIRED there)
objects[].asset.typefacestringtext3d kind: a 3D typeface catalog slug or family name (e.g. "bebas-neue", "Playfair Display"); unknown names fall back to the house face
objects[].asset.material"standard" | "metal" | "glass" | "neon"text3d kind: fleet-audited material preset (single-sided, no dispersion; default standard)
objects[].asset.depthnumber 0.02..1text3d kind: extrusion as a fraction of the glyph height (default 0.25)
objects[].asset.bevelbooleantext3d kind: beveled edges (default true)
objects[].spanobjectOUTPUT-time visibility with soft edge fades; absent = whole timeline
objects[].span.startnumber 0..
objects[].span.durationnumber
objects[].transform3d*object
objects[].transform3d.x*numberFRACTION of the frame width [0..1] (0.5 = center; NOT pixels)
objects[].transform3d.y*numberFRACTION of the frame height [0..1]
objects[].transform3d.znumber -2..2.5world units toward the camera from the card plane (0 = on it; 0.5 floats clearly in front)
objects[].transform3d.rxnumber
objects[].transform3d.rynumber
objects[].transform3d.rznumber
objects[].transform3d.scalenumber ..1fraction of the frame height (default 0.18)
objects[].animobjectTHE animation vocabulary, one for every visual primitive. The card enters by tilt-in | pull-out | rise | fade | slide and leaves by recede | fade; words enter by none | fade | rise | pop | blur | typewriter and leave by none | fade | rise; an image or video clip enters and leaves by none | fade | rise; a prop idles by spin | float; a FOOTAGE clip meets the clip beside it by slide | fade | scale (a transition at the boundary: the incoming clip plays live while the outgoing, frozen on its last frame, moves away; a slide names its side). Absent = the house default; none = an explicit nothing
objects[].anim.entervariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
objects[].anim.exitvariantone step of an animation: a bare kind is the house motion; an object adds its seconds and, for words, the per-unit grammar (a fade or rise that names its unit, stagger or seconds animates per unit)
objects[].anim.idle"spin" | "float" | nullprops only: what it does while it stays
objects[].fromstringthe template that placed this prop (see the overlay clip's from)
objects[].animation"spin" | "float" | nullDEPRECATED: anim.idle (migrated on read)
objects[].motionarrayPose keyframes (MO) over transform3d, CLIP-LOCAL time (seconds from span.start; 0 when span-less). Values interpolate across the gap between poses; absent pose fields inherit transform3d; spin/float presets compose additively. Pure f(t).
objects[].motion[].at*number 0..CLIP-LOCAL seconds from the clip's span start
objects[].motion[].xnumberframe fraction (the transform3d.x convention). Absent = the base.
objects[].motion[].ynumberframe fraction. Absent = the base.
objects[].motion[].znumber -2..2.5world units toward the camera. Absent = the base.
objects[].motion[].rxnumberEuler degrees
objects[].motion[].rynumberEuler degrees
objects[].motion[].rznumberEuler degrees
objects[].motion[].scalenumber ..1fraction of the frame height. Absent = the base.
objects[].motion[].easestringarrival ease (@vosjs/timeline EASINGS name). Absent = 'power2.inOut'.
export*object
export.resolution*"720p" | "1080p" | "2k" | "4k"names the SHORT edge. Presets above the footage's capture width upscale (validate warns).
export.fps*30 | 60
export.format*"mp4"doc-level default; the CLI render's --format flag (webm|mp4) overrides
endCardobjectthe end card the clip closes on: a hold on the last frame (default 2.5 s) while the card recedes and the words rise as the house title, caption and label overlays. The clip's last frame is its poster
endCard.secondsnumber 1..8
endCard.headlinestring
endCard.substring
endCard.wordmarkstring
endCard.markobjectthe brand's mark as an image (a take-dir path or a URL the render page can load) with its width-over-height, drawn above the wordmark; a wide mark (a stylised wordmark asset) stands alone. vos deliver fetches BRAND.md's logoUrl into the take's brand/ folder and sets this
endCard.mark.key*string
endCard.mark.aspect*number
mediaarrayThe take's OTHER media (concat, a layer's source): each the `source` shape (a media plus the facts only a capture has) with an `id` a segment or a source-anchored span names. The primary stays `source`. Absent = one media. Each may carry `frame`, its own card.
media[].id*stringstable identity a segment or a span names (m1, m2, …)
media[].frameobjectThis media's OWN card over the take's frame while it plays: its placement and size (`inset`), the browser bar, the corner, the shadow, the border, the cover fit and its focus. Frame-wide fields (aspectRatio, padding, background, backgroundMedia, anim) stay the take's. Absent fields fall to the take's frame; the bar names this media's recorded page, and a media with no page wears no bar unless `browserBar` says so.
media[].frame.browserBar
media[].frame.radius
media[].frame.shadow
media[].frame.shadowContact
media[].frame.shadowColor
media[].frame.inset
media[].frame.border
media[].frame.borderWidth
media[].frame.borderColor
media[].frame.fit
media[].frame.camera
media[].frame.focus
media[].frame.focusFollow
stillnumber 0..The frame that stands for this take, in OUTPUT seconds: the cover the shelf shows, the still a kit leads with (vos plan --still <t>, LAUNCH.md still:). Absent, the take's own last freeze (a poster's rest; a template's freezes never count), else the hero moment after the card and the opening clips have entered.
program*objectThe anchor.
program.configobjectThe user's VosConfigJson, as authored. Present on the wire; omitted on disk in a program directory (config.json is the config).
program.tweenEditsobjectRetimes over the config's recorded tweens, keyed by spec index: { startTime?, duration?, ease?, to?, from? }. Baked into the composed config's createTimeline; the authored config keeps its own.
program.durationnumberThe anchor's own output length in seconds, when the config's duration is a placeholder. Absent = the config's.
  • Verbs for the --set overrides that test any of these fields from the CLI.
  • The take directory for where this document lives.