Using Kling Motion Control through one API call
Updated 2026-10-02
Motion Control is a different job from text-to-video. You supply two things: a character image and a reference video. The motion and expression from the video are transferred onto the character in the photo. Think of it as puppeteering a still. It is available as Kling v3.0 Motion Control through POST /videos, with no separate endpoint. This guide explains the two ways to send the reference video, how long the output will be, and what to watch for.
What it is, and what it is not
VideoRouter's docs distinguish it from neighbouring features. Lip sync changes mouth movement from audio. Face swap replaces a face. Reference-to-video uses reference files as a style guide with no promise of preserving motion. Motion Control is the one where a driving video's actual movement is carried onto the character in your photo. If your goal is "make this person in this photo do what the person in that clip does", this is the model family to look at.
Model ids
Motion Control ships as two tiers. VideoRouter's documentation shows them as kling/motion-control-std/novita and kling/motion-control-pro/novita, with Novita named as the host in the docs. The catalog also lists a kling-motion-control model page; the live price table shows which hosts currently serve it, and hosts can change, so treat the docs' Novita pinning as the known-good form and check the model page for others. The Std and Pro split follows the same pattern as Kling v3.0 (see Std vs Pro vs 4K), and the two tiers are priced differently.
Orientation 1: reference in input_video_references
Here the reference motion video goes into input_video_references with exactly one entry, alongside start_image_url. This is the one documented case where a reference array and a start image are allowed together. The output is a fixed length of about four seconds and follows the photo's own composition rather than the reference video's length.
import requests
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "kling/motion-control-std/novita",
"start_image_url": "https://example.com/character.jpg",
"input_video_references": [
{"type": "video_url", "video_url": {"url": "https://example.com/reference-motion.mp4"} }
],
},
).json()
# {"id": "video_...", "status": "queued", ...}
Orientation 2: reference in input_video_url
The second form puts the reference video in input_video_url, the same field used for video editing, still paired with start_image_url. Here the output duration matches the reference video's own probed length, up to 30 seconds.
job = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_..."},
json={
"model": "kling/motion-control-pro/novita",
"start_image_url": "https://example.com/character.jpg",
"input_video_url": "https://example.com/reference-motion.mp4",
},
).json()
Unlike every other editing model on the platform, neither form requires a prompt. The reference video is the instruction.
Choosing between the two
input_video_references | input_video_url | |
|---|---|---|
| Output length | Fixed, about 4 seconds | Matches the reference video, up to 30 seconds |
| Composition | Follows the photo | Documented as matching the reference's length; check output framing on your own inputs |
| Good for | Short loops, avatar reactions, quick tests | Full performances, longer gestures, dance or speech clips |
Because billing is by duration, the second form's cost scales with the length of the reference clip. The docs state both orientations bill the same per-second rate regardless of which field carries the reference, and that the rate already includes the 2% platform fee. Trim your reference clip to the part you need before uploading, rather than paying for seconds you will cut. For real numbers see the live table; this page does not quote rates.
Preparing good inputs
- Character image. A clear, front-facing or three-quarter shot with the full body or torso you want animated visible. If the reference video shows a full-body dance and your photo is a headshot, there is nothing to carry the body motion onto. Match framing between photo and reference as closely as you can.
- Reference video. One subject, steady camera, unobstructed body. Busy multi-person footage gives the model nothing unambiguous to follow.
- Rights. You are responsible for having the right to use both the photo and the performance in the reference. The platform ships no gallery of ready-made reference clips; you upload your own.
Limits and gotchas
A few practical notes before the list. Both orientations are ordinary jobs under POST /videos, so everything else you know about the endpoint applies: bearer-key auth, OpenAI-style error envelopes, and a failed status when the upstream cannot complete the job. Input validation errors, such as a missing start image or a reference URL the host cannot fetch, are most likely to show up at creation or very early in the job, which is another reason to test with a short clip first.
A reasonable rollout looks like this: run three or four character images against one trusted reference clip on Std, look for failure patterns (cropped limbs, face drift, hands), fix the inputs, and only then spend on Pro. The reference clip usually matters more than the model tier for the final look, because it defines the motion the model has to reproduce.
- Single host today. The docs describe one provider for these models with no cross-provider failover pool, so an unpinned request has no second host to fall back to if that one fails. Failed upstream jobs are not billed, so retry rather than resubmitting in parallel.
- Mutual exclusivity elsewhere. The start-image-plus-reference combination is a carve-out. Do not assume other models accept it.
- Poll normally. It is a standard async job:
GET /videos/{id}untilcompletedorfailed, and polling is free. - Test short first. Run the 4-second orientation on Std with a trimmed reference before committing to a 30-second Pro render.
For authentication and job basics see the quickstart. When you are ready to run your own character image, sign up for a key. Full parameter details are in the Motion Control docs.
Frequently asked questions
What does Kling Motion Control do?
It transfers the motion and expression from a reference video onto a character in a still image, so the photo performs what the video shows.
How long is the output?
With the reference in input_video_references the output is a fixed length of about four seconds. With it in input_video_url the output matches the reference video's length, up to 30 seconds.
Do I need a text prompt?
No. Per the docs neither orientation requires a prompt; the reference video acts as the instruction.
Which model ids do I use?
The docs use kling/motion-control-std/novita and kling/motion-control-pro/novita. Check the model page for the current host list.
Keep reading
- Kling v3.0 Std vs Pro vs 4K — Which Tier Should You Use?
- Kling Image-to-Video API: Animate a Start Image (Code)
- Kling O3 (Omni) vs Kling v3: Which Model Should You Call?
- How to Call the Kling API With curl: Step by Step
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →