Home › Guides › Motion Control

Using Kling Motion Control through one API call

Updated 2026-10-02

Motion Control is a different job from text-to-video. You supply two things: a character image and a reference video. The motion and expression from the video are transferred onto the character in the photo. Think of it as puppeteering a still. It is available as Kling v3.0 Motion Control through POST /videos, with no separate endpoint. This guide explains the two ways to send the reference video, how long the output will be, and what to watch for.

What it is, and what it is not

VideoRouter's docs distinguish it from neighbouring features. Lip sync changes mouth movement from audio. Face swap replaces a face. Reference-to-video uses reference files as a style guide with no promise of preserving motion. Motion Control is the one where a driving video's actual movement is carried onto the character in your photo. If your goal is "make this person in this photo do what the person in that clip does", this is the model family to look at.

Model ids

Motion Control ships as two tiers. VideoRouter's documentation shows them as kling/motion-control-std/novita and kling/motion-control-pro/novita, with Novita named as the host in the docs. The catalog also lists a kling-motion-control model page; the live price table shows which hosts currently serve it, and hosts can change, so treat the docs' Novita pinning as the known-good form and check the model page for others. The Std and Pro split follows the same pattern as Kling v3.0 (see Std vs Pro vs 4K), and the two tiers are priced differently.

Orientation 1: reference in input_video_references

Here the reference motion video goes into input_video_references with exactly one entry, alongside start_image_url. This is the one documented case where a reference array and a start image are allowed together. The output is a fixed length of about four seconds and follows the photo's own composition rather than the reference video's length.

import requests

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
    json={
        "model": "kling/motion-control-std/novita",
        "start_image_url": "https://example.com/character.jpg",
        "input_video_references": [
            {"type": "video_url", "video_url": {"url": "https://example.com/reference-motion.mp4"} }
        ],
    },
).json()
# {"id": "video_...", "status": "queued", ...}

Orientation 2: reference in input_video_url

The second form puts the reference video in input_video_url, the same field used for video editing, still paired with start_image_url. Here the output duration matches the reference video's own probed length, up to 30 seconds.

job = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_..."},
    json={
        "model": "kling/motion-control-pro/novita",
        "start_image_url": "https://example.com/character.jpg",
        "input_video_url": "https://example.com/reference-motion.mp4",
    },
).json()

Unlike every other editing model on the platform, neither form requires a prompt. The reference video is the instruction.

Choosing between the two

input_video_referencesinput_video_url
Output lengthFixed, about 4 secondsMatches the reference video, up to 30 seconds
CompositionFollows the photoDocumented as matching the reference's length; check output framing on your own inputs
Good forShort loops, avatar reactions, quick testsFull performances, longer gestures, dance or speech clips

Because billing is by duration, the second form's cost scales with the length of the reference clip. The docs state both orientations bill the same per-second rate regardless of which field carries the reference, and that the rate already includes the 2% platform fee. Trim your reference clip to the part you need before uploading, rather than paying for seconds you will cut. For real numbers see the live table; this page does not quote rates.

Preparing good inputs

Limits and gotchas

A few practical notes before the list. Both orientations are ordinary jobs under POST /videos, so everything else you know about the endpoint applies: bearer-key auth, OpenAI-style error envelopes, and a failed status when the upstream cannot complete the job. Input validation errors, such as a missing start image or a reference URL the host cannot fetch, are most likely to show up at creation or very early in the job, which is another reason to test with a short clip first.

A reasonable rollout looks like this: run three or four character images against one trusted reference clip on Std, look for failure patterns (cropped limbs, face drift, hands), fix the inputs, and only then spend on Pro. The reference clip usually matters more than the model tier for the final look, because it defines the motion the model has to reproduce.

For authentication and job basics see the quickstart. When you are ready to run your own character image, sign up for a key. Full parameter details are in the Motion Control docs.

Frequently asked questions

What does Kling Motion Control do?

It transfers the motion and expression from a reference video onto a character in a still image, so the photo performs what the video shows.

How long is the output?

With the reference in input_video_references the output is a fixed length of about four seconds. With it in input_video_url the output matches the reference video's length, up to 30 seconds.

Do I need a text prompt?

No. Per the docs neither orientation requires a prompt; the reference video acts as the instruction.

Which model ids do I use?

The docs use kling/motion-control-std/novita and kling/motion-control-pro/novita. Check the model page for the current host list.

Keep reading

Using Kling is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →