Kling: Kling Lip Sync API
Kling lip sync: puts an audio track on one person in an existing video and animates the mouth, cheeks and jaw to match. Videos of 2–60 s, one person per task; pick the speaker in multi-person shots with face_id. Through the unified API you only send the video and the audio, and the gateway detects faces and aligns the timing.
- Input: audio, video
- Output: video
- Released: 2025-12-04
Frequently Asked Questions
How do I choose who speaks in a multi-person video?
Without `face_id` the person on screen the longest is used. To pick someone else, call `POST /kling/v1/videos/identify-face` to see each person's `face_id` and thumbnail, then pass `face_id` in input. One person per task.
What if the audio is longer than the person is on screen?
Without timing parameters the audio starts when the person appears and is cut where the person leaves. To use part of the audio, crop it with `sound_start_time` / `sound_end_time` in milliseconds.