Cross-language speech synthesis method and device, storage medium and electronic device
By extracting speech information from the original video and constructing a personalized timbre model, combined with a large language model for cross-language translation and duration adaptation, target language speech data is generated, and lip-syncing and subtitle rendering are performed. This solves the problem of timbre preservation and multimodal synchronization in cross-language video translation, and improves the realism and naturalness of the video.
Patent Information
- Application Number
- CN202610720905.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for cross-language video translation suffer from difficulties in preserving timbre and synchronizing multimodal expressions, resulting in translated videos that lack realism and immersion, and are also costly to produce and have long processing cycles.
By extracting audio information from the original video, a personalized timbre model is constructed. Combined with a large language model, semantic translation and duration adaptation are performed to generate target language audio data. Lip alignment and subtitle rendering are then performed to achieve synergistic consistency between audio, subtitles, and lip movements.
It enhances the realism and naturalness of cross-language video translation, reduces production costs, shortens processing cycles, and achieves synchronized collaboration between speech, subtitles, and lip movements.
Smart Images

Figure CN122416982A_ABST