Cross-language speech synthesis method and device, storage medium and electronic device

By extracting speech information from the original video and constructing a personalized timbre model, combined with a large language model for cross-language translation and duration adaptation, target language speech data is generated, and lip-syncing and subtitle rendering are performed. This solves the problem of timbre preservation and multimodal synchronization in cross-language video translation, and improves the realism and naturalness of the video.

CN122416982APending Publication Date: 2026-07-17SHENZHEN SKYWORTH DISPLAY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610720905.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for cross-language video translation suffer from difficulties in preserving timbre and synchronizing multimodal expressions, resulting in translated videos that lack realism and immersion, and are also costly to produce and have long processing cycles.

Method used

By extracting audio information from the original video, a personalized timbre model is constructed. Combined with a large language model, semantic translation and duration adaptation are performed to generate target language audio data. Lip alignment and subtitle rendering are then performed to achieve synergistic consistency between audio, subtitles, and lip movements.

Benefits of technology

It enhances the realism and naturalness of cross-language video translation, reduces production costs, shortens processing cycles, and achieves synchronized collaboration between speech, subtitles, and lip movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416982A_ABST
    Figure CN122416982A_ABST
Patent Text Reader

Abstract

本申请涉及一种跨语言的语音合成方法、装置、存储介质以及电子设备。该方法包括:对原始视频进行语音信息提取处理,得到目标人物的原始语音数据及其源语言文本、韵律特征以及音色特征,并构建与目标人物对应的个性化音色模型;对源语言文本进行目标语言的语义翻译处理,并进行时长适配处理,得到目标语言文本;将韵律特征、个性化音色模型以及目标语言文本输入至语音合成模型中,以生成目标语言语音数据,并对原始视频中的目标人物进行口型对齐及语音数据替换处理,得到第一视频;对第一视频的原始字幕区域进行擦除处理,并将目标语言文本渲染至对应区域,得到目标语言视频。本申请解决了跨语言视频翻译中音色保留与多模态同步难的技术问题。
Need to check novelty before this filing date? Find Prior Art