Real-time video translation and audio and picture synchronization method and system based on multi-modal large model

Through multimodal large model and cross-modal attention mechanism, video translation is combined with multimodal features, and the problem of visual context information being ignored and audio-visual synchronization in traditional technology is solved, and video translation with high semantic accuracy is achieved.

CN120218091APending Publication Date: 2025-06-27SHANGHAI YINGZHUO INFORMATION TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510380695.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional video translation methods ignore the visual context information in the video, resulting in the disconnection of the translation results from the picture situation and fail to effectively solve the problem of timbre consistency and lip sync.

Method used

Real-time video translation and audio-visual synchronization method based on multimodal large model is adopted, and context information is dynamically aligned through multimodal features and cross-modal attention mechanisms, context semantic vectors are generated, and timbre and lip features are synchronized.

Benefits of technology

It significantly improves the semantic accuracy of the translation and realizes the four-in-one video translation of "semantic-tooth-lip-scene", solving the problems of audio and video fragmentation, high delay, and scene disconnection in traditional technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218091A_ABST
    Figure CN120218091A_ABST
Patent Text Reader

Abstract

The invention provides a real-time video translation and audio and picture synchronization method and system based on a multi-modal large model, and relates to the technical field of video translations, and the method comprises the steps: obtaining a source video; extracting the source video based on the multi-modal large model to obtain a multi-modal feature; fusing the multi-modal features through a cross-modal attention mechanism to generate a context semantic vector; translating into a target language text in real time based on the context semantic vector, and processing the translated language text based on the multi-modal features to obtain a translated language sound source; and performing mouth shape adjustment on the source video based on the translation language sound source, and merging the translation language sound source and the mouth shape animation video to obtain a real-time translation video with synchronous sound and picture. According to the method, the limitation of traditional single-modal translation is broken through, and the semantic accuracy of translation is remarkably improved by dynamically aligning the context information through the multi-modal features in combination with a cross-modal attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video translation, and particularly to a real-time video translation and audio-visual synchronization method and system based on a multimodal large model. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, video translation technology has been widely applied in fields such as film and television localization, cross-border live streaming, and online education.

[0003] Traditional video translation methods mainly rely on speech recognition (ASR) or subtitle text for unimodal translation, such as speech translation systems based on recurrent neural networks (RNNs) (e.g., Google Translate) or pure text translation engines (e.g., Transformer). Such methods completely ignore the visual context information in the video (such as the speaker's expressions, gestures, and scene objects), resulting in a disconnection between the translation result and the picture context. When dealing with translated speech, the prior art only simply replaces the original audio or overlays subtitles, without solving the problems of timbre consistency and lip synchronization.

[0004] Therefore, a real-time video translation and audio-visual synchronization method and system based on a multimodal large model are proposed. Summary of the Invention

[0005] This specification provides a real-time video translation and audio-visual synchronization method and system based on a multimodal large model, which breaks through the limitations of traditional single-modal translation. Through multimodal features and combined with a cross-modal attention mechanism to dynamically align context information, the semantic accuracy of the translation is significantly improved.

[0006] This specification provides a real-time video translation and audio-visual synchronization method based on a multimodal large model, including:

[0007] Obtain a source video;

[0008] Extract the multimodal features from the source video based on the multimodal large model;

[0009] Fuse the multimodal features through a cross-modal attention mechanism to generate a context semantic vector;

[0010] Translate the context semantic vector into a target language text in real time, and process the translated language text based on the multimodal features to obtain a translated language sound source;

[0011] Adjust the lip movements of the source video based on the translated language sound source, and merge the translated language sound source and the lip animation video to obtain a real-time translated video with audio-visual synchronization.

[0012] Optionally, the multi-modal features are fused through a cross-modal attention mechanism to generate a context semantic vector, including:

[0013] The multi-modal features include speech duration features and lip key frame features;

[0014] Based on the speech duration features and the lip key frame features, a cross-modal time similarity matrix is determined;

[0015] The cross-modal time similarity matrix is superimposed with a mask matrix, and attention weights are generated through Softmax;

[0016] The speech features are weighted and fused to generate a context semantic vector.

[0017] Optionally, the real-time translation of the context semantic vector into the target language text includes:

[0018] Based on the context semantic vector, a domain knowledge base is dynamically loaded through a sliding window mechanism to generate a translation text in the target language.

[0019] Optionally, the processing of the translated language text based on the multi-modal features to obtain a translated language sound source includes:

[0020] The multi-modal features include the source speaker's voiceprint features;

[0021] Based on the source speaker's voiceprint features and the translated language text, a text-to-speech model is driven to clone the timbre of the source speaker to obtain a translated language sound source.

[0022] Optionally, the processing of the translated language text based on the multi-modal features to obtain a translated language sound source further includes:

[0023] The multi-modal features include speech duration features;

[0024] Based on the speech duration features and the translated language text, a text-to-speech model is driven to adjust the speaking speed of the translated speaker to obtain a translated language sound source.

[0025] Optionally, the lip synchronization of the source video is adjusted based on the translated language sound source, and the translated language sound source and the lip animation video are merged to obtain a real-time translated video with lip-sync sound, including:

[0026] The translated language sound source is input into a diffusion model, and the per-frame lip key point offset is output;

[0027] The per-frame lip key point offset is superimposed on the source video through bilinear interpolation superposition values.

[0028] Optionally, superimposing the per-frame lip key-point offset on the source video through the bilinear interpolation superposition value includes:

[0029] Align the time stamps of the translated language sound source, and according to the frame-level alignment relationship output by the cross-modal attention mechanism, compensate for the timing deviation between the translated language sound source and the lip animation video through the dynamic time warping algorithm;

[0030] Perform motion compensation on the non-lip region of the source video through optical flow estimation;

[0031] Map the lip animation to the source video through bilinear interpolation to generate a lip mask;

[0032] Synthesize a real-time translated video with synchronized sound and picture through a residual fusion model.

[0033] This specification provides a video translation and sound-picture synchronization system based on a multi-modal large model, including:

[0034] An acquisition module for acquiring a source video;

[0035] An extraction module for extracting multi-modal features from the source video based on the multi-modal large model;

[0036] A fusion module for fusing the multi-modal features through a cross-modal attention mechanism to generate a context semantic vector;

[0037] A translation module for real-time translating the context semantic vector into a target language text and processing the translated language text based on the multi-modal features to obtain a translated language sound source;

[0038] An adaptation module for adjusting the lip shape of the source video based on the translated language sound source and merging the translated language sound source and the lip animation video to obtain a real-time translated video with synchronized sound and picture.

[0039] Optionally, the fusion module includes:

[0040] The multi-modal features include speech duration features and lip key-frame features;

[0041] Based on the speech duration features and the lip key-frame features, determine a cross-modal time similarity matrix;

[0042] Superimpose the cross-modal time similarity matrix with a mask matrix and generate attention weights through Softmax;

[0043] Perform weighted fusion on the speech features to generate a context semantic vector.

[0044] Optionally, the translation module includes:

[0045] Based on the context semantic vector, the domain knowledge base is dynamically loaded through a sliding window mechanism to generate the translated text in the target language.

[0046] Optionally, the translation module includes:

[0047] The multi-modal features include the voiceprint feature of the source speaker;

[0048] Based on the voiceprint feature of the source speaker and the translated language text, drive the text-to-speech model to clone the timbre of the source speaker to obtain the translated language sound source.

[0049] Optionally, the translation module further includes:

[0050] The multi-modal features include the speech duration feature;

[0051] Based on the speech duration feature and the translated language text, drive the text-to-speech model to adjust the speaking speed of the translated speaker to obtain the translated language sound source.

[0052] Optionally, the adaptation module includes:

[0053] Input the translated language sound source into the diffusion model to output the offset of lip key points frame by frame;

[0054] Superimpose the offset of lip key points frame by frame on the source video through bilinear interpolation superposition values.

[0055] Optionally, the step of superimposing the offset of lip key points frame by frame on the source video includes:

[0056] Align the time stamps of the translated language sound source, and according to the frame-level alignment relationship output by the cross-modal attention mechanism, compensate for the timing deviation between the translated language sound source and the mouth shape animation video through the dynamic time warping algorithm;

[0057] Perform motion compensation on the non-mouth shape area of the source video through optical flow estimation;

[0058] Map the mouth shape animation to the source video through bilinear interpolation to generate a mouth shape mask;

[0059] Synthesize a real-time translated video with lip-sync through a residual fusion model.

[0060] In the present invention, the limitation of traditional single-modal translation is broken through. Through multi-modal features and the cross-modal attention mechanism to dynamically align context information, the semantic accuracy of translation is significantly improved. Through the deep integration of multi-modal large models and cross-modal temporal alignment technology, "semantics-timbre-lip shape-scene" four-in-one video translation is achieved, overcoming the core pain points such as audio-visual disconnection, high latency, and scene disconnection in traditional technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0062] Figure 1 Schematic diagram of the principle of a real-time video translation and audio-visual synchronization method based on a multi-modal large model provided by an embodiment of this specification;

[0063] Figure 2 Schematic diagram of the principle of a video translation and audio-visual synchronization system based on a multi-modal large model provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations. The basic principles defined in the following description can be applied to other implementation schemes, variation schemes, improvement schemes, equivalent schemes, and other technical schemes that do not deviate from the spirit and scope of the present invention.

[0065] The following combines the attached Figure 1-2 Describe the exemplary embodiments of the present invention more comprehensively. However, the exemplary embodiments can be implemented in various forms and should not be understood that the present invention is limited to the embodiments described herein. On the contrary, providing these exemplary embodiments can make the present invention more comprehensive and complete, and more convenient to convey the inventive concept to those skilled in the art. The same reference numerals in the figures represent the same or similar elements, components, or parts, and thus the repeated description of them will be omitted.

[0066] On the premise of conforming to the technical concept of the present invention, the features, structures, characteristics, or other details described in a specific embodiment do not exclude being combined in a suitable manner in one or more other embodiments.

[0067] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are intended to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.

[0068] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0069] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0070] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.

[0071] Figure 1 A schematic diagram of the principle of a method for real-time video translation and audio-visual synchronization based on a multimodal large model provided in an embodiment of this specification, the method may include:

[0072] S110: Acquire source video;

[0073] In a specific implementation of the present specification, a source video is obtained from a local file, a real-time live stream, and a cloud storage, and the source video includes an audio track, a video track, and an optional subtitle track.

[0074] The input video is formatted in a unified way, including:

[0075] If the video is in compressed format, a hardware decoder is used for real-time decoding; if the video contains multiple audio tracks, the audio of the main speaker is extracted through sound source separation (such as the Spleeter model).

[0076] S120: extracting the source video based on the multimodal large model to obtain multimodal features;

[0077] S130: Fusing the multimodal features through a cross-modal attention mechanism to generate a contextual semantic vector;

[0078] In the specific embodiments of this specification, a pre-trained HRNet model is used to extract the key frame features of lip movements in video frames, and a 3D-CNN model is used to capture the micro-expression features of consecutive frames (such as eyebrow movements and blink frequencies); a CLIP-ViT model is used to extract the scene context features to identify the objects, backgrounds, and limb movement semantics in the video; a Wav2Vec 2.0 model is used to extract the acoustic features of speech, and a VQ-VAE encoder is used to compress and generate duration encodings; if there are embedded subtitles, a BERT model is directly used to generate text semantic vectors; if there are no subtitles, an OCR-TransNet model is used to scan the embedded text in the video (such as signs and bullet screens in the scene) to generate visual text features.

[0079] Optionally, S130 includes:

[0080] The multi-modal features include speech duration features and key frame features of lip movements;

[0081] Based on the speech duration features and the key frame features of lip movements, a cross-modal temporal similarity matrix is determined;

[0082] The cross-modal temporal similarity matrix is superimposed with a mask matrix, and attention weights are generated through Softmax;

[0083] The speech features are weighted and fused to generate a context semantic vector.

[0084] In the specific embodiments of this specification, the speech duration features of the language signal in the source video stream are extracted (T is the time step, d a is the feature dimension); the key frame features of the visual lip movements in the source video stream are extracted where the key frames are generated through lip key point detection (20-point model) and optical flow tracking; calculate the language feature A t and the lip movement feature V k of the cross-modal temporal similarity matrix S ∈ R T×T , specifically:

[0085]

[0086] Introduce a temporal mask matrix M ∈ R T×T , and the maximum temporal offset between language and lip movements is constrained to ±τ frames (τ = 10), and the masking rule is:

[0087]

[0088] The similarity matrix S is superimposed with the mask matrix M, and attention weights W ∈ R T×T are generated through Softmax:

[0089]

[0090] Perform weighted fusion on the speech feature A t to generate a context semantic vector

[0091] C = W·V k

[0092] S140: Real-time translate the context semantic vector into the target language text, and process the translated language text based on the multimodal features to obtain the translated language sound source;

[0093] Optionally, the S140 includes:

[0094] Based on the context semantic vector, dynamically load the domain knowledge base through the sliding window mechanism to generate the translation text in the target language.

[0095] In the specific implementation of this specification, the context semantic vector is input into the pre-trained domain classification model (BERT-Based), and the scene category probabilities of the current video segment are respectively P ∈ R K (K is the number of domains, such as medical, legal, film and television); select the domain with a probability exceeding the preset value to trigger the loading of the corresponding domain knowledge base, and the knowledge base includes a glossary, a translation style module and a domain entity library. Set the sliding window size to W = 10 seconds and the window step size to S = 5 seconds, and cache the historical semantic vectors {C t-W ,..., C t} within the window; calculate the association weights between the current semantic C t and the historical context through the self-attention mechanism to generate the enhanced semantic vector

[0096]

[0097] According to the loaded domain knowledge base, construct a dynamic prompt template, for example: [Domain pattern] Translate the following {source language} text into {target language}, using {glossary}, style requirements: {style description}. Text:

[0098] {Current window text}

[0099] Input the enhanced semantic vector and the prompt template into the multimodal large model (such as GPT-4) to generate the target language text T out , and at the same time inject the proper nouns in the domain entity library (such as "CT" - "Computed Tomography"); compare the translation text T outWith visual context (such as objects in a scene, human actions), if a conflict is detected (such as the video shows "shaking hands" but is translated as "refusing"), trigger confidence-weighted correction:

[0100] T final = α·T out +(1 - α)·T backup (α = visual matching degree)

[0101] Optionally, the S140 includes:

[0102] The multi-modal features include the voiceprint features of the source speaker;

[0103] Based on the voiceprint features of the source speaker and the translated language text, drive the text-to-speech model to clone the timbre of the source speaker to obtain the translated language sound source.

[0104] In the specific implementation of this specification, a multi-task learning voiceprint encoder (such as an improved d-vector model) is used to extract the voiceprint feature vector of the speaker from the original speech. This model enhances the robustness to complex environments (such as live background sounds, multi-person conversations) through joint training of voiceprint recognition and background noise classification tasks, ensuring that the voiceprint features are not disturbed by the environment.

[0105] Convert the translated text into a phoneme sequence and extract semantic content features; inject the voiceprint features of the source speaker to control the timbre, intonation, and rhythm of the synthesized speech; the generator (such as an improved Tacotron 2) generates a Mel spectrogram based on the content and voiceprint features, and the discriminator forces the synthesized speech to be highly consistent with the source speaker in timbre by comparing the spectral details of the real and synthesized speech (such as fundamental frequency, formants).

[0106] For the phoneme difference problem in cross-language timbre cloning (such as the difference in pronunciation mouth shapes between Chinese and English), construct a phoneme-phoneme mapping table to dynamically adjust the phoneme duration and intensity of the target speech. For example, map the fricative sound "th" in English to the "s" sound in Chinese, and predict the natural rhythm transition between different languages through an LSTM network.

[0107] Generate a coarse-grained Mel spectrogram to ensure overall timbre consistency; then optimize the details (such as plosives, liaisons) through a local attention mechanism to improve the naturalness of the speech. Use low-latency models such as Parallel WaveGAN to convert the Mel spectrogram into a waveform, supporting streaming output in real-time scenarios (delay ≤ 50ms). Automatically adjust the frequency response curve of the synthesized speech according to the speech characteristics of the target language (such as more high frequencies in Japanese) to avoid timbre distortion. Perform secondary alignment of the duration of the synthesized speech and the mouth shape animation sequence, and ensure that the audio-visual synchronization accuracy error ≤ 3 frames by fine-tuning the length of the speech mute segment.

[0108] Optionally, the S140 further includes:

[0109] The multimodal features include speech duration features;

[0110] The speech speed of the translator is adjusted based on the speech duration feature and the translated language text to drive the text-to-speech model to obtain a translated language sound source.

[0111] In a specific embodiment of the present specification, the duration features at the phoneme level are extracted from the source speech, including the pronunciation duration of each phoneme, the pause interval between syllables, and the overall rhythm pattern of the sentence, and are encoded into a structured vector through a bidirectional LSTM network to capture the speaker's personalized speaking habits (such as fast-speech short sentences, emphatic extended sounds). The translated target text is predicted for phoneme boundaries, and the phoneme segmentation markers are generated in combination with the pronunciation rules of the target language (such as English linking, Japanese glottalization) to provide an alignment benchmark for cross-language speech speed adaptation. The duration features of the source speech are aligned with the phoneme segmentation of the target text in a cross-language time sequence, and the duration ratio relationship between the two at the phoneme, word, and sentence levels is calculated, and the speed scaling factor of the target speech is dynamically generated. For example, if the phoneme at the end of the interrogative sentence in the source speech is extended by 20%, the target speech is synchronously extended to the corresponding position phoneme. In view of cross-language pronunciation differences (such as Chinese monosyllables vs. English polysyllabic words), the phoneme merging and splitting strategies are used to adjust the local speech speed to ensure that the target speech remains natural and fluent when the number of syllables changes. The target speech baseline rate is set according to the average speech rate of the source speaker (such as 4.5 syllables / second) to prevent the overall speech from being too fast or too slow. At emotional expression nodes such as interrogative sentences and exclamatory sentences, the phoneme duration is dynamically adjusted according to the emphasis mode of the source speech to retain the emotional expression habits of the original speaker. The adversarial training mechanism is introduced, and the discriminator distinguishes the rhythm naturalness of real and synthetic speech, forcing the generator to learn the speech rate change curve that conforms to the target language habits. The sliding window incremental processing is adopted to update the speech rate parameters every 500ms, and the output rate is dynamically fine-tuned in combination with the context semantics (such as the need to speed up emergency broadcasts). A smooth transition is achieved through a lightweight speech buffer pool (capacity = 2 seconds), eliminating the mechanical feeling caused by sudden changes in speech rate and ensuring the coherence of synthesized speech in real-time scenarios. The synthesized speech and the lip animation sequence are aligned and checked at the millisecond level. If the deviation between the speech duration and the lip movement exceeds the threshold (such as ±50ms), the speech rate fine-tuning is automatically triggered and the speech segment is regenerated. Based on user feedback data (such as real-time ratings in live broadcast scenarios), a speech rate adaptive learning mechanism is established to continuously optimize the long-tail performance of the speech rate control model.

[0112] S150: adjusting the lip movements of the source video based on the translated language sound source, and merging the translated language sound source and the lip movement animation video to obtain a real-time translated video with synchronized audio and video.

[0113] Optionally, the S150 includes:

[0114] Input the source audio of the translation language into the diffusion model to output the offset of lip key points frame by frame;

[0115] Superimpose the offset of the lip key points frame by frame on the source video through the bilinear interpolation superposition value.

[0116] Optionally, superimposing the offset of the lip key points frame by frame on the source video through the bilinear interpolation superposition value includes:

[0117] Perform timestamp alignment on the source audio of the translation language, and compensate for the timing deviation between the source audio of the translation language and the mouth animation video according to the frame-level alignment relationship output by the cross-modal attention mechanism through the dynamic time warping algorithm;

[0118] Perform motion compensation on the non-mouth regions of the source video through optical flow estimation;

[0119] Map the mouth animation to the source video through bilinear interpolation to generate a mouth mask;

[0120] Synthesize a real-time translation video with synchronized sound and picture through a residual fusion model.

[0121] In the specific embodiments of this specification, based on the frame-level alignment relationship output by the cross-modal attention mechanism (such as the i-th frame of speech corresponding to the j-th frame of lip movement), the dynamic time warping algorithm (DTW) is adopted, with a precision unit of 10 ms, to perform microsecond-level alignment compensation on the audio stream of the translated speech and the lip animation sequence. For local deviations (such as the speech tail sound being prolonged but the lip closing in advance), a silent segment is inserted or a smooth transition frame is used to ensure that the audio-visual deviation is small. The correlation between the peak value of the speech energy and the lip opening degree is synchronously detected. If a continuous deviation is detected (such as the speech has finished playing but the lips are still moving), a real-time regeneration process is triggered to locally replace the abnormal segment. A lightweight optical flow model (such as PWC-Net) is used to predict the pixel motion vectors of non-lip regions (background, limb movements) in the original video, and motion compensation warping is performed on adjacent frames to eliminate the background jitter or tearing caused by the superposition of lip animations. For occluded regions (such as a hand passing over the lips), motion inference filling is performed based on the content of the previous and subsequent frames. The optical flow calculation area is dynamically cropped to only process the 200 px range around the lip animation, reducing the computational load while retaining the background integrity. The generated lip animation key points (20-point lip model) are mapped from the low-resolution feature space (such as 256×256) to the resolution of the original video (such as 1080p), and a smooth lip region mask is generated through bilinear interpolation to accurately cover the lip movement range. Gaussian blur and Alpha channel gradient are applied to the mask edge to eliminate the jagged effect caused by resolution scaling, enabling the synthesized lips to naturally transition with the skin color of the original video. The lip mask is decomposed into a lip retention area and a background retention area, and a residual fusion formula is adopted; the background retention area directly reuses the pixels of the original video; the lip retention area overlays the generated lip animation frames, and the brightness and hue of the original video are matched through a color correction module (such as cold tone / warm tone adaptation). A GPU-accelerated circular buffer (capacity = 500 ms) is used to temporarily store the synthesized frames, and the output resolution and bit rate are dynamically adjusted in combination with adaptive bitrate control (ABR) to ensure real-time rendering of the 4K video stream (delay ≤ 200 ms).

[0122] In the present invention, the limitations of traditional single-modal (speech / text) translation are broken through. By integrating multi-modal features of speech, text, vision (lip movement, expression, scene), and combining cross-modal attention mechanism to dynamically align context information, the semantic accuracy of translation is significantly improved. Through the deep integration of multi-modal large models and cross-modal temporal alignment technology, a "semantics-timbre-lip movement-scene" four-in-one video translation is achieved, overcoming the core pain points such as audio-visual disconnection, high latency, and scene disconnection in traditional technologies.

[0123] Figure 2 The following is a schematic diagram of the principle of a video translation and audio-visual synchronization system based on a multi-modal large model provided by an embodiment of this specification. The system may include:

[0124] An acquisition module 10, configured to acquire a source video;

[0125] An extraction module 20, configured to extract the source video based on the multimodal large model to obtain multimodal features;

[0126] A fusion module 30, configured to fuse the multimodal features through a cross-modal attention mechanism to generate a context semantic vector;

[0127] A translation module 40, configured to translate the context semantic vector into a target language text in real time, and process the translated language text based on the multimodal features to obtain a translated language sound source;

[0128] An adaptation module 50, configured to adjust the lip movement of the source video based on the translated language sound source, and merge the translated language sound source and the lip-sync animation video to obtain a real-time translated video with synchronized sound and picture.

[0129] Optionally, the fusion module 30 includes:

[0130] The multimodal features include speech duration features and lip key frame features;

[0131] Based on the speech duration features and the lip key frame features, determine a cross-modal time similarity matrix;

[0132] Overlay the cross-modal time similarity matrix with a mask matrix, and generate attention weights through Softmax;

[0133] Perform weighted fusion on the speech features to generate a context semantic vector.

[0134] Optionally, the translation module 40 includes:

[0135] Based on the context semantic vector, dynamically load a domain knowledge base through a sliding window mechanism to generate a translated text in the target language.

[0136] Optionally, the translation module 40 includes:

[0137] The multimodal features include the source speaker's voiceprint features;

[0138] Based on the source speaker's voiceprint features and the translated language text, drive a text-to-speech model to clone the timbre of the source speaker to obtain a translated language sound source.

[0139] Optionally, the translation module 40 further includes:

[0140] The multimodal features include speech duration features;

[0141] Adjust the speaking speed of the translated speaker based on the speech duration feature and the translated language text to drive the text-to-speech model, and obtain the translated language sound source.

[0142] Optionally, the adaptation module 50 includes:

[0143] Input the translated language sound source into the diffusion model, and output the frame-by-frame lip key point offsets;

[0144] Superimpose the frame-by-frame lip key point offsets on the source video through bilinear interpolation superposition values.

[0145] Optionally, the superimposing the frame-by-frame lip key point offsets on the source video through bilinear interpolation superposition values includes:

[0146] Perform timestamp alignment on the translated language sound source, and compensate for the timing deviation between the translated language sound source and the lip animation video through the dynamic time warping algorithm according to the frame-level alignment relationship output by the cross-modal attention mechanism;

[0147] Perform motion compensation on the non-lip regions of the source video through optical flow estimation;

[0148] Map the lip animation to the source video through bilinear interpolation to generate a lip mask;

[0149] Synthesize a real-time translated video with lip-sync through a residual fusion model.

[0150] The functions of the system according to the embodiments of the present invention have been described in the above method embodiments. Therefore, for the details not described in this embodiment, reference may be made to the relevant descriptions in the foregoing embodiments, and details will not be repeated here.

[0151] The present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (for example, a computer program and a computer program product) for executing some or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0152] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0153] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments.

[0154] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A real-time video translation and audio-visual synchronization method based on a multimodal large model, characterized in that: include: Get the source video; Extracting the source video based on the multimodal large model to obtain multimodal features; The multimodal features are fused through a cross-modal attention mechanism to generate a contextual semantic vector; Translating into a target language text in real time based on the context semantic vector, and processing the translated language text based on the multimodal features to obtain a translated language sound source; The source video is lip-synced based on the translated language sound source, and the translated language sound source and the lip-sync animation video are merged to obtain a real-time translated video with synchronized audio and video.

2. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 1, characterized in that: The multimodal features are fused through a cross-modal attention mechanism to generate a contextual semantic vector, including: The multimodal features include speech duration features and lip shape key frame features; Determining a cross-modal temporal similarity matrix based on the speech duration features and the lip key frame features; The cross-modal temporal similarity matrix is ​​superimposed with the mask matrix, and attention weights are generated through Softmax; The speech features are weighted and fused to generate a contextual semantic vector.

3. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 2, characterized in that: The real-time translation into a target language text based on the context semantic vector includes: dynamically loading a domain knowledge base through a sliding window mechanism based on the context semantic vector to generate a translation text in the target language.

4. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 3, characterized in that: The step of processing the translated language text based on the multimodal features to obtain a translated language sound source includes: The multimodal features include source speaker voiceprint features; The timbre of the source speaker is cloned based on the voiceprint feature of the source speaker and the translated language text-driven text-to-speech model to obtain a translated language sound source.

5. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 4, characterized in that: The step of processing the translated language text based on the multimodal features to obtain a translated language sound source further includes: The multimodal features include speech duration features; The speech speed of the translator is adjusted based on the speech duration feature and the translated language text to drive the text-to-speech model to obtain a translated language sound source.

6. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 5, characterized in that: The lip-sync adjustment of the source video based on the translated language sound source and merging the translated language sound source and the lip-sync animation video to obtain a real-time translated video with synchronized audio and video include: Input the translated language sound source into the diffusion model, and output the frame-by-frame lip key point offset; The frame-by-frame lip key point offsets are superimposed on the source video by bilinear interpolation superposition values.

7. The method for real-time video translation and audio-visual synchronization based on a multimodal large model as claimed in claim 6, characterized in that: The step of superimposing the frame-by-frame lip key point offsets onto the source video by bilinear interpolation superimposition values ​​includes: Performing timestamp alignment on the translated language sound source, and compensating for the timing deviation between the translated language sound source and the lip-sync animation video through a dynamic time warping algorithm according to the frame-level alignment relationship output by the cross-modal attention mechanism; Performing motion compensation on the non-lip-syncing area of ​​the source video by optical flow estimation; Mapping the lip animation to the source video through bilinear interpolation to generate a lip mask; Synthesize real-time translation video with synchronized audio and video through residual fusion model.

8. A video translation and audio-visual synchronization system based on a multimodal large model, characterized in that: include: An acquisition module, used to acquire source video; An extraction module, used for extracting the source video based on the multimodal large model to obtain multimodal features; A fusion module, used to fuse the multimodal features through a cross-modal attention mechanism to generate a contextual semantic vector; A translation module, configured to translate into a target language text in real time based on the context semantic vector, and process the translated language text based on the multimodal features to obtain a translated language sound source; The adaptation module is used to adjust the lip shape of the source video based on the translation language sound source, and merge the translation language sound source and the lip shape animation video to obtain a real-time translation video with synchronized audio and video.

Citation Information

Cited By

  • Multi-language intelligent translation method and system based on short video

    CN121078267A

  • Automatic subtitle translation method and device based on large model and storage medium

    CN121189338A

  • Method and device for automatic translation of subtitles based on large models, and storage medium

    CN121189338B

  • Multi-mode video subtitle and audio collaborative translation and dynamic shunting method and system

    CN121390089A

  • Audio and video real-time intelligent shunting translation method and system combined with buffering strategy

    CN121615662A