Video dubbing language conversion method and system and related equipment

By extracting audio track data from videos, performing voice classification and speech processing, and combining timbre models and translation technology, we have achieved diversified language conversion for video dubbing, solved the language barrier problem, and improved the internationalization of videos and user experience.

CN120932629APending Publication Date: 2025-11-11SHENZHEN MAIFENG TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511035459.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Current technology lacks a method to convert video dubbing language by incorporating the speaker's accent, which prevents videos about language barriers from effectively helping those in need.

Method used

By extracting audio track data from the video to be converted, performing voice extraction and role classification, generating single-speaker audio, performing speech-to-text and voice cloning, achieving target language translation and text-to-speech, and finally replacing the video audio track to obtain a dubbed converted video.

Benefits of technology

It enables video dubbing that translates language based on the speaker's voice timbre, increasing the diversity of videos, meeting user needs, and reducing production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932629A_ABST
    Figure CN120932629A_ABST
Patent Text Reader

Abstract

The invention provides a video dubbing language conversion method, a video dubbing language conversion system and related equipment. The method comprises the following steps: acquiring audio track data from a video to be converted; carrying out human voice extraction on the audio track data and classifying according to roles to obtain a single speaker audio of each role; performing voice-to-text conversion on the single speaker audio of each role to obtain an original language copywriting of each role; performing sound cloning on the single speaker audio of each role to obtain a timbre model of each role; performing target language translation on the original language copywriting of each role to obtain a translated copywriting of each role; based on the translation copywriting of each role and the tone model of each role, performing text-to-voice conversion to obtain a translation audio of each role; and performing replacement of each role translation audio on the audio track data in the to-be-converted video to obtain a dubbing conversion video. According to the technical scheme, language video dubbing conversion combined with the tone of the speaker is achieved, the video is more diversified, and the user requirements can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text-to-speech conversion technology, and in particular to a method, system, and related equipment for converting video dubbing language. Background Technology

[0002] Many high-quality videos on the market have their effectiveness significantly reduced due to language barriers. Many educational or explanatory videos fail to provide genuine help to those who need it because the corresponding language is incomprehensible. If it were possible to automatically translate and generate corresponding multilingual audio based on each person's accent and automatically dub it, this problem could be solved. Multilingual translation combining speaker accents with video would make videos more diverse. Current technology lacks a method for converting video dubbing languages ​​based on speaker accents.

[0003] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0004] This invention provides a method, system, and related equipment for converting video dubbing languages. The main objective of this invention is to solve the technical problems mentioned in the background section of the prior art.

[0005] The first aspect of this invention provides a method for converting video dubbing languages, comprising: Obtain the audio track data from the video to be converted; The audio track data is processed to extract human voices and categorized by role to obtain individual speaker audio for each role; The audio of each character's single speaker is converted into text to obtain the original language script for each character; The voice cloning of the single-speaker audio of each character is performed to obtain the timbre model of each character; The original language texts of each character are translated into the target language to obtain the translated texts for each character; Based on the translation texts and voice models of each character, text-to-speech conversion is performed to obtain the translated audio of each character; The audio track data in the video to be converted is replaced with the translated audio for each character to obtain a dubbed converted video.

[0006] In an optional embodiment of the first aspect of the present invention, the step of extracting human voices from the audio track data and classifying them by role to obtain individual speaker audio for each role includes: The audio track data is subjected to VAD (Voice Activity Detection) to identify audio segments containing human voices; Voiceprint features are extracted from the identified audio segments to obtain the voiceprint features of each speaker. Based on the voiceprint features, cluster analysis is performed on the speakers of each role, and the audio segments of the same speaker are grouped into one category; The audio segments of each category are sorted chronologically to generate single-speaker audio for each character.

[0007] In an optional embodiment of the first aspect of the present invention, the step of converting the single-speaker audio of each character into speech-to-text to obtain the original language script of each character includes: The single-speaker audio of each character is preprocessed, including noise reduction, audio format conversion, and audio normalization. A deep learning-based speech recognition model is used to perform speech-to-text conversion on the preprocessed single-speaker audio to obtain preliminary text. The initial text is then post-processed, including adding punctuation marks, semantic correction, and formatting. Generate structured raw language copy containing character identifiers and corresponding text content.

[0008] In an optional embodiment of the first aspect of the present invention, the step of cloning the voice of each character's single-speaker audio to obtain a timbre model for each character includes: Phonic features are extracted from the single-speaker audio of each character, and the timbre features include fundamental frequency, formants and spectral envelope; The vocoder model is trained based on the extracted timbre features to generate an initial timbre model; The initial timbre model is optimized using a variational autoencoder or a generative adversarial network to improve timbre fidelity and stability, thereby obtaining timbre models for each character. A timbre model library is established to store the timbre models of each character.

[0009] In an optional embodiment of the first aspect of the present invention, the step of translating the original language text of each character into the target language to obtain the translated text of each character includes: The original language texts of each character are preprocessed, including word segmentation, part-of-speech tagging, and syntactic analysis. A neural network-based multilingual translation model is used to translate the original language texts of each character after the text preprocessing, and to obtain preliminary translation results. The preliminary translation results are then subjected to further adjustments, including grammatical correction, semantic optimization, and format adjustment. Generate structured translation text that includes character identifiers and corresponding translation content.

[0010] In an optional embodiment of the first aspect of the present invention, the step of performing text-to-speech conversion based on the translated text of each character and the timbre model of each character to obtain the translated audio of each character includes: The translated texts for each character are preprocessed using speech synthesis, which includes text cleaning, phoneme conversion, and prosodic analysis. Based on the timbre model of each character, a neural network vocoder is used to generate the speech waveform of each character; The generated speech waveforms of each character are then subjected to further optimization processing, including noise reduction, volume normalization, and audio format conversion, to obtain the translated audio of each character. Generate a structured audio file containing character identifiers and corresponding translated audio.

[0011] In an optional embodiment of the first aspect of the present invention, replacing the translated audio of each character in the audio track data of the video to be converted to obtain a dubbed converted video includes: The audio track data of the video to be converted is segmented to obtain the original audio segments of each character and their corresponding timestamps; Based on the timestamp, the translated audio of each character is aligned with the corresponding original audio segment; Using audio editing technology, the translated audio of each character replaces the corresponding original audio segment, while maintaining synchronization with the video footage; The translated audio of each character after replacement is processed with volume matching and reverb to obtain a dubbing conversion video that is synchronized with the video and has natural sound effect transitions.

[0012] A second aspect of the present invention provides a video dubbing language conversion system, the video dubbing language conversion system comprising: The audio track data acquisition module is used to obtain audio track data from the video to be converted; The character audio acquisition module is used to extract human voices from the audio track data and classify them according to the character to obtain the single-speaker audio of each character. The original text acquisition module is used to convert the single-speaker audio of each character into text to obtain the original language text of each character. The character voice cloning module is used to clone the voice of each character's single-speaker audio to obtain the timbre model of each character. The text translation and conversion module is used to translate the original language texts of each character into the target language to obtain the translated texts for each character. The translation audio acquisition module is used to convert text to speech based on the translation text of each character and the timbre model of each character to obtain the translation audio of each character; The translation and dubbing replacement module is used to replace the translated audio of each character in the audio track data of the video to be converted, so as to obtain a dubbed converted video.

[0013] A third aspect of the present invention provides a video dubbing language conversion device, the video dubbing language conversion device comprising: a memory and at least one processor, the memory storing instructions, and the memory and the at least one processor being interconnected via a line; The at least one processor invokes the instructions in the memory to cause the video dubbing language conversion device to perform the video dubbing language conversion method as described in any one of the first aspects of the present invention.

[0014] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a video dubbing language conversion method as described in any one of the first aspects of the present invention.

[0015] Beneficial Effects: This invention provides a method, system, and related equipment for converting video dubbing into spoken language. The method includes: acquiring audio track data from the video to be converted; extracting human voices from the audio track data and classifying them by role to obtain individual speaker audio for each role; converting the individual speaker audio for each role into speech-to-text to obtain the original language script for each role; cloning the individual speaker audio for each role to obtain the timbre model for each role; translating the original language script for each role into the target language to obtain the translated script for each role; converting text into speech based on the translated script and the timbre model for each role to obtain the translated audio for each role; and replacing the translated audio for each role in the audio track data of the video to be converted to obtain the dubbed video. This invention achieves language-switching video dubbing by combining speaker timbre, resulting in more diverse videos and better meeting user needs. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of an embodiment of a video dubbing language conversion method according to the present invention; Figure 2 This is a schematic diagram of an embodiment of a video dubbing language conversion system according to the present invention; Figure 3 This is a schematic diagram of an embodiment of a video dubbing language conversion device according to the present invention. Detailed Implementation

[0017] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0018] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 The first aspect of this invention provides a method for converting video dubbing languages, comprising: S100. Obtain audio track data from the video to be converted. In this invention, the video to be converted can be a single-speaker video or a multi-speaker video. The inventive point of this invention is to convert the dubbing of the video to be converted from the original language (e.g., English) to the target language (e.g., Chinese) while retaining the original voice of the dubbing speaker, so that the transition after the video language translation is more natural and better accepted by users. The video translation mainly relies on the audio in the video file, that is, in step S100, this invention needs to obtain audio track data with time parameters from the video to be converted.

[0019] S200: Extract human voices from the audio track data and classify them by role to obtain the single-speaker audio for each role. In this invention, after obtaining the audio track data from the video to be converted, in order to better perform dubbing, it is necessary to extract the human voice from the audio track data. When there are multiple speakers in the audio track data, different speakers will have different timbres, so it is also necessary to classify the extracted human voices by role.

[0020] In an optional embodiment of step S200 of the present invention, the step of extracting human voices from the audio track data and classifying them by role to obtain single-speaker audio for each role includes: performing VAD (Voice Activity Detection) on the audio track data to identify audio segments containing human voices; extracting voiceprint features from the identified audio segments to obtain voiceprint features of each role's speaker; performing cluster analysis on the speakers of each role based on the voiceprint features to group the audio segments of the same speaker into one category; and performing temporal sorting on the audio segments of each category to generate single-speaker audio for each role.

[0021] Specifically, in this invention, a pre-trained deep learning VAD model (e.g., an end-to-end network based on CNN-LSTM or a deep learning Spleeter audio separation tool based on U-Net) can be used to analyze audio frames in real time. The model takes the Mel-spectrum of the audio as input and outputs a binary classification label (0: non-human voice, 1: human voice) for each frame. Frames with consecutive outputs of 1 are merged into human voice audio segments, and their start and end timestamps are recorded. Then, for all the obtained human voice audio segments, a deep neural network voiceprint encoder (such as a ResNet34-TDNN structure) is used to extract fixed-dimensional voiceprint embedding vectors. The encoder takes the linear frequency spectrum of the audio segments as input and extracts biometric features related to the speaker's identity through convolutional and temporal delay layers. The output vector satisfies the following conditions: cosine similarity between segments of the same speaker > 0.8, and similarity between segments of different speakers < 0.2. Then, the voiceprint embedding vectors of all audio segments are input into an unsupervised clustering algorithm (such as improved spectral clustering or density-based DBSCAN algorithm). By calculating the cosine distance between vectors and setting an adaptive threshold, audio segments belonging to the same speaker are automatically classified. Temporal continuity constraints are introduced to prioritize merging segments within adjacent time windows to avoid clustering errors caused by short-term speech breaks. Finally, for each clustered audio segment set, the segments are arranged in ascending order according to the original timestamps. A continuous single-role speaker audio stream is generated through cross-gradient splicing technology and output as an independent audio file, ultimately obtaining a structured audio library with the same number of speaker roles as the original audio track.

[0022] S300. The individual speaker audio of each character is converted to text to obtain the original speech text for each character. In this invention, after extracting the individual speaker audio of each character, the Speech To Text interface is used for audio-to-text conversion. The text of each character's audio in the video is obtained, along with the time parameters of each dialogue segment (for example, in a video where Xiaoming and Xiaohong are talking, using Speech To Text will produce text fragments like "Character A [00:00 - 00:05]: Hello, what were you doing this morning? Character B [00:05-00:15]: I'm at work, I have a lot of work today."). This allows us to know the time, character, and text content of each dialogue segment. The text is then categorized by character, time, and number of segments and displayed on the software, allowing uploaders and producers to double-check and reduce errors in audio-to-text conversion.

[0023] In an optional embodiment of step S300 of the present invention, the step of converting the single-speaker audio of each character into speech-to-text to obtain the original language text of each character includes: preprocessing the single-speaker audio of each character, the preprocessing including noise reduction, audio format conversion and audio standardization; performing speech-to-text processing on the preprocessed single-speaker audio based on a deep learning-based speech recognition model to obtain preliminary text; performing postprocessing on the preliminary text, the postprocessing including punctuation addition, semantic correction and formatting; and generating a structured original language text containing character identifiers and corresponding text content.

[0024] Specifically, in this invention, noise reduction of the single-speaker audio of each character can be achieved using spectral subtraction based on deep noise suppression (DNS). For example, a pre-trained U-Net structure model can be used, with input Mel-frequency cepstral coefficients (MFCCs) containing noise frequencies, and the output noise mask can be spectrally subtracted in the frequency domain to preserve the fundamental frequency and harmonic structure. Speech-to-text processing can be achieved using an end-to-end speech recognition model with a Conformer-Transformer hybrid structure (convolutional layers extract local features + self-attention layers capture long dependencies), and a preliminary text sequence can be generated through a CTC / Attention joint decoding mechanism. Semantic correction can be achieved by using an N-gram language model to correct homophones and by using a named entity recognition module to correct proper nouns.

[0025] S400. Perform voice cloning on the individual speaker audio of each character to obtain the timbre model of each character. In this invention, in order to ensure that the dubbing of the translated text also has the timbre of the original speaker, this step requires using the previously separated individual speaker audio of each character to perform voice cloning to obtain the corresponding character's voice model (here, the So-VITS-SVCAI cloning scheme can be used, first extracting Hubert features based on the separated voice, then using a timbre encoder to extract the target speaker's timbre, and training the voice model of the specified voice by calculating pitch characteristics, etc.). This allows for effect previewing and supports retrying cloning to correct for voice discrepancies.

[0026] In an optional embodiment of step S400 of the present invention, the step of cloning the voice of the single speaker audio of each character to obtain the timbre model of each character includes: extracting timbre features from the single speaker audio of each character, wherein the timbre features include fundamental frequency, formants and spectral envelope; training a vocoder model based on the extracted timbre features to generate an initial timbre model; optimizing the initial timbre model using a variational autoencoder or a generative adversarial network to improve timbre fidelity and stability, thereby obtaining the timbre model of each character; and establishing a timbre model library to store the timbre models of each character.

[0027] Specifically, the core principle of voice cloning in this invention is to decouple timbre from pronunciation content and reconstruct the target timbre using a generative model. Timbre feature extraction and decoupling can be achieved using a SoftVC encoder with an improved VITS structure. The content encoding part extracts pronunciation features (phonemes / pitch / rhythm) that are unrelated to timbre, while the timbre encoding part extracts speaker identity features (spectral envelope / formant distribution). Speaker timbre transfer can be achieved using a conditional vocoder based on generative adversarial networks (GANs) to reconstruct the target timbre. The input is a latent variable of the content plus a target timbre vector. The generator uses HiFi-GAN or DiffWave, injecting timbre conditions (such as AdaIN) through an adaptive normalization layer. A discriminator distinguishes between real speech and generated speech. The loss function includes multi-scale STFT loss, feature matching loss, and adversarial loss.

[0028] S500: Translate the original language text for each character into the target language to obtain the translated text for each character. In this invention, this step involves using a translation engine (supporting multiple languages) to translate the original language text into the target language. After translation, manual review can be conducted to ensure the accuracy of the translation. Additionally, excessively long translated texts can be trimmed to ensure that the length of the translated audio does not deviate significantly.

[0029] In an optional embodiment of step S500 of the present invention, the step of translating the original language text of each character into the target language to obtain the translated text of each character includes: preprocessing the original language text of each character, the text preprocessing including word segmentation, part-of-speech tagging and syntactic analysis; translating the preprocessed original language text of each character using a multilingual translation model based on a neural network to obtain a preliminary translation result; performing subsequent adjustment processing on the preliminary translation result, the subsequent adjustment processing including grammatical correction, semantic optimization and format adjustment; and generating a structured translated text containing character identifiers and corresponding translated content.

[0030] Specifically, in this invention, for word segmentation, the jieba segmenter can be used for Chinese, with a custom dictionary containing technical terms; the MeCab segmenter can be used for Japanese, with character names added as proper nouns; the NLTK segmenter can be used for English, handling abbreviations and hyphens; part-of-speech tagging can use the BERT-POS model, with the tagging format being Universal POSTagset; syntactic analysis can use BERT-Parser for dependency parsing, generating a dependency tree structure to facilitate maintaining semantic relationships during subsequent translation; the translation model can adopt a dynamic routing translation model architecture; the encoder can use the XLM-RoBERTa base model + domain adaptation layer; and the decoder can use the Transformer-XL structure + character context caching mechanism.

[0031] S600: Based on the translated text and timbre model of each character, perform text-to-speech conversion to obtain the translated audio of each character. In this step of the invention, the timbre model obtained in step S400 and the translated text obtained in step S500 are used to perform text-to-speech audio conversion with timbre. This step can also use the VITS TTS scheme. First, the obtained timbre model is used to know the phonemes and timbre characteristics of the target timbre. Then, the VITS variational autoencoder (VAE-VITS) is used to generate the speech waveform of the target timbre. Finally, the HiFi-GAN vocoder is used to convert the Mel spectrum obtained in the VAE into the real audio waveform to obtain the sound effect of the character's language translation. Of course, the audio generated in this step of the invention can also support listening and re-conversion, allowing users to convert multiple times to correct the effect.

[0032] In an optional embodiment of step S600 of the present invention, the step of converting text to speech based on the translation text of each character and the timbre model of each character to obtain the translated audio of each character includes: performing speech synthesis preprocessing on the translation text of each character, the speech synthesis preprocessing including text cleaning, phoneme conversion and prosodic analysis; generating speech waveforms of each character using a neural network vocoder based on the timbre model of each character; performing subsequent optimization processing on the generated speech waveforms of each character, the subsequent optimization processing including noise reduction, volume normalization and audio format conversion to obtain the translated audio of each character; and generating a structured audio file containing character identifiers and corresponding translated audio.

[0033] Specifically, the text cleaning in the preprocessing of this invention mainly includes removing special characters, processing abbreviations, and normalizing numbers; phoneme conversion can use a G2P model based on the Transformer architecture, combined with the IPA phoneme set; prosodic analysis can use Tacotron2 to extract prosodic features, and the timbre model contains the timbre features of the characters. By weighting and incorporating the timbre features of the timbre model during the process of encoding the translated text into audio, speech waveforms with the timbre of each character can be output; the vocoder architecture can adopt the HiFi-GAN model, which includes a multi-scale generator and a multi-scale discriminator, to further improve the output quality of the audio after incorporating the timbre.

[0034] S700: Replace the translated audio of each character in the audio track data of the video to be converted to obtain a dubbed converted video. In this invention, this step involves optimizing the duration of the dialogue segments of the acquired translated audio of each character, adjusting the speech rate to correct the audio-visual asynchrony caused by the length of multiple languages ​​after conversion, retaining the background sound while replacing the character's pronunciation, and generating a video to complete the video translation function.

[0035] In an optional embodiment of step S700 of the present invention, replacing the translated audio of each character in the audio track data of the video to be converted to obtain a dubbing-converted video includes: performing audio segmentation on the audio track data of the video to be converted to obtain the original audio segments of each character and their corresponding timestamps; aligning the translated audio of each character with the corresponding original audio segments based on the timestamps; using audio editing technology to replace the corresponding original audio segments with the translated audio of each character while maintaining synchronization with the video screen; and performing volume matching and reverb processing on the replaced translated audio of each character to obtain a dubbing-converted video with synchronized video and natural sound effect transitions.

[0036] Specifically, the audio segmentation step in this invention can also use a deep learning model (such as the VAD model) to perform real-time speech detection on the video audio track, and use a role recognition algorithm (such as a DNN model based on voiceprint recognition) to distinguish different speakers, generate timestamp information accurate to the millisecond level, and record the speaking interval of each role; example output format: [{Role ID:1, Start Time:0.000, End Time:2.345}, …]; Audio alignment uses timestamp information as the alignment benchmark and applies the Dynamic Time Warping (DTW) algorithm to ensure the temporal correspondence between the translated audio and the original audio, achieving millisecond-level precise alignment while preserving the rhythm and rhyme characteristics of the original audio; Audio replacement can employ seamless audio splicing technology and use fade-in / fade-out effects to handle audio boundaries, enabling independent audio replacement based on characters, supporting simultaneous processing of multiple characters, and maintaining the synchronization between video and audio; Sound effects processing can use an adaptive volume matching algorithm to adjust the translated audio based on the volume distribution of the original audio, apply intelligent reverb processing to simulate the original recording environment, and support multi-channel processing to maintain the spatial sense of the audio, achieve smooth audio transitions, and avoid abruptness.

[0037] In simple terms, the main steps of this invention include: (1) Step 1: First, it is necessary to separate the human voice and background sound in the clip based on the original video audio track. Extract the character's voice and use the Speech To Text interface to convert the voice to text. Obtain the text of each character's audio in the video, and obtain the time parameters of each dialogue segment, etc., classify them by character, time, and number of segments, and display them on the software so that the uploader and producer can double-check and reduce the error in voice-to-text conversion.

[0038] (2) Step 2: Extract the voiceprint of the character. Use the voices of each character that were previously separated to clone the voices and obtain the voice model of the corresponding character. The audition effect is given so that users can audition the effect. At the same time, it supports retrying the cloning to correct the problem of the voices not being similar.

[0039] (3) Step 3: Based on the previously separated and proofread text, use translation APIs (such as Google or others) to perform multilingual translation, and display it on the interface for users to check and verify, ensuring that the translation is correct. At the same time, excessively long texts can be trimmed to ensure that the length of the translated audio does not deviate too much.

[0040] (4) Step 4: Using the translated text and the character's voice model obtained above, perform Text To Speech audio conversion to obtain the character's multilingual translation sound effect, which supports listening and re-conversion, allowing users to convert multiple times to correct the effect.

[0041] (5) Step 5: Optimize the duration of each dialogue obtained in Step 1, fine-tune the speech speed to correct the audio-visual asynchrony caused by the length of multiple languages ​​after conversion, retain the background sound to replace the character's pronunciation, generate the video, and the video translation function can be completed.

[0042] In summary, the technical solution of this invention solves the problems of video dubbing translation and internationalization, greatly reduces production costs, liberates productivity, and allows video production to focus more on quality and content.

[0043] See Figure 2 The second aspect of the present invention provides a video dubbing language conversion system, the video dubbing language conversion system comprising: Audio track data acquisition module 10 is used to obtain audio track data from the video to be converted; The character audio acquisition module 20 is used to extract human voices from the audio track data and classify them according to the character to obtain the single-speaker audio of each character. The original text acquisition module 30 is used to convert the single speaker audio of each character into text to obtain the original language text of each character. The character voice cloning module 40 is used to clone the voice of each character's single speaker audio to obtain the timbre model of each character. The text translation and conversion module 50 is used to translate the original language texts of each character into the target language to obtain the translated texts of each character. The translation audio acquisition module 60 is used to convert text to speech based on the translation text of each character and the timbre model of each character to obtain the translation audio of each character; The translation and dubbing replacement module 70 is used to replace the translated audio of each character in the audio track data of the video to be converted, so as to obtain a dubbing-converted video.

[0044] In an optional embodiment of the second aspect of the present invention, the character audio acquisition module includes: A voice activity detection unit is used to perform VAD voice activity detection on the audio track data and identify audio segments containing human voices. The voiceprint feature extraction unit is used to extract voiceprint features from the identified audio segments and obtain the voiceprint features of each speaker. The clustering analysis unit is used to perform clustering analysis on the speakers of each role based on the voiceprint features, and to group the audio segments of the same speaker into one category. An audio sorting unit is used to sort the audio segments of each category in chronological order to generate single-speaker audio for each character.

[0045] In an optional embodiment of the second aspect of the present invention, the original text acquisition module includes: An audio preprocessing unit is used to preprocess the single-speaker audio of each character, the preprocessing including noise reduction, audio format conversion and audio normalization; The text conversion unit is used to perform speech-to-text processing on the preprocessed single-speaker audio based on a deep learning-based speech recognition model to obtain preliminary text. The post-processing unit is used to perform post-processing on the preliminary text, including adding punctuation marks, semantic correction, and formatting. The original text generation unit is used to generate structured original language text that includes role identifiers and corresponding text content.

[0046] In an optional embodiment of the second aspect of the present invention, the character voice cloning module includes: Timbre feature extraction is used to extract timbre features from the single-speaker audio of each character. The timbre features include fundamental frequency, formants, and spectral envelope. The model training unit is used to train the vocoder model based on the extracted timbre features and generate an initial timbre model. The model optimization unit is used to optimize the initial timbre model using a variational autoencoder or a generative adversarial network to improve timbre fidelity and stability, and obtain the timbre model of each character. A model storage unit is used to establish a timbre model library to store the timbre models of each character.

[0047] In an optional embodiment of the second aspect of the present invention, the text translation and conversion module includes: The text preprocessing unit is used to preprocess the original language text of each character. The text preprocessing includes word segmentation, part-of-speech tagging, and syntactic analysis. The translation processing unit is used to translate the original language texts of each character after the text preprocessing based on a neural network-based multilingual translation model to obtain preliminary translation results. The result adjustment unit is used to perform subsequent adjustment processing on the preliminary translation result, including grammatical correction, semantic optimization and format adjustment. The translation copy generation unit is used to generate structured translation copy that includes character identifiers and corresponding translation content.

[0048] In an optional embodiment of the second aspect of the present invention, the translated audio acquisition module includes: The preprocessing unit is used to perform speech synthesis preprocessing on the translated texts of each character. The speech synthesis preprocessing includes text cleaning, phoneme conversion and prosodic analysis. A speech synthesis unit is used to generate speech waveforms for each character using a neural network vocoder based on the timbre model of each character. A waveform optimization unit is used to perform subsequent optimization processing on the generated speech waveforms of each character. The subsequent optimization processing includes noise reduction, volume normalization, and audio format conversion to obtain the translated audio of each character. An audio file generation unit is used to generate a structured audio file containing a character identifier and the corresponding translated audio.

[0049] In an optional embodiment of the second aspect of the present invention, the translation dubbing replacement module includes: The audio segmentation unit is used to segment the audio track data of the video to be converted and obtain the original audio segments of each character and their corresponding timestamps. Time alignment is used to align the translated audio of each character with the corresponding original audio segment based on the timestamp; Audio replacement is used to replace the original audio segments with the translated audio of each character using audio editing techniques, while maintaining synchronization with the video footage; The audio processing unit is used to perform volume matching and reverb processing on the translated audio of each character after replacement, so as to obtain a dubbing conversion video with video synchronization and natural sound effect transition.

[0050] Figure 3This is a schematic diagram of the structure of a video dubbing language conversion device provided in an embodiment of the present invention. The video dubbing language conversion device can vary significantly due to differences in configuration or performance, and may include one or more processors 80 (central processing units, CPUs) (e.g., one or more processors) and memory 90, and one or more storage media 100 (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the video dubbing language conversion device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the video dubbing language conversion device.

[0051] The video dubbing language conversion device of the present invention may further include one or more power supplies 110, one or more wired or wireless network interfaces 120, one or more input / output interfaces 130, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated structure of the video dubbing language conversion device does not constitute a limitation on the video dubbing language conversion device. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0052] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the video dubbing language conversion system.

[0053] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system or system / unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0054] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a video dubbing language conversion device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0055] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for converting video dubbing languages, characterized in that, include: Obtain the audio track data from the video to be converted; The audio track data is processed to extract human voices and categorized by role to obtain individual speaker audio for each role; The audio of each character's single speaker is converted into text to obtain the original language script for each character; The voice cloning of the single-speaker audio of each character is performed to obtain the timbre model of each character; The original language texts of each character are translated into the target language to obtain the translated texts for each character; Based on the translation texts and voice models of each character, text-to-speech conversion is performed to obtain the translated audio of each character; The audio track data in the video to be converted is replaced with the translated audio for each character to obtain a dubbed converted video.

2. The video dubbing language conversion method according to claim 1, characterized in that, The step of extracting human voices from the audio track data and classifying them by role to obtain individual speaker audio for each role includes: The audio track data is subjected to VAD (Voice Activity Detection) to identify audio segments containing human voices; Voiceprint features are extracted from the identified audio segments to obtain the voiceprint features of each speaker. Based on the voiceprint features, cluster analysis is performed on the speakers of each role, and the audio segments of the same speaker are grouped into one category; The audio segments of each category are sorted chronologically to generate single-speaker audio for each character.

3. The video dubbing language conversion method according to claim 1, characterized in that, The process of converting the individual speaker audio of each character into text to obtain the original language script for each character includes: The single-speaker audio of each character is preprocessed, including noise reduction, audio format conversion, and audio normalization. A deep learning-based speech recognition model is used to perform speech-to-text conversion on the preprocessed single-speaker audio to obtain preliminary text. The initial text is then post-processed, including adding punctuation marks, semantic correction, and formatting. Generate structured raw language copy containing character identifiers and corresponding text content.

4. The video dubbing language conversion method according to claim 1, characterized in that, The step of cloning the individual speaker audio of each character to obtain the timbre model of each character includes: Phonic features are extracted from the single-speaker audio of each character, and the timbre features include fundamental frequency, formants and spectral envelope; The vocoder model is trained based on the extracted timbre features to generate an initial timbre model; The initial timbre model is optimized using a variational autoencoder or a generative adversarial network to improve timbre fidelity and stability, thereby obtaining timbre models for each character. A timbre model library is established to store the timbre models of each character.

5. The video dubbing language conversion method according to claim 1, characterized in that, The process of translating the original language texts for each character into their target language to obtain the translated texts for each character includes: The original language texts of each character are preprocessed, including word segmentation, part-of-speech tagging, and syntactic analysis. A neural network-based multilingual translation model is used to translate the original language texts of each character after the text preprocessing, and to obtain preliminary translation results. The preliminary translation results are then subjected to further adjustments, including grammatical correction, semantic optimization, and format adjustment. Generate structured translation text that includes character identifiers and corresponding translation content.

6. The video dubbing language conversion method according to claim 1, characterized in that, The text-to-speech conversion based on the translated text and voice model of each character to obtain the translated audio for each character includes: The translated texts for each character are preprocessed using speech synthesis, which includes text cleaning, phoneme conversion, and prosodic analysis. Based on the timbre model of each character, a neural network vocoder is used to generate the speech waveform of each character. The generated speech waveforms of each character are then subjected to further optimization processing, including noise reduction, volume normalization, and audio format conversion, to obtain the translated audio of each character. Generate a structured audio file containing character identifiers and corresponding translated audio.

7. The video dubbing language conversion method according to claim 1, characterized in that, The step of replacing the audio track data of each character in the video to be converted to obtain the dubbed video includes: The audio track data of the video to be converted is segmented to obtain the original audio segments of each character and their corresponding timestamps; Based on the timestamp, the translated audio of each character is aligned with the corresponding original audio segment; Using audio editing technology, the translated audio of each character is replaced with the corresponding original audio segment, while maintaining synchronization with the video footage; The translated audio of each character after replacement is processed with volume matching and reverb to obtain a dubbing conversion video that is synchronized with the video and has natural sound effect transitions.

8. A video dubbing language conversion system, characterized in that, The video dubbing language conversion system includes: The audio track data acquisition module is used to obtain audio track data from the video to be converted; The character audio acquisition module is used to extract human voices from the audio track data and classify them according to the character to obtain the single-speaker audio of each character. The original text acquisition module is used to convert the single-speaker audio of each character into text to obtain the original language text of each character. The character voice cloning module is used to clone the voices of each character's single-speaker audio to obtain the timbre model of each character. The text translation and conversion module is used to translate the original language texts of each character into the target language to obtain the translated texts for each character. The translation audio acquisition module is used to convert text to speech based on the translation text of each character and the timbre model of each character to obtain the translation audio of each character; The translation and dubbing replacement module is used to replace the translated audio of each character in the audio track data of the video to be converted, so as to obtain a dubbed converted video.

9. A video dubbing language conversion device, characterized in that, The video dubbing language conversion device includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor invokes the instructions in the memory to cause the video dubbing language conversion device to perform the video dubbing language conversion method as described in any one of claims 1-7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the video dubbing language conversion method as described in any one of claims 1-7.

Citation Information

Cited By

  • Video translation method and device

    CN121814981A

  • Video decoding method and apparatus

    CN121814981B