A dubbing method, apparatus, electronic device and storage medium

By using speech synthesis and timbre conversion technologies, automated cross-language dubbing of film and television works can be achieved, solving the problems of high dubbing costs and talent shortage in film and television works, and improving dubbing efficiency and quality.

CN119815149BActive Publication Date: 2025-10-31YOUKU CULTURE TECH (BEIJING) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510025048.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-10-31
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently achieve cross-language dubbing for film and television works, especially dubbing for less commonly spoken languages, resulting in high costs for manual dubbing and a shortage of dubbing talent.

Method used

By using speech synthesis, timbre conversion, and audio-visual synchronization processing, cross-language audio data is generated and synchronized with video to achieve automated dubbing.

Benefits of technology

It has reduced the cost of dubbing for film and television works, improved the efficiency of dubbing production, and met the needs of large-scale film and television works going overseas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815149B_ABST
    Figure CN119815149B_ABST
Patent Text Reader

Abstract

This disclosure relates to a dubbing method, apparatus, electronic device, and storage medium. The method includes: acquiring a target dubbing script for a video to be dubbed; wherein the target dubbing script includes at least one line of dialogue and the corresponding time and character for each line; the language of the target dubbing script is different from the language of the video to be dubbed; generating target audio data by performing speech synthesis on the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script; performing timbre conversion on the target audio data; and performing audio-visual synchronization processing between the timbre-converted audio data and the video to be dubbed to generate a dubbed video. In this disclosure, through speech synthesis, timbre conversion, and audio-visual synchronization, automated cross-language dubbing of videos is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech processing technology, and in particular to a dubbing method, apparatus, electronic device, and storage medium. Background Technology

[0002] To facilitate the dissemination and promotion of films and television programs in different language countries, dubbing is necessary. However, relying solely on human dubbing methods is insufficient to meet the large demand for dubbing in film and television productions. Summary of the Invention

[0003] In view of this, the present disclosure provides a dubbing method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to one aspect of this disclosure, a dubbing method is provided, the method comprising:

[0005] Obtain the target dubbing script for the video to be dubbed; wherein, the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line of dialogue; the language of the target dubbing script is different from the language of the video to be dubbed;

[0006] Target audio data is generated by synthesizing the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script.

[0007] Perform timbre conversion on the target audio data;

[0008] The audio data after timbre conversion is synchronized with the video to be dubbed to generate a dubbed video.

[0009] In one possible implementation, generating target audio data by performing speech synthesis on the target dubbing script includes:

[0010] Extract the prosodic and phonemic features of the target dubbing script;

[0011] The prosodic features and phonemic features of the target dubbing script are fused to generate text features;

[0012] Speech synthesis is performed based on the text features to generate the target audio data.

[0013] In one possible implementation, the method further includes:

[0014] Extract the emotional features of the target dubbing script;

[0015] The step of generating the target audio data by performing speech synthesis based on the text features includes:

[0016] Based on the emotional features and text features of the target dubbing script, speech synthesis is performed to generate the target audio data.

[0017] In one possible implementation, the timbre conversion of the target audio data includes:

[0018] Among a plurality of preset reference timbres, a target timbre that matches at least one character corresponding to the target audio data is determined; the languages ​​corresponding to the plurality of reference timbres are the same as the languages ​​corresponding to the target audio data.

[0019] Based on the target timbre matched to the at least one character, the target audio data is timbre converted.

[0020] In one possible implementation, extracting the phoneme features of the target dubbing script includes:

[0021] Based on a preset phoneme dictionary, a phoneme sequence corresponding to the target dubbing script is generated;

[0022] Extract the phoneme features corresponding to each phoneme in the phoneme sequence;

[0023] The process of fusing the prosodic and phonemic features of the target dubbing script to generate text features includes:

[0024] The prosodic features of the target dubbing script are processed to determine the prosodic features corresponding to each phoneme in the phoneme sequence;

[0025] The prosodic features corresponding to each phoneme in the phoneme sequence are fused with the phoneme features corresponding to each phoneme to generate the text features.

[0026] In one possible implementation, extracting the emotional features of the target dubbing script includes:

[0027] The emotional features of the target dubbing script are extracted using an emotion recognition model. The emotion recognition model is trained using a transfer learning algorithm based on a preset emotion space, which includes at least the following emotions: joy, sadness, surprise, anger, disappointment, pride, jealousy, rage, bittersweetness, and happiness.

[0028] In one possible implementation, the timbre conversion of the target audio data based on the target timbre matched with the at least one role includes:

[0029] Obtain the preset timbre features corresponding to the target timbre matched by the at least one character;

[0030] Extract semantic features corresponding to at least one character from the target audio data;

[0031] By fusing the preset timbre features with the semantic features, audio data after timbre conversion is generated.

[0032] In one possible implementation, the step of performing audio-visual synchronization processing between the timbre-converted audio data and the video to be dubbed, to generate the dubbed video, includes:

[0033] The audio data after timbre conversion is decomposed into multiple audio segments; each audio segment corresponds to a line of dialogue.

[0034] Determine the processing method for each audio segment; wherein the processing method includes one or more of the following: remain unchanged, speed up, axis merging, and axis shifting, where axis merging means merging two adjacent audio segments into one audio segment, axis shifting means moving the start and / or end time of the audio segment, and speeding up means shortening the duration of the audio segment.

[0035] The multiple audio segments are adjusted based on the processing method of each audio segment, and the adjusted audio segments are aligned with the scenes in the video to be dubbed in chronological order to generate the dubbed video.

[0036] In one possible implementation, the method further includes:

[0037] Obtain the source audio data of the video to be dubbed; the language of the source audio data is the same as the language of the video to be dubbed.

[0038] The source audio data is subjected to speech recognition to generate a source dubbing script;

[0039] Translate the source dubbing script to generate the target dubbing script.

[0040] According to another aspect of this disclosure, a dubbing device is provided, the device comprising:

[0041] The acquisition module is used to acquire the target dubbing script of the video to be dubbed; wherein, the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line of dialogue; the language of the target dubbing script is different from the language of the video to be dubbed;

[0042] The speech synthesis module is used to generate target audio data by performing speech synthesis on the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script.

[0043] A timbre conversion module is used to convert the timbre of the target audio data;

[0044] The audio-visual synchronization module is used to perform audio-visual synchronization processing between the audio data after timbre conversion and the video to be dubbed, so as to generate the dubbed video.

[0045] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-described method when executing instructions stored in the memory.

[0046] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the above-described method.

[0047] According to another aspect of this disclosure, a computer program product is provided, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0048] Through the above aspects of this disclosure, a target dubbing script for a video to be dubbed is obtained; wherein, the target dubbing script includes at least one line of dialogue and the corresponding time and character for each line; the language of the target dubbing script is different from the language of the video to be dubbed; target audio data is generated by synthesizing speech from the target dubbing script; wherein, the language of the target audio data is the same as the language of the target dubbing script; the target audio data undergoes timbre conversion; the timbre-converted audio data is then synchronized with the video to be dubbed to generate a dubbed video. In this way, through speech synthesis, timbre conversion, and audio-visual synchronization, automated cross-language dubbing of videos is achieved, eliminating the need for expensive manual dubbing, significantly reducing the cost of video dubbing, and increasing dubbing production capacity; it provides convenience for the export of a large number of film and television works and other content, and significantly reduces the cost of export.

[0049] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0050] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0051] Figure 1 A flowchart illustrating a dubbing method according to an embodiment of the present disclosure is shown;

[0052] Figure 2This diagram illustrates the structure of a speech synthesis model according to an embodiment of the present disclosure.

[0053] Figure 3 A schematic diagram illustrating timbre conversion according to an embodiment of the present disclosure is shown;

[0054] Figure 4 A schematic diagram illustrating the adjustment of timing information of an audio segment according to an embodiment of the present disclosure is shown.

[0055] Figure 5 A schematic diagram illustrating an audio-visual hard synchronization strategy according to an embodiment of the present disclosure is shown;

[0056] Figure 6 A schematic diagram illustrating an audio-visual soft synchronization strategy according to an embodiment of the present disclosure is shown;

[0057] Figure 7 A flowchart of a speech synthesis method according to an embodiment of the present disclosure is shown;

[0058] Figure 8 A schematic diagram illustrating an extraction of prosodic features and phoneme features according to an embodiment of the present disclosure is shown.

[0059] Figure 9 A schematic diagram illustrating the extraction of phoneme features and prosodic features according to an embodiment of the present disclosure is shown.

[0060] Figure 10 A schematic diagram of a preset emotional space according to an embodiment of the present disclosure is shown;

[0061] Figure 11 A schematic diagram illustrating the determination of subtitle text corresponding to audio according to an embodiment of the present disclosure is shown;

[0062] Figure 12 A schematic diagram illustrating speech-text alignment according to an embodiment of the present disclosure is shown;

[0063] Figure 13 This diagram shows a structural diagram of a dubbing device according to an embodiment of the present disclosure;

[0064] Figure 14 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. Detailed Implementation

[0065] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0066] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this disclosure include a particular feature, structure, or characteristic described in connection with that embodiment. Therefore, the terms "exemplary," "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in various parts of this specification, do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including, but not limited to," unless otherwise specifically emphasized.

[0067] In this disclosure, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0068] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0069] Film and television works are a powerful mass communication medium in an international context, making localized dubbing for film and television dramas particularly important. With a massive number of domestic film and television works needing export in recent years, and a typical TV series having dozens of episodes, the workload for manual foreign language dubbing is enormous. Furthermore, the emotional expressiveness and depth of understanding required for film and television works far exceed those of general radio dramas. This demands that voice actors possess a higher level of skill and in-depth understanding of film and television dramas; however, qualified voice actors are relatively scarce, especially for dubbing in less commonly spoken languages, where finding suitable talent is particularly difficult.

[0070] To address the aforementioned issues, this disclosure proposes a dubbing method (detailed description below). By performing text-to-speech (TTS), voice conversion (VC), and audio-visual synchronization, it achieves automated cross-language dubbing of videos, effectively reducing manual workload and eliminating the need for expensive manual dubbing, thus significantly lowering the cost of video dubbing.

[0071] For example, this method can be executed by electronic devices such as terminal devices or servers. The terminal device can be a desktop or mobile terminal, such as a laptop, tablet, desktop computer, smartphone, smart speaker, smartwatch, smart TV, in-vehicle terminal, or other types of electronic devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0072] Figure 1 A flowchart illustrating a dubbing method according to an embodiment of the present disclosure is shown, such as... Figure 1 As shown, it includes the following steps:

[0073] Step 101: Obtain the target dubbing script for the video to be dubbed; wherein, the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line of dialogue; the language of the target dubbing script is different from the language of the video to be dubbed.

[0074] The videos to be dubbed can be of various types, including movies, TV series, animations, music videos, TV programs, sports events, documentaries, social media videos, and advertisements. The language of the video to be dubbed can also be called the source language, which is the language used by the characters during the filming or production of the video, such as Chinese, English, Japanese, Thai, etc.

[0075] A target dubbing script is a text prepared for voice actors to guide them in delivering the lines of a designated character at corresponding moments in the video to be dubbed. The language of the target dubbing script can also be called the target language, i.e., the language the voice actor will use to dub the video. For example, the target dubbing script includes the target language lines for each character, and the time corresponding to any line in the target dubbing script can include the start time of that line. For example, the target dubbing script may also include information such as the age and gender of each character.

[0076] For example, in a scenario where a Chinese TV drama needs to be dubbed before being broadcast in Thailand, the video to be dubbed is a Chinese TV drama in Chinese language, the target dubbing script is in Thai language, and the target dubbing script contains Thai lines for different characters in the Chinese TV drama, as well as the time corresponding to each line of Thai lines.

[0077] For example, the target dubbing script can be generated from the source audio data of the video to be dubbed. The source audio data corresponds to the same language as the video to be dubbed. The source audio data can be audio recordings of actors speaking during the filming of the video, or audio recordings of dubbing actors in a recording studio after the video is dubbed. For example, the source audio data can include audio clips of the source language lines for each character, along with the start time and duration of each audio clip. The source audio data and the visuals in the video to be dubbed are synchronized. Displaying the source language lines from the source audio data on the video screen according to the start time and duration of the corresponding audio clips forms subtitles. In one possible implementation, the source audio data of the video to be dubbed can be obtained; speech recognition can be performed on the source audio data to generate a source dubbing script; and the source dubbing script can be translated to generate the target dubbing script. The source dubbing script corresponds to the same language as the video to be dubbed. Speech recognition and translation can be implemented using existing technologies, and are not limited thereto. As an example, different speech recognition models can be pre-trained for different languages. These models can identify the speaker's identity and the content being spoken in the audio. Then, for the language of the video to be dubbed, a pre-trained language recognition model corresponding to that language is selected. The source audio data of the video to be dubbed is input into this model. The model distinguishes different roles in the source audio data through voiceprint recognition and semantic recognition of the lines spoken by different roles, recording the time corresponding to each line. The identified roles are matched with the identified lines to obtain the source dubbing script. For example, the source dubbing script may also include information such as the age and gender of each role. Then, a machine translation model can be used to translate the source dubbing script, translating the lines of each role from the source language to the target language, generating the target dubbing script. In another possible implementation, the source dubbing script of the video to be dubbed can be obtained and translated to generate the target dubbing script.

[0078] Step 102: Generate target audio data by performing speech synthesis on the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script.

[0079] For example, speech synthesis models corresponding to different languages ​​can be pre-trained. These trained models are used to synthesize speech in the same language based on text. For the language corresponding to the target audio data, a pre-trained speech synthesis model corresponding to that language is selected. The target dubbing script is input into this model, and the target audio data is output. The target audio data consists of the audio of each character speaking in the target language in the video to be dubbed, and may include audio clips of each character's lines in the target language, as well as the start time and duration of each audio clip.

[0080] As an example, a speech synthesis model can be a model built on a conditional variational autoencoder (VITS) with adversarial learning for end-to-end text-to-speech. Figure 2 A schematic diagram of the structure of a speech synthesis model according to an embodiment of the present disclosure is shown, as follows: Figure 2 As shown, a trained speech synthesis model can include: a text encoder, a projection layer, a stochastic duration predictor, a flow model, and a decoder. The text encoder encodes the target dubbing script to obtain the corresponding textual intermediate features. The projection layer maps these intermediate features to a latent space. The stochastic duration predictor presets the duration of each phoneme in the target dubbing script and maps it to the latent space. The flow model generates latent variables, and the decoder decodes these latent variables to generate target audio data corresponding to the target dubbing script. Thus, the trained speech synthesis model can generate high-quality speech that sounds natural and matches the target dubbing script. The training process of the speech synthesis model can be implemented using existing technologies and is not limited thereto.

[0081] Step 103: Perform timbre conversion on the target audio data.

[0082] Timbre, in particular, describes the vocal characteristics of different people speaking, and is typically related to factors such as the vibration pattern of the vocal cords, the shape and size of the oral cavity, and the speaker's gender and age. In this step, by converting the timbre of the target audio data, the timbre differentiation of each character within the target audio data is improved, thereby enhancing the user's experience in watching and understanding the plot.

[0083] In one possible implementation, the timbre conversion of the target audio data includes: determining, from a plurality of preset reference timbres, a target timbre that matches at least one character corresponding to the target audio data; the languages ​​corresponding to the plurality of reference timbres are the same as the languages ​​corresponding to the target audio data; and performing timbre conversion on the target audio data based on the target timbre that matches the at least one character.

[0084] Considering that a TV series often features hundreds of characters, including elderly people, playful young girls, domineering CEOs, and gentle ladies, accurately portraying these character traits during dubbing is extremely challenging. For example, a voice library containing different languages ​​can be pre-constructed based on existing corpora. Each language's voice library includes multiple reference voices, as well as the corresponding voice features and character information (such as age (or age range), gender, etc.) for each reference voice. For instance, for the same language, the voices of actors of different genders and ages can be used as reference voices, and voice features can be extracted from the actors' audio data as the corresponding voice features for the reference voices. This pre-constructs a voice library containing different dimensions of features such as language, gender, age, and voice. Then, voice matching can be performed on different characters involved in the target audio data within the voice library. For example, determining the target timbre that matches at least one character corresponding to the target audio data from a preset set of reference timbres may include: selecting a timbre library corresponding to the target language; extracting the timbre features of each character from the source audio data of the video to be dubbed, and obtaining the age and gender of each character from the target dubbing script; for any character, comparing the extracted timbre features of that character with the timbre features of each reference timbre in the timbre library, and combining the similarity of age and gender, comprehensively selecting the reference timbre with the highest similarity as the target timbre matching that character, for example, targeting... For any reference timbre, the similarity between the character's timbre features and the reference timbre features, the similarity between the character's age and the corresponding age of the reference timbre, and the similarity between the character's gender and the corresponding gender of the reference timbre are calculated. The sum of these three similarities is then calculated, and the reference timbre with the highest sum of the three similarities is selected as the target timbre. In essence, if the character corresponding to the target timbre data is the same as the character corresponding to the source audio data, then the target timbre is the target timbre matched to that character in the target audio data, thus achieving a high similarity match between the character and the target timbre in the target audio data. Furthermore, based on the target timbre matched to that character, timbre conversion can be performed on the audio segments corresponding to that character in the target audio data.

[0085] Considering that the age of the same character may change as the plot develops in the same film or television work, and the timbre of the same character may differ at different ages (for example, the timbre of a teenager may differ from that of an adult), the target timbre for each age can be obtained based on the age involved in each character in the target dubbing script. The audio segments corresponding to the character in the target timbre data can then be divided into audio segments corresponding to different ages of the character. Based on the target timbre matching the character at different ages, timbre conversion can be performed on the corresponding age-related audio segments. For example, when obtaining the target timbre matching any age of the character, the audio data corresponding to that age of the character can be obtained from the source audio data of the video to be dubbed. Timbre features are extracted from this audio data, and then the similarity between the timbre features and the timbre features of each reference timbre, the similarity between the age and the corresponding age of each reference timbre, and the similarity between the character's gender and the corresponding gender of each reference timbre are calculated. The sum of these three similarities is then calculated, and the reference timbre with the highest sum of the three similarities is taken as the target timbre matching that age of the character.

[0086] For example, the timbre conversion of the target audio data based on the target timbre matched with the at least one character may include: obtaining preset timbre features corresponding to the target timbre matched with the at least one character; extracting semantic features corresponding to the at least one character in the target audio data; and generating timbre-converted audio data by fusing the preset timbre features with the semantic features. The semantic features are features in the target audio data that are unrelated to the character's timbre but related to content information, emotional information, prosodic information, etc. In this way, by extracting semantic features, timbre conversion is performed while maintaining the semantic features unchanged, avoiding the impact of changes in the timbre features of the target audio data on the semantic information, emotional information, prosodic information, etc. of the target audio data; thus, intelligent conversion of the character's timbre can be performed without losing emotional expression characteristics and maintaining the same semantics, to accurately shape the character's vocal characteristics in the target audio data.

[0087] Figure 3 A schematic diagram illustrating timbre conversion according to an embodiment of the present disclosure is shown, such as... Figure 3As shown, for any given character, the target audio data contains the corresponding semantics and timbre. The semantics of the character are encoded, and semantic features are extracted. Based on the target timbre matched to the character, preset timbre features corresponding to the target timbre are obtained from a preset timbre library. Then, the preset timbre features and semantic features are decoded and fused to generate timbre-converted audio data. Since the semantic features do not change during the timbre conversion process, it ensures that the semantics of the character in the timbre-converted audio data are the same as those in the target audio data before conversion. At the same time, the timbre features change during the timbre conversion process, ensuring that the timbre of the character in the timbre-converted audio data is the same as the target timbre, thus achieving timbre conversion within the same language semantic space.

[0088] In one possible implementation, a timbre conversion model can be pre-trained. The trained timbre conversion model has the ability to convert timbre within the semantic space of the same language. The source audio data and target audio data of the video to be dubbed are input into the timbre conversion model, and the output is the timbre-converted audio data. The language corresponding to the timbre-converted audio data is still the target language, and the semantics of each character in the target audio data before timbre conversion are the same, but the timbre of each character has changed. As an example, a timbre conversion model can include a timbre extraction network, a timbre removal network, and a vocoder. The timbre extraction network can extract the timbre features of each character in the target audio data from the source audio data of the video to be dubbed. Based on the similarity between the timbre features of each character and the timbre features of each reference timbre in the timbre library, the similarity between the age of each character and the corresponding age of each reference timbre, and the similarity between the gender of each character and the corresponding gender of each reference timbre, a target timbre matching each character is selected from the timbre library. The timbre removal network can obtain the semantic features of each character in the target audio data, achieving accurate acquisition of features in the target audio data that are unrelated to the timbre of each character but are related to content information, emotional information, and prosodic information. The vocoder performs fusion processing based on the semantic features and the preset timbre features corresponding to the target timbre to generate timbre-converted audio data, thereby converting the timbre of each character in the target audio data into the target timbre while preserving the original semantics. As another example, a timbre conversion model can include a timbre extraction network, a timbre removal network, and a vocoder. The timbre extraction network can extract the timbre features of each character at different ages from the source audio data of the video to be dubbed. Then, based on the similarity between the timbre features of each character at different ages and the timbre features of each reference timbre in the timbre library, the similarity between different ages and the corresponding ages of each reference timbre, and the similarity between the gender of each character and the corresponding gender of each reference timbre, the target timbre matching different ages of each character can be selected from the timbre library. The timbre removal network can obtain the semantic features corresponding to different ages of each character in the target audio data. The vocoder can then perform fusion processing based on the semantic features of the same character at the same age and the preset timbre features corresponding to the matched target timbre to generate the audio data after timbre conversion for different ages of the character.

[0089] For example, the audio data after timbre conversion includes audio segments of the target language lines of each character and the start time and duration of each audio segment; it can be understood that the start time and duration of each audio segment in the audio data after timbre conversion are consistent with the start time and duration of the corresponding audio segment in the target audio data.

[0090] Step 104: Perform audio-visual synchronization processing on the audio data after timbre conversion and the video to be dubbed to generate the dubbed video.

[0091] Considering the differences in pronunciation systems between the target and source languages, the timing information of the converted audio data and the source audio data of the video to be dubbed may differ. This timing information can include the duration and start time of each audio segment. For example, for the same line spoken by the same character, the duration of the audio segment corresponding to the target language line in the converted audio data may differ from the duration of the audio segment corresponding to the source language line in the source audio data. Therefore, audio-visual synchronization processing can be used to align the timing information of the converted audio data with that of the source audio data of the video to be dubbed, ensuring that the sound and visuals are consistent and guaranteeing the synchronization of audio and video content. This ensures an immersive experience for viewers watching the dubbed video. For instance, the timing information in the converted audio data can be adjusted to ensure consistency between the sound and the visuals in the video to be dubbed, thus achieving audio-visual synchronization processing.

[0092] For example, the audio data after timbre conversion can be decomposed into multiple audio segments, each corresponding to a line of dialogue in the target language. For each audio segment, the corresponding audio segment of the source language dialogue in the source audio data is determined. The duration of the audio segment is compared to the duration of the audio segment of the source language dialogue. If it exceeds the duration, it is further determined whether the excess duration is greater than a preset threshold. If it is greater than the preset threshold, the audio segment is accelerated, for example, at 1.3x speed or 1.5x speed, to reduce the difference between the duration of the audio segment of the target language dialogue and the duration of the corresponding audio segment of the source language dialogue. The specific value of the preset threshold can be set according to the requirements.

[0093] In one possible implementation, the step of performing audio-visual synchronization processing between the timbre-converted audio data and the video to be dubbed to generate the dubbed video includes: decomposing the timbre-converted audio data into multiple audio segments; wherein each audio segment corresponds to a line of dialogue; determining a processing method for each audio segment; wherein the processing method includes one or more of the following: keeping it unchanged, speeding up, aligning, and tilt-shifting; adjusting the multiple audio segments based on the processing method of each audio segment, and aligning the adjusted audio segments with the scenes in the video to be dubbed in time sequence to generate the dubbed video.

[0094] "Keep it unchanged" means that the timing information of the audio segment is not adjusted, while "speed up," "align," and "shift" all mean that the timing information of the audio segment is adjusted. Figure 4 This diagram illustrates the adjustment of timing information of an audio segment according to an embodiment of the present disclosure. The term "alignment" refers to merging two adjacent audio segments into one, "shift" refers to moving the start and / or end times of the audio segment, and "acceleration" refers to shortening the duration of the audio segment.

[0095] For example, when determining the processing method for each audio segment, multiple audio segments can be considered comprehensively. Based on the duration of the multiple audio segments and the duration of the corresponding audio segments in the source audio data, the processing method for each audio segment among the multiple audio segments can be adaptively determined. During this adaptive determination process, parameters such as the speedup ratio, tilt-shift offset, and number of combined tilt-shifts corresponding to different processing method combinations can be precisely calculated, ultimately determining the optimal processing method. The tilt-shift offset can include the sum of the time differences between the target start time after the tilt-shift and the original start time before the shift, and / or the target end time after the shift relative to the time before the shift. The sum of time differences at the original termination time; as an example, different weight values ​​can be assigned to speedup ratio, tilt-shift offset, and number of axes. Based on the weight values ​​and the speedup ratio, tilt-shift offset, and number of axes corresponding to each processing method combination, the magnitude of the modification value is calculated, and the processing method combination with the smallest modification value is taken as the optimal processing method combination, that is, the optimal processing method for multiple audio segments, thereby achieving precise synchronization of audio and video and ensuring perfect integration of video audio and video; among which, the weight values ​​assigned to speedup ratio, tilt-shift offset, and number of axes can be set according to requirements, and the value obtained by weighted summation of the assigned weight values ​​with respect to speedup ratio, tilt-shift offset, and number of axes can be used as the modification value.

[0096] For example, if any audio segment has been sped up, the processing method for that audio segment is determined to be one of keeping it unchanged, aligning the axes, or shifting the axes, in order to avoid the problem of unclear pronunciation caused by sping up again.

[0097] Based on the adaptive determination of the processing method corresponding to each audio segment in multiple audio segments, the number of roles to which the multiple audio segments belong can be divided into audio-visual hard synchronization strategy and audio-visual soft synchronization strategy.

[0098] As an example, a hard audio-visual synchronization strategy involves adjusting multiple audio segments of a character's consecutive lines of dialogue in the target language. This means treating the character's consecutive lines of dialogue as a whole and adaptively determining the processing method for each audio segment within those segments. The goal is to ensure that the total duration of the processed target language audio segments matches the total duration of the corresponding source language audio segments. For instance, let's take a scenario where the source language is Chinese and the target language is Thai. Figure 5 A schematic diagram of an audio-visual hard synchronization strategy according to an embodiment of the present disclosure is shown, such as... Figure 5As shown, the Chinese subtitle timeline, Chinese subtitle text, Thai subtitle text, and Thai audio queue can be mapped according to time. The Chinese subtitle timeline represents the audio segments of each Chinese line in the Chinese audio data arranged chronologically. The Chinese subtitle text represents the Chinese lines arranged chronologically according to their display time in the video frame. The Thai subtitle text represents the Thai lines arranged chronologically according to their display time in the video frame. The Thai audio queue is the Thai audio data obtained through the above processing and timbre conversion. For any audio segment of a Chinese line, the duration, middle... The duration of the Chinese subtitle text corresponding to the Chinese dialogue, the duration of the Thai subtitle text corresponding to the Thai dialogue, and the duration of the audio segments corresponding to the Thai dialogue in the Thai audio queue are the same, but they may differ from the duration of the audio segments corresponding to the Thai dialogue in the Thai audio queue. For the five audio segments of speaker 1 in the Thai audio queue (i.e., 1, 2, 3, 4, and 5 in the figure), the processing method for each of the five audio segments is adaptively determined. Among them, the duration of audio segment 1 is shorter than the duration of the corresponding Chinese subtitle line "auspicious day", and correspondingly, there will be a blank space in the audio queue. Audio segments 3 and 4 have been accelerated. The processing methods for these five audio segments were determined through adaptive processing as follows: Audio segment 1 was processed to remain unchanged; Audio segment 2 was processed to be accelerated, meaning that character 1 spoke the line "Miss" in Thai faster; Audio segment 3 was processed to be shifted forward and backward in both directions, meaning that character 1 spoke the line "It's time to get married" in Thai slower; Audio segments 4 and 5 were processed by combining the axes of 4 and 5, meaning that character 1 combined the two lines "Miss" and "It's time to get in the sedan chair" into one line in Thai, thus reducing the time gap between the two lines. The audio segments were processed according to their respective processing methods to generate the voice track shown in the figure. This voice track is the audio data after audio-visual synchronization processing. The duration reduced by combining the axes of 4 and 5 is the same as the duration of the backward axis shift in audio segment 3, and the duration reduced by accelerating audio segment 2 is the same as the duration of the forward axis shift in audio segment 3. In the voice track, the total duration of audio clips 1-5 in which character 1 speaks lines in Thai is consistent with the total duration of the corresponding audio clips in the Chinese subtitles, thus achieving time alignment.

[0099] As another example, the audio-visual soft synchronization strategy means adjusting multiple audio segments corresponding to adjacent target language lines of different characters together. That is, treating multiple consecutive lines of target language lines of multiple characters as a whole, and adaptively determining the processing method for each audio segment corresponding to the multiple lines of target language lines of these multiple characters, so that the total duration of the multiple audio segments processed according to the processing method is consistent with the total duration of the multiple audio segments obtained from the corresponding source language lines. Figure 6A schematic diagram of an audio-visual soft synchronization strategy according to an embodiment of the present disclosure is shown, such as... Figure 6 As shown, for the four audio segments of speaker 1 (i.e., 1, 2, 3, and 4 in the figure) and one audio segment of speaker 2 (i.e., 5 in the figure) in the Thai audio queue, the processing method for these five audio segments is adaptively determined. Audio segment 1 has a shorter duration than the corresponding Chinese subtitle line "Huang Dao Ji Ri" (auspicious day), resulting in a gap in the audio queue. Audio segment 5 has been sped up. Audio segments 4 and 5 belong to two different speakers and overlap in time, meaning there is a segment in the Thai audio queue where speaker 1 and speaker 2 speak simultaneously. The adaptive processing method for these five audio segments is determined as follows: the processing method for audio segment 1 is to maintain... The processing method for audio segments 2 and 3 is to combine the two lines, i.e., character 1 combines the two lines "Miss" and "It's time to get married" into one line in Thai, thus reducing the time gap between the two lines; the processing method for audio segment 4 is to speed up, i.e., character 1 speeds up the line "Miss" in Thai; the processing method for audio segment 5 is to shift the axis forward, i.e., character 2 says the line "You should get in the sedan chair" in Thai in advance; the audio segments are processed according to the processing method corresponding to each audio segment to generate the human voice track shown in the figure. This human voice track is the audio data after audio-visual synchronization and processing; among them, the forward duration of audio segment 5, in addition to the duration reduced by the acceleration of audio segment 4 and the duration reduced by the combination of the two lines, also occupies a part of the blank time. In the audio track, after character 1 says "Miss" in Thai, character 2 then says "You should get in the sedan chair" in Thai, thus avoiding the overlap of speaking times between different characters. At the same time, the sum of the duration of audio segments 1-4 spoken by character 1 and the duration of audio segment 5 spoken by character 2 is consistent with the total duration of the audio segments corresponding to the Chinese subtitle timeline, thus achieving time alignment.

[0100] Understandably, one can flexibly choose the appropriate audio-visual synchronization strategy according to the needs. For example, if the relative independence of the speech between different characters is to be maintained, a hard audio-visual synchronization strategy can be chosen. If the original rhythm of the speech content is to be maintained as much as possible, a soft audio-visual synchronization strategy can be adopted.

[0101] In this embodiment, a target dubbing script for a video to be dubbed is obtained. The target dubbing script includes at least one line of dialogue and the corresponding time and character for each line. The language of the target dubbing script is different from the language of the video to be dubbed. Target audio data is generated by synthesizing speech from the target dubbing script. The language of the target audio data is the same as the language of the target dubbing script. The target audio data undergoes timbre conversion. The timbre-converted audio data is then synchronized with the video to be dubbed to generate a dubbed video. Thus, by performing speech synthesis, timbre conversion, and audio-visual synchronization, automated cross-language dubbing of videos is achieved, eliminating the need for expensive manual dubbing, significantly reducing the cost of video dubbing, and increasing dubbing production capacity. This facilitates the export of a large number of film and television works and other content, greatly reducing the cost of export.

[0102] The following provides an exemplary description of possible implementation methods for generating target audio data by performing speech synthesis on the target dubbing script in step 102 above.

[0103] Figure 7 A flowchart of a speech synthesis method according to an embodiment of the present disclosure is shown, as follows: Figure 7 As shown, it includes the following steps:

[0104] Step 701: Extract the prosody embedding and phoneme embedding of the target dubbing script.

[0105] Among them, a phoneme is the smallest unit of speech that is divided according to the natural attributes of speech. Phoneme features are used to represent pronunciation information. Prosody refers to the sound patterns in a language, including stress, intonation, pauses, etc. Prosodic features are used to characterize the intonation, rhythm, and intensity of speech.

[0106] For example, the aforementioned speech synthesis model can be used to extract prosodic and phoneme features from the target dubbing script. For instance, a pre-trained prosodic feature extraction model and a pre-trained phoneme feature extraction model can be configured within the speech synthesis model. As an example, the phoneme feature extraction model can be a grapheme-to-phoneme (G2P) model, and the prosodic feature extraction model can be a BERT model using bidirectional encoder representations from Transformers. BERT is a pre-trained model based on the Transformer architecture, exhibiting extremely high performance through pre-training on large-scale unlabeled text and then fine-tuning on specific tasks. For example, taking a VITS-based speech synthesis model as an example, a pre-trained VITS-based model can be configured within the speech synthesis model. Figure 2 The VITS text encoder shown is configured with both the BERT and G2P models. Figure 8 This diagram illustrates an embodiment of the present disclosure for extracting prosodic features and phonemic features, as shown below. Figure 8 As shown, the BERT model is used to extract prosodic features, and the G2P model is used to extract phoneme features.

[0107] In one possible implementation, extracting the phoneme features of the target dubbing script includes: generating a phoneme sequence corresponding to the target dubbing script based on a preset phoneme dictionary; and extracting the phoneme features corresponding to each phoneme in the phoneme sequence. The phoneme dictionary contains the correspondence between keywords and phonemes in the target dubbing script. Thus, the target dubbing script is divided into phonemes based on the preset phoneme dictionary, thereby generating a phoneme sequence. It is understood that for the same target dubbing script, the phoneme sequences obtained by using different phoneme dictionaries may be different. For example, phoneme dictionaries for different languages ​​can be pre-constructed according to the characteristics of different languages. The preset phoneme dictionary can include more granular phonemes, thereby making the pronunciation of characters in the generated target audio data more natural. For example, for the same line, more phonemes can be divided according to the preset phoneme dictionary.

[0108] Figure 9 This diagram illustrates an embodiment of the present disclosure for extracting phoneme features and prosodic features, as shown below. Figure 9 As shown, the G2P model is equipped with a word segmenter. Based on a pre-defined Thai phoneme dictionary, the word segmenter divides each line in the target dubbing script into a phoneme sequence. The concatenation of the phoneme sequences corresponding to each line forms the phoneme sequence corresponding to the target dubbing script. The G2P model is pre-trained to extract the phoneme features of each phoneme; for example... Figure 9 As shown, regarding the Thai dialogue " The related technologies are divided into terms " The phoneme of “sawatdi” is “sawatdi”. In this embodiment of the disclosure, the G2P model is configured with two word segmenters, which further segment the dialogue based on different phoneme dictionaries to obtain more fine-grained phonemes. Among them, word segmenter 1, based on the corresponding phoneme dictionary, segments “sawatdi” into “sawatdi”. "Word segmentation into " "", "", Three keywords, among which, The corresponding phoneme is "sawa". The corresponding phoneme is " "", The corresponding phoneme is "di". Using word segmenter 2, based on another phoneme dictionary, "di" is segmented... "Word segmentation into " "and" Two keywords; among them, " The corresponding phoneme is "sawat". The corresponding phoneme is "di".

[0109] In one possible implementation, the method further includes: extracting the emotional features of the target dubbing script.

[0110] Among them, emotional characteristics are used to represent the type of emotion (such as happy, sad, excited, depressed, etc.) and the degree of attitude (such as affirmation, denial, praise, criticism, irony, etc.).

[0111] For example, a pre-trained emotion recognition model can be configured within the speech synthesis model to extract emotional features from the target dubbing script. Taking a VITS-based speech synthesis model as an example, the emotion recognition model can be configured within the text encoder of the VITS model. As an example, the emotion recognition model can be a Hubert-based speech emotion recognition model (HuBert with speech emotion recognization, Hubert-SER), where Hubert is a BERT-based self-supervised speech representation learning model.

[0112] As an example, the extraction of emotional features from the target dubbing script includes: extracting emotional features from the target dubbing script using an emotion recognition model; wherein the emotion recognition model is trained using a transfer learning algorithm based on a preset emotion space, and the preset emotion space includes at least: joy, sadness, surprise, anger, disappointment, pride, jealousy, rage, bittersweetness, and happiness.

[0113] Figure 10This diagram illustrates a preset emotional space according to an embodiment of the present disclosure, such as... Figure 10 As shown, the left side represents the existing emotional space, and the right side represents the preset emotional space used in this embodiment. In the preset emotional space used in this embodiment, yellow represents "Joy", blue represents "Sadness", light blue represents "Surprise", red represents "Anger", dark blue represents "Disappointment", orange represents "Pride", purple represents "Envy", dark purple represents "Outrage", olive green represents "Bittersweetness", and green represents "Delight". Compared with the existing emotional space, it has been further refined, thereby making it easier to extract the subtle emotional features of the target dubbing script.

[0114] Transfer learning is a deep learning technique that allows a model trained on one task to be fine-tuned on another related task. This method can improve the model's performance on new tasks while reducing training time and data requirements. For example, samples of dubbing scripts in the source language can be used beforehand. Figure 10 Within the preset emotion space shown, the emotion recognition model is trained to accurately capture subtle emotions in the source language, and then uses dubbing script samples in the target language. Figure 10 The emotion recognition model is then trained further within the preset emotion space to ultimately obtain a trained emotion recognition model. This trained model possesses cross-lingual emotion recognition and transfer capabilities, enabling emotion mapping across different languages ​​and thus enhancing the emotional richness and impact of the synthesized target audio data.

[0115] Step 702: The prosodic features and phoneme features of the target dubbing script are fused to generate text features.

[0116] For example, as described above Figure 8 As shown, the prosodic features output by the BERT model can be fused with the phoneme features output by the G2P model to generate text features.

[0117] In one possible implementation, the step of fusing the prosodic features and phoneme features of the target dubbing script to generate text features includes: processing the prosodic features of the target dubbing script to determine the prosodic features corresponding to each phoneme in the phoneme sequence; and fusing the prosodic features corresponding to each phoneme in the phoneme sequence with the phoneme features corresponding to each phoneme to generate the text features.

[0118] Because the phonemes are further subdivided using a pre-defined phoneme dictionary, the trained phoneme feature extraction model configured in the speech synthesis model has the ability to extract the phoneme features of each subdivided phoneme, while the trained prosodic feature extraction model has the ability to extract the prosodic features of the target dubbing script. The prosodic features of the target dubbing script contain the phoneme features of all the phonemes corresponding to the target dubbing script. Therefore, the prosodic features of the target dubbing script can be processed to determine the prosodic features corresponding to each phoneme in the phoneme sequence, thereby aligning the feature dimensions of the prosodic features and the phoneme features, so as to facilitate the fusion of the prosodic features corresponding to each phoneme in the phoneme sequence with the phoneme features corresponding to each phoneme.

[0119] For example, such as Figure 9 As shown, the dimensional alignment of prosodic features and phoneme features can be achieved through three alignment methods. Alignment method 1: Average and broadcast. This method averages the prosodic features across all dimensions and propagates the average value to every phoneme in the entire phoneme sequence, resulting in identical prosodic features for each phoneme. Alignment method 2: Greedy repeat. This method distributes the prosodic features evenly to each phoneme according to the order of dimensions and phonemes in the phoneme sequence. For the remaining prosodic features, the remaining dimensions are still assigned to the phonemes with the highest phoneme order. For example... Figure 9 As shown, the prosodic feature has a dimension of 7, and the number of phonemes is 3. Therefore, the prosodic feature dimension corresponding to each phoneme is 2, and the remaining one dimension of the prosodic feature is assigned to the first phoneme. Alignment method 3: Accurate alignment. This method assigns the prosodic feature to each phoneme according to the length of each phoneme. As shown in the figure, the length of the phoneme "sawat" is 5, the length of "di" is 2, and the prosodic feature dimension is 7. Therefore, the first 5 dimensions of the prosodic feature can be assigned to the phoneme "sawat", and the last 2 dimensions of the prosodic feature can be assigned to the phoneme "di".

[0120] Step 703: Perform speech synthesis based on the text features to generate the target audio data.

[0121] For example, it can be done through the above Figure 2 The speech synthesis model shown generates target audio data by synthesizing speech based on text features generated by a text encoder.

[0122] For example, speech synthesis is performed based on text features fused with prosodic and phonemic features. The generated target audio data contains prosodic and phonemic information, which effectively improves the pronunciation of each character in the target audio data.

[0123] In one possible implementation, the step of generating the target audio data by performing speech synthesis based on the text features includes: performing speech synthesis based on the emotional features of the target dubbing script and the text features to generate the target audio data. In this way, the generated target audio data contains prosodic information, phonemic information, and emotional information, realizing cross-language joint emotion transfer speech synthesis (Emotinal Text-to-Speech, Emotional-TTS), greatly enhancing the realism and prosodic expression of each character's target language. For example, taking a VITS-based speech synthesis model as an example, it is possible to... Figure 2 The VITS text encoder shown is configured with a Hubert-SER model, a BERT model, and a G2P model. The BERT model is used to extract prosodic features, the G2P model is used to extract phoneme features, and the Hubert-SER model is used to extract emotional features. The text encoder fuses the prosodic features and phoneme features to generate text features, and then fuses the text features and emotional features to generate intermediate text features. These intermediate text features are then passed to the mapping layer, which maps them to the latent space. The streaming model generates latent variables in the latent space, and the decoder decodes the latent variables to generate the target audio data.

[0124] In this embodiment, prosodic features and phonemic features of the target dubbing script are extracted; the prosodic features and phonemic features of the target dubbing script are fused to generate text features; speech synthesis is performed based on the text features to generate the target audio data; thus, by extracting prosodic features to improve pronunciation, film-level dubbing synthesis can be achieved; as an example, emotional features of the target dubbing script are extracted, and speech synthesis is performed based on the emotional features of the target dubbing script and the text features to generate the target audio data, thereby enabling speech synthesis with localized prosodic features and cross-language emotional transfer, and the generated target audio data is full of rich emotions.

[0125] Furthermore, the obtained target audio data is further processed by timbre conversion and audio-visual synchronization, thereby realizing the dubbing process of speech synthesis, emotion transfer, timbre matching, and audio-visual synchronization, and automatically producing dubbing audio with multiple timbres and emotions.

[0126] For example, the audio channels of the dubbed video can be further optimized as needed. For instance, audio separation and restoration techniques can be used to perform detailed analysis and processing of the audio data in the dubbed video to restore or enhance specific emotional sounds, including but not limited to crying and laughing, to achieve greater realism.

[0127] Furthermore, this embodiment fully utilizes the platform's extensive dubbing resources and provides a method for generating training data. Through audio preprocessing, speaker separation, and text annotation, high-quality training data is rapidly acquired. This not only enhances the quality and applicability of the training data but also expands the diversity of the corpus, covering multiple language domains such as Thai, Vietnamese, Spanish, and English. Therefore, the models involved in the aforementioned dubbing methods can be pre-trained using this training data.

[0128] For example, audio-text alignment (phonetic-text alignment) can be performed when generating training data. Figure 11 This diagram illustrates the determination of subtitle text corresponding to audio according to an embodiment of the present disclosure, such as... Figure 11 As shown, for any audio segment, which can be dialogue audio in film and television works, audio tracks in videos, etc., automatic speech recognition is performed on the audio to generate speech recognition text (ASR Text). Then, the similarity calculation is performed between the speech recognition text and the subtitle text (ASR Text). The subtitle text can be pre-made based on all film and television works containing the audio. In this way, the subtitle text corresponding to the audio segment is found.

[0129] For example, speech-to-text alignment can be performed when generating training data. Figure 12 This diagram illustrates a speech-to-text alignment according to an embodiment of the present disclosure, such as... Figure 12 As shown, for any audio segment, speech-text alignment processing is performed based on the audio and its corresponding subtitle text. For example, alignment processing can be performed using the Montreal Forced Aligner (MFA) to generate speech-text alignment data.

[0130] The "fusion" between features involved in the embodiments of this disclosure can be achieved by selecting specific fusion methods in related technologies as needed, such as splicing, summation, multiplication, etc., and this disclosure does not limit this.

[0131] Based on the same inventive concept as the above method embodiments, the present disclosure also provides a dubbing device, which can be used to execute the technical solutions described in the above method embodiments.

[0132] Figure 13 This diagram illustrates a structural diagram of a dubbing device according to an embodiment of the present disclosure, such as... Figure 13As shown, the device includes: an acquisition module 1301, used to acquire a target dubbing script for a video to be dubbed; wherein the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line; the language of the target dubbing script is different from the language of the video to be dubbed; a speech synthesis module 1302, used to generate target audio data by performing speech synthesis on the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script; a timbre conversion module 1303, used to perform timbre conversion on the target audio data; and an audio-visual synchronization module 1304, used to perform audio-visual synchronization processing between the timbre-converted audio data and the video to be dubbed to generate a dubbed video.

[0133] In this embodiment, a target dubbing script for a video to be dubbed is obtained. The target dubbing script includes at least one line of dialogue and the corresponding time and character for each line. The language of the target dubbing script is different from the language of the video to be dubbed. Target audio data is generated by synthesizing speech from the target dubbing script. The language of the target audio data is the same as the language of the target dubbing script. The target audio data undergoes timbre conversion. The timbre-converted audio data is then synchronized with the video to be dubbed to generate a dubbed video. Thus, by performing speech synthesis, timbre conversion, and audio-visual synchronization, automated cross-language dubbing of videos is achieved, eliminating the need for expensive manual dubbing, significantly reducing the cost of video dubbing, and increasing dubbing production capacity. This facilitates the export of a large number of film and television works and other content, greatly reducing the cost of export.

[0134] In one possible implementation, the speech synthesis module 1302 is further configured to: extract prosodic features and phoneme features of the target dubbing script; fuse the prosodic features and phoneme features of the target dubbing script to generate text features; and perform speech synthesis based on the text features to generate the target audio data.

[0135] In one possible implementation, the speech synthesis module 1302 is further configured to: extract the emotional features of the target dubbing script; and perform speech synthesis based on the emotional features of the target dubbing script and the text features to generate the target audio data.

[0136] In one possible implementation, the timbre conversion module 1303 is further configured to: determine, among a plurality of preset reference timbres, a target timbre that matches at least one character corresponding to the target audio data; the language corresponding to the plurality of reference timbres is the same as the language corresponding to the target audio data; and perform timbre conversion on the target audio data based on the target timbre that matches at least one character.

[0137] In one possible implementation, the speech synthesis module 1302 is further configured to: generate a phoneme sequence corresponding to the target dubbing script based on a preset phoneme dictionary; extract phoneme features corresponding to each phoneme in the phoneme sequence; process the prosodic features of the target dubbing script to determine the prosodic features corresponding to each phoneme in the phoneme sequence; and fuse the prosodic features corresponding to each phoneme in the phoneme sequence with the phoneme features corresponding to each phoneme to generate the text features.

[0138] In one possible implementation, the speech synthesis module 1302 is further configured to: extract the emotional features of the target dubbing script through an emotion recognition model; wherein the emotion recognition model is trained using a transfer learning algorithm based on a preset emotion space, and the preset emotion space includes at least: joy, sadness, surprise, anger, disappointment, pride, jealousy, rage, bittersweetness, and happiness.

[0139] In one possible implementation, the timbre conversion module 1303 is further configured to: obtain preset timbre features corresponding to the target timbre matched by the at least one character; extract semantic features corresponding to the at least one character in the target audio data; and generate timbre-converted audio data by fusing the preset timbre features with the semantic features.

[0140] In one possible implementation, the audio-visual synchronization module 1304 is further configured to: decompose the audio data after timbre conversion into multiple audio segments; wherein each audio segment corresponds to a line of dialogue; determine the processing method corresponding to each audio segment; wherein the processing method includes one or more of the following: remain unchanged, speed up, axis merging, and axis shifting, wherein axis merging means merging two adjacent audio segments into one audio segment, axis shifting means moving the start time and / or end time of the audio segment, and speeding up means shortening the duration of the audio segment; adjust the multiple audio segments based on the processing method of each audio segment, and align the adjusted audio segments with the scenes in the video to be dubbed in chronological order to generate the dubbed video.

[0141] In one possible implementation, the acquisition module 1301 is further configured to: acquire source audio data of the video to be dubbed; the language corresponding to the source audio data is the same as the language corresponding to the video to be dubbed; perform speech recognition on the source audio data to generate a source dubbing script; and translate the source dubbing script to generate the target dubbing script.

[0142] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0143] This disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the above-described method. The computer-readable storage medium can be volatile or non-volatile.

[0144] This disclosure also proposes an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0145] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0146] Figure 14 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 14 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0147] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). Electronic device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Or similar.

[0148] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0149] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.

[0150] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0151] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0152] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0153] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0154] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0155] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0157] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A dubbing method, characterized in that, The method includes: Obtain the target dubbing script for the video to be dubbed; wherein, the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line of dialogue; the language of the target dubbing script is different from the language of the video to be dubbed; Target audio data is generated by synthesizing the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script. The target audio data is subjected to timbre conversion; The audio data after timbre conversion is synchronized with the video to be dubbed to generate a dubbed video. The step of performing audio-visual synchronization processing between the timbre-converted audio data and the video to be dubbed, to generate the dubbed video, includes: The audio data after timbre conversion is decomposed into multiple audio segments; each audio segment corresponds to a line of dialogue. Determine the processing method for each audio segment; wherein the processing method includes one or more of the following: remain unchanged, speed up, axis merging, and axis shifting, where axis merging means merging two adjacent audio segments into one audio segment, axis shifting means moving the start and / or end time of the audio segment, and speeding up means shortening the duration of the audio segment. The multiple audio segments are adjusted based on the processing method of each audio segment, and the adjusted audio segments are aligned with the scenes in the video to be dubbed in chronological order to generate the dubbed video.

2. The method according to claim 1, characterized in that, The step of generating target audio data by synthesizing speech from the target dubbing script includes: Extract the prosodic and phonemic features of the target dubbing script; The prosodic features and phonemic features of the target dubbing script are fused to generate text features; Speech synthesis is performed based on the text features to generate the target audio data.

3. The method according to claim 2, characterized in that, The method further includes: Extract the emotional features of the target dubbing script; The step of generating the target audio data by performing speech synthesis based on the text features includes: Based on the emotional features and text features of the target dubbing script, speech synthesis is performed to generate the target audio data.

4. The method according to claim 1, characterized in that, The timbre conversion of the target audio data includes: Among a plurality of preset reference timbres, a target timbre that matches at least one character corresponding to the target audio data is determined; the languages ​​corresponding to the plurality of reference timbres are the same as the languages ​​corresponding to the target audio data. Based on the target timbre matched to the at least one character, the target audio data is timbre converted.

5. The method according to claim 2, characterized in that, The extraction of phoneme features from the target dubbing script includes: Based on a preset phoneme dictionary, a phoneme sequence corresponding to the target dubbing script is generated; Extract the phoneme features corresponding to each phoneme in the phoneme sequence; The process of fusing the prosodic and phonemic features of the target dubbing script to generate text features includes: The prosodic features of the target dubbing script are processed to determine the prosodic features corresponding to each phoneme in the phoneme sequence; The prosodic features corresponding to each phoneme in the phoneme sequence are fused with the phoneme features corresponding to each phoneme to generate the text features.

6. The method according to claim 3, characterized in that, The extraction of emotional features from the target dubbing script includes: The emotional features of the target dubbing script are extracted using an emotion recognition model. The emotion recognition model is trained using a transfer learning algorithm based on a preset emotion space, which includes at least the following emotions: joy, sadness, surprise, anger, disappointment, pride, jealousy, rage, bittersweetness, and happiness.

7. The method according to claim 4, characterized in that, The step of converting the target audio data based on the target timbre matched with the at least one role includes: Obtain the preset timbre features corresponding to the target timbre matched by the at least one character; Extract semantic features corresponding to at least one character from the target audio data; By fusing the preset timbre features with the semantic features, audio data after timbre conversion is generated.

8. The method according to claim 1, characterized in that, The method further includes: Obtain the source audio data of the video to be dubbed; the language of the source audio data is the same as the language of the video to be dubbed. The source audio data is subjected to speech recognition to generate a source dubbing script; Translate the source dubbing script to generate the target dubbing script.

9. A dubbing device, characterized in that, The device includes: The acquisition module is used to acquire the target dubbing script of the video to be dubbed; wherein, the target dubbing script includes at least one line of dialogue and the time and character corresponding to each line of dialogue; the language of the target dubbing script is different from the language of the video to be dubbed; The speech synthesis module is used to generate target audio data by performing speech synthesis on the target dubbing script; wherein the language of the target audio data is the same as the language of the target dubbing script. A timbre conversion module is used to convert the timbre of the target audio data; The audio-visual synchronization module is used to perform audio-visual synchronization processing between the audio data after timbre conversion and the video to be dubbed, and generate the dubbed video; The audio-visual synchronization module is specifically used for: decomposing the audio data after timbre conversion into multiple audio segments; wherein each audio segment corresponds to a line of dialogue; determining the processing method corresponding to each audio segment; wherein the processing method includes one or more of the following: keeping it unchanged, speeding up, merging, and shifting, where merging means merging two adjacent audio segments into one audio segment, shifting means moving the start time and / or end time of the audio segment, and speeding up means shortening the duration of the audio segment; adjusting the multiple audio segments based on the processing method of each audio segment, and aligning the adjusted audio segments with the scenes in the video to be dubbed in chronological order to generate the dubbed video.

10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 8 when executing instructions stored in the memory.

11. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video dubbing method and device, computer device and computer readable storage medium

    CN110933330A

  • Speech synthesis method and device, medium and electronic equipment

    CN114242035A

  • Audio processing method and device, electronic equipment and storage medium

    CN114842858A

  • Speech synthesis method, emotion migration method, interaction method, storage medium and program product

    CN114882868A

  • Timbre selection method and device, electronic equipment, readable storage medium and program product

    CN116110366A