Speech synthesis method and device, electronic equipment, storage medium and program product

By extracting the audio emotional characteristics of the source language dubbing audio in the source video and fusing it with the text encoding characteristics of the target language subtitle text, high-quality target language audio is generated, and the problems of high cost and low efficiency in the overseasization of film and television dramas in the existing technology are solved, and automatic and efficient cross-language speech synthesis is achieved.

CN120199228AActive Publication Date: 2025-06-24YOUKU CULTURE TECH (BEIJING) CO LTD

Patent Information

Application Number
CN202510358710.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing technology relies on professional voice actors in the process of overseas film and television dramas, resulting in high cost, low efficiency, and difficulty in achieving high-quality cross-lingual voice synthesis.

Method used

A speech synthesis method is proposed. By obtaining the source language subtitle text and the target language subtitle text of the source video, the emotional extractor is used to extract the audio emotional characteristics of the source language dubbing audio, and fuse it with the text encoding characteristics of the target language subtitle text to generate high-quality target language audio with the source language dubbing audio emotions.

Benefits of technology

It realizes automatic and efficient generation of high-quality target language audio, reducing costs and improving efficiency, eliminating professional dubbing equipment and conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199228A_ABST
    Figure CN120199228A_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: acquiring a source language dubbing audio corresponding to a source language subtitle text of a source video and a target language subtitle text translated by the source language subtitle text; extracting audio emotion features from the source language dubbing audio by using an emotion extractor, wherein the audio emotion features represent emotions expressed by the source language dubbing audio; converting the target language subtitle text into a phoneme sequence, and encoding the phoneme sequence by using a text encoder to obtain text encoding features; fusing the audio emotion features with the text coding features to obtain emotion text features; and generating a target language audio based on the emotion text features by using a decoder, the target language audio being used as a dubbing audio of the source video in the target language. Therefore, the high-quality target language audio with the source dubbing audio emotion can be automatically and efficiently generated, the cost is low, and the efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech synthesis, and in particular, to a speech synthesis method and apparatus, an electronic device, a storage medium, and a program product. Background Art

[0002] In recent years, with the continuous improvement of the quality of domestic film and television dramas, many popular domestic film and television dramas have also been broadcast on foreign TV screens and overseas platforms. In order to better promote the overseas distribution of film and television dramas, the subtitles of the film and television dramas can be translated into the languages of the target markets and dubbed. However, at present, the overseas distribution of film and television dramas mainly relies on professional voice actors to complete the dubbing episode by episode. However, this method requires both professional dubbing capabilities and dubbing conditions, and also consumes a large amount of manpower and financial resources. Summary of the Invention

[0003] In view of this, the present disclosure provides a speech synthesis method and apparatus, an electronic device, a storage medium, and a program product, which can automatically and efficiently generate high-quality target language audio with the emotions of the source dubbing audio, with low cost and high efficiency.

[0004] According to an aspect of the present disclosure, there is provided a speech synthesis method, including: obtaining source language dubbing audio corresponding to source language subtitle text of a source video, and target language subtitle text translated from the source language subtitle text, where the source language is different from the target language; using an emotion extractor to extract audio emotion features from the source language dubbing audio, where the audio emotion features represent the emotions expressed by the source language dubbing audio; converting the target language subtitle text into a phoneme sequence, and encoding the phoneme sequence using a text encoder to obtain text encoding features; fusing the audio emotion features with the text encoding features to obtain emotion text features; and using a decoder to generate target language audio based on the emotion text features, where the target language audio is used as the dubbing audio of the source video in the target language.

[0005] In a possible implementation, the audio emotion features include emotion features of each frame of audio in the L frames of audio of the source language dubbing audio, where L represents the number of frames of the source language dubbing audio; where the fusing the audio emotion features with the text encoding features to obtain emotion text features includes: calculating an average value of the emotion features of the L frames of audio in the audio emotion features to obtain a first emotion feature; performing high-dimensional mapping on the first emotion feature to obtain a second emotion feature, where the dimension of the second emotion feature is the same as the dimension of the text encoding features; and adding the second emotion feature to the text encoding features to obtain emotion text features.

[0006] In a possible implementation, the method further includes: extracting semantic features from the phoneme sequence by using a semantic extractor, where the semantic features characterize the semantics of the target language subtitle text; the using a decoder to generate a target language audio based on the emotional text features includes: fusing the emotional text features with the semantic features to obtain comprehensive features; using the decoder to generate a target language audio based on the comprehensive features.

[0007] In a possible implementation, after generating the target language audio, the method further includes: converting the target language audio into a recognized text by using the speech recognition technology corresponding to the target language; determining a character error rate of the recognized text based on the difference between the recognized text and the target language subtitle text, where the character error rate characterizes the pronunciation error rate of the target language audio; in a case where the character error rate is higher than a preset threshold, screening out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library, so as to regenerate the target language audio by using the audio emotion features of the target reference audio; where the source language audio library includes reference audios of the source language with various emotions.

[0008] In a possible implementation, the source language audio library includes a first audio library, where the reference audios in the first audio library are labeled with emotion categories, and where screening out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library includes: using an emotion recognition model to perform emotion recognition on the source language dubbed audio to obtain an emotion classification result of the source language dubbed audio, where the emotion classification result characterizes the emotion category expressed by the source language dubbed audio; based on the emotion categories labeled for each reference audio in the first audio library, using the reference audio whose emotion category matches the emotion classification result as the target reference audio.

[0009] In a possible implementation, the source language audio library includes a second audio library, where the reference audios in the second audio library are labeled with the emotion audio features of the reference audios extracted by using the emotion extractor, and where screening out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library includes: respectively calculating the similarity between the audio emotion features of the source language dubbed audio and the emotion audio features labeled for each reference audio in the second audio library, and using the reference audio with the highest similarity as the target reference audio.

[0010] In a possible implementation, the training process of the emotion extractor, text encoder, and decoder includes: in the first training stage, training the initial text encoder and the initial decoder using a first dataset to obtain the trained first text encoder and the first decoder, where the first dataset includes: multiple first target language text samples and the corresponding first target language audio samples for each first target language text sample; in the second training stage, training the first text encoder, the first decoder, and the initial emotion extractor using a second dataset to obtain the trained second text encoder, the second decoder, and the first emotion extractor; where the second dataset includes: multiple source language audio samples, the corresponding second target language text samples and the corresponding second target language audio samples for each source language audio sample; in the third training stage, performing fine-tuning training on the second text encoder, the second decoder, and the first emotion extractor using a third dataset to obtain the trained text encoder, decoder, and emotion extractor, where the third dataset includes: multiple third target language texts and the corresponding third target language audio samples for each third target language text, and the third target language audio samples are obtained by collecting the audio of reading the third target language text in a standard pronunciation.

[0011] In a possible implementation, the training of the initial text encoder and the initial decoder using the first dataset to obtain the trained first text encoder and the first decoder includes: for any first target language text sample, encoding the phoneme sequence of the first target language text sample using the initial text encoder to obtain the text encoding feature of the first target language text sample; using the initial decoder to generate the predicted target language audio corresponding to the first target language text sample based on the text encoding feature of the first target language text sample; adjusting the parameters of the initial text encoder and the initial decoder based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample to obtain the trained first text encoder and the first decoder.

[0012] In a possible implementation, the fine-tuning training of the second text encoder, the second decoder, and the first emotion extractor using the third data set to obtain a trained text encoder includes: for any third target language text sample, encoding the phoneme sequence of the third target language text sample using the second text encoder to obtain the text encoding feature of the third target language text sample; extracting the audio emotion feature of the third target language audio sample from the third target language audio sample corresponding to the third target language text sample using the first emotion extractor; fusing the text encoding feature of the third target language text sample with the audio emotion feature of the corresponding third target language audio sample to obtain the emotion text feature corresponding to the third target language audio sample; generating a predicted target language audio corresponding to the third target language text sample using the second decoder based on the emotion text feature corresponding to the third target language audio sample; and adjusting the parameters of the second text encoder, the second decoder, and the first emotion extractor based on the difference between the predicted target language audio corresponding to the third target language text sample and the third target language audio sample to obtain a trained text encoder, emotion extractor, and decoder.

[0013] In a possible implementation, the construction process of the second data set includes: obtaining a bilingual subtitle file corresponding to the original long video data, source language dubbed audio data, and target language dubbed audio data, where the bilingual subtitle file includes the start time and end time of each bilingual subtitle text display, and the bilingual subtitle text includes a target language subtitle text and a source language subtitle text; based on the start time and end time of each bilingual subtitle text display in the bilingual subtitle file, segmenting the source language dubbed audio data and the target language dubbed audio data to obtain multiple first audio groups after segmentation, each first audio group including a segmented source language audio segment and a target language audio segment; using the speech recognition technology corresponding to the target language to convert the target language audio segment in each first audio group into a recognized text, and determining the character error rate of the recognized text converted from each target language audio segment based on the difference between the recognized text converted from each target language audio segment and the target language subtitle text corresponding to each target language audio segment; screening out the first audio groups where the target language audio segments with a character error rate higher than a preset threshold are located in the multiple first audio groups to obtain multiple second audio groups; separating the human voices from the source language audio segments and the target language audio segments in each second audio group to obtain multiple third audio groups, where the source language audio segments and the target language audio segments in each third audio group include human voice tracks; detecting the audio quality of the source language audio segments and the target language audio segments in each third audio group, and screening out the third audio groups with an audio quality lower than a preset quality threshold in the multiple third audio groups to obtain multiple fourth audio groups, where the audio quality includes naturalness and / or noise value; obtaining the second data set based on the source language audio segments and the target language audio segments in the multiple fourth audio groups, and the target language subtitle text corresponding to the target language audio segments in each fourth audio group.

[0014] According to another aspect of the present disclosure, there is provided a speech synthesis device, including: an acquisition module, configured to acquire source language dubbed audio corresponding to the source language subtitle text of the source video, and the target language subtitle text translated from the source language subtitle text, where the source language is different from the target language; a feature extraction module, configured to extract audio emotion features from the source language dubbed audio by using an emotion extractor, and the audio emotion features represent the emotion expressed by the source language dubbed audio; a text encoding module, configured to convert the target language subtitle text into a phoneme sequence, and encode the phoneme sequence by using a text encoder to obtain text encoding features; a feature fusion module, configured to fuse the audio emotion features and the text encoding features to obtain emotion text features; an audio generation module, configured to generate target language audio by using a decoder based on the emotion text features, and the target language audio is used as the dubbed audio of the source video in the target language.

[0015] According to another aspect of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0016] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0017] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0018] According to various aspects of the present disclosure, by extracting the audio emotion features of the source language dubbed audio and fusing them with the text encoding features of the target language subtitle text, emotional text features with the emotional information of the source language dubbed audio and the text information of the target language subtitle text are obtained. In this way, based on the emotional text features, a target language audio with the emotional expression of the source language dubbed audio can be generated, that is, it is possible to automatically and efficiently generate a high-quality target language audio with the emotion of the source language dubbed audio, without the need for professional dubbing equipment and dubbing conditions, nor the consumption of a large amount of manpower and financial resources, making the generation cost of the dubbed audio from the source language to the target language lower and the efficiency higher.

[0019] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings included in the specification and constituting a part of the specification illustrate the exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.

[0021] Figure 1 A flowchart showing a speech synthesis method according to an embodiment of the present disclosure.

[0022] Figure 2 A flowchart showing another speech synthesis method according to an embodiment of the present disclosure.

[0023] Figure 3 A schematic diagram showing a speech synthesis process of Chinese-Thai according to an embodiment of the present disclosure.

[0024] Figure 4 A flowchart showing another speech synthesis method according to an embodiment of the present disclosure.

[0025] Figure 5The flowchart of a three-stage training method according to an embodiment of the present disclosure is shown.

[0026] Figure 6 The block diagram of a speech synthesis device according to an embodiment of the present disclosure is shown.

[0027] Figure 7 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0028] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0029] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.

[0030] When an element is referred to as being "connected", "coupled", "responsive", or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.

[0031] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, the first element / operation in some embodiments may be referred to as the second element / operation in other embodiments.

[0032] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" should not necessarily be construed as superior to or better than other embodiments.

[0033] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0034] As described above, the current overseas adaptation of film and television dramas mainly relies on professional voice actors to complete the dubbing episode by episode. However, this method requires both professional dubbing capabilities and dubbing conditions, and also consumes a large amount of manpower and financial resources. With the development of artificial intelligence (AI) technology, the use of AI technology can achieve seamless cross-lingual voice conversion of actors in film and television works, enabling audiences of different languages to hear the "original" dialogue, that is, to conform to the tone, intonation, and emotions of the character dialogue in the original dubbing. This article mainly focuses on the implementation and realization of a voice synthesis project at the level of film and television dramas, so that the synthesized voice can reach a level that satisfies local people in terms of tone, intonation, emotions, etc. Voice synthesis (Text-to-Speech, TTS), as one of the most important technologies in artificial intelligence, has been successfully applied to various fields, including human-computer interaction, intelligent customer service, virtual assistants, etc. However, at present, the voice synthesis technology is rarely applied to the AI dubbing of film and television dramas. The main reason is that the current voice synthesis technology cannot achieve the voice synthesis of delicate emotions, let alone achieve the "original" emotional expression of film and television dramas.

[0035] Therefore, in order to achieve the professional-level AI dubbing effect for film and television dramas, the embodiments of the present disclosure propose a voice synthesis method that can seamlessly convert domestic Chinese film and television dramas into overseas film and television works, realizing seamless cross-lingual voice and emotion conversion of actors, enabling audiences of different languages to hear the "original" dialogue. Combining the characteristics of film and television dramas, the voice synthesis method proposed in the embodiments of the present disclosure is equivalent to realizing a cross-lingual source language-target language joint emotion transfer voice synthesis method. The delicate emotional expression of the source language dubbing audio (such as Chinese dubbing audio) is extracted through an emotion extractor, and the emotion of the source language dubbing audio is transferred to the voice synthesis of the target language to achieve a synthesized audio with rich emotions.

[0036] The speech synthesis method according to the embodiments of the present disclosure can be deployed on various terminal devices through software or hardware transformation. The terminal devices involved in the embodiments of the present disclosure can refer to devices with wireless connection functions and / or wired connection functions. The wireless connection function means that it can be connected to other devices through wireless connection methods such as Wi-Fi and Bluetooth. The terminal devices involved in the embodiments of the present disclosure can also communicate with other devices through the wired connection function. The terminal devices involved in the embodiments of the present disclosure can be touch-screen, non-touch-screen, or without a screen. Touch-screen devices can be controlled by clicking, swiping, etc. on the display screen with fingers, styluses, etc. Non-touch-screen devices can be connected to input devices such as mice, keyboards, and touch panels to control the terminal devices. Devices without a screen can be, for example, Bluetooth speakers without a screen. For example, the terminal devices of the present application can include but are not limited to user equipment (UE), mobile devices, user terminals, terminals, handheld devices, tablet computers, laptop computers, palmtop computers, computing devices, etc.

[0037] The speech synthesis method according to the embodiments of the present disclosure can also be deployed on a server. The server can be located in the cloud or locally, and can be a physical device or a virtual device such as a virtual machine or a container, and has a wireless communication function. Among them, the wireless communication function can be set in the chip (system) or other components or assemblies of the server. It can refer to a device with a wireless connection function. The wireless connection function means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server involved in the embodiments of the present disclosure can also have the function of communicating through a wired connection. For example, the server according to the embodiments of the present disclosure can be located in the cloud, communicate with the terminal device, receive the source language dubbing audio corresponding to the source language subtitle text sent by the terminal device and the target language subtitle text translated from the source language dubbing audio, and use the speech synthesis method deployed on the server to generate the target language audio based on the target language subtitle text and the source language dubbing audio, and return it to the terminal device to generate the target language audio for the user in the terminal device.

[0038] Figure 1 The flowchart showing the speech synthesis method according to an embodiment of the present disclosure is as follows Figure 1 As shown, the method includes: step S11 to step S15.

[0039] In step S11, obtain the source language dubbing audio corresponding to the source language subtitle text of the source video, and the target language subtitle text translated from the source language subtitle text, where the source language is different from the target language.

[0040] Among them, the source video can be any video for which a dubbed audio in the target language to be generated is required. For example, the source video can be videos such as movies, short dramas, variety shows, etc. The source video can be a complete video or a clip from a video. The embodiments of the present disclosure do not limit this.

[0041] Among them, the source language subtitle text and the target language subtitle text of the source video are texts in different languages. For example, the source language subtitle text is Chinese subtitle text, and the target language subtitle text is Thai subtitle text, English subtitle text, etc.; of course, the source language subtitle text can also be Thai subtitle text or English subtitle text, etc., and the target language subtitle text is Chinese subtitle text. It should be understood that those skilled in the art can translate the source language subtitle text into subtitle text in any target language according to actual needs. The embodiments of the present disclosure do not limit this.

[0042] Considering that the source video usually corresponds to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. To improve the efficiency of generating the target language audio subsequently, the corresponding target language audio can be generated for each target language subtitle text respectively. Thus, in one possible implementation, a subtitle file corresponding to the source video can be obtained. The subtitle file includes multiple source language subtitle texts corresponding to the source video and a subtitle timeline. The subtitle timeline can represent the subtitle display time of each source language subtitle text. The subtitle display time includes the start time and the end time when the subtitle text is displayed. The subtitle timeline of the source language subtitle text is also the subtitle timeline that the target language subtitle text should use; then, according to the subtitle display time of each source language subtitle text indicated by the subtitle timeline, the audio segment corresponding to each source language subtitle text can be cut out from the complete dubbed audio file corresponding to the source video as the source language dubbed audio corresponding to each source language subtitle text. In this way, based on the source language dubbed audio corresponding to each source language subtitle text and the target language subtitle text translated from each source language subtitle text, the target language audio corresponding to each target language subtitle text can be generated respectively. Among them, a "piece" of subtitle text can refer to the subtitle text between a start time and a corresponding end time in the subtitle timeline, that is, the subtitle text that is displayed on the screen at the same time period.

[0043] In step S12, an emotion extractor is used to extract the audio emotion features from the source language dubbed audio. The audio emotion features represent the emotion expressed by the source language dubbed audio.

[0044] As described above, the source video may correspond to multiple source language subtitle texts. Each source language subtitle text can be segmented into its corresponding source language dubbed audio. That is, multiple source language subtitle texts can correspond to multiple source language dubbed audios. Therefore, an emotion extractor can be used to extract the audio emotion features in the source language dubbed audio corresponding to each source language subtitle text respectively.

[0045] In practical applications, those skilled in the art can adopt the artificial intelligence models publicly available in the art for extracting audio emotion features. For example, a Speech Emotion Recognition (SER) model can be used to extract audio emotion features. Specifically, the implicit emotion encoding vector output by the middle layer of the SER model can be used as the audio emotion features extracted from the source language dubbed audio. Of course, a self-developed model can also be adopted as long as it can achieve the functions that the emotion extractor can achieve. The embodiments of the present disclosure do not limit this.

[0046] Among them, there are two encoding methods for the emotion extractor, namely, the Wav-level (i.e., audio wave dimension) encoding method and the Frame-level (audio frame dimension) encoding method. The emotion extractor under the Wav-level encoding method outputs a one-dimensional feature vector, such as a feature vector of (1, 1024). This method is equivalent to extracting the overall emotion features of the entire source language dubbed audio. The emotion extractor under the Frame-level encoding method outputs a two-dimensional feature vector, such as a feature vector of (L, 1024), where the first dimension "L" represents the dimension of the frame and also represents the number of frames of the source language dubbed audio. This method is equivalent to extracting the emotion features of each frame of the L-frame audio of the source language dubbed audio. That is to say, the audio emotion features can include the overall emotion features of the source language dubbed audio and can also include the emotion features of each frame of the L-frame audio of the source language dubbed audio.

[0047] It should be understood that when using the Wav-level encoding method to extract audio emotion features, due to the slight noise in the low-dimensional space, it may significantly affect the accuracy of the extracted one-dimensional audio emotion features. However, the high-dimensional audio emotion features extracted by the Frame-level encoding method can smooth the influence of noise in the high-dimensional space and improve the accuracy of emotion feature extraction.

[0048] In step S13, the target language subtitle text is converted into a phoneme sequence, and the phoneme sequence is encoded by a text encoder to obtain text encoding features.

[0049] As described above, the source video may correspond to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. Therefore, each target language subtitle text can be converted into a phoneme sequence, and the text encoder can be used to encode the phoneme sequence converted from each target language subtitle text respectively to obtain the text encoding features of each target language subtitle text.

[0050] Among them, based on the phoneme dictionary corresponding to the target language, the conversion of the target language subtitle text into a phoneme sequence can be realized, and the embodiments of the present disclosure do not limit the conversion method from text to phoneme.

[0051] Among them, those skilled in the art can adopt the text encoder disclosed in the field of speech synthesis. For example, the text encoder in the VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) model can be adopted to encode the phoneme sequence converted from the target language subtitle text to obtain the text encoding features. Of course, a self-developed model can also be adopted as long as it can realize the functions that the text encoder can achieve, and the embodiments of the present disclosure do not limit this.

[0052] In step S14, the audio emotion feature and the text encoding feature are fused to obtain the emotion text feature.

[0053] As described above, the source video may correspond to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. In step S12, the audio emotion feature in the source language dubbed audio corresponding to each source language subtitle text can be extracted by using the emotion extractor respectively, and in step S13, the text encoder can be used to encode the phoneme sequence of the target language subtitle text translated from each source language subtitle text respectively to obtain the text encoding features of each target language subtitle text. Thus, in step S14, for each target language subtitle text, the text encoding feature of each target language subtitle text can be fused with the audio emotion feature in the source language dubbed audio of the source language subtitle text corresponding to the target language subtitle text to obtain the emotion text feature corresponding to each target language subtitle text, which is equivalent to embedding the emotion feature of the source language dubbed audio of each source language subtitle text into the text encoding feature of the corresponding each target language subtitle text, so as to obtain the emotion text feature with emotion information corresponding to each target language subtitle text.

[0054] As described above, the audio emotion feature can be a one-dimensional feature vector. Considering that the dimension of the text encoding feature is usually high, the one-dimensional audio emotion feature can be mapped to the feature dimension of the text encoding feature, and then the high-dimensional emotion feature mapped to the feature dimension of the text encoding feature is added to the text encoding feature to obtain the emotion text feature.

[0055] As described above, the audio emotion feature can also be a two-dimensional feature vector, that is, the audio emotion feature can include the emotion feature of each frame of the L-frame audio of the source language dubbed audio. Thus, fusing the audio emotion feature with the text encoding feature to obtain the emotion text feature may include:

[0056] By calculating the mean value of the emotion features of the L-frame audio in the audio emotion feature, the first emotion feature is obtained;

[0057] Performing high-dimensional mapping on the first emotion feature to obtain a second emotion feature, and the dimension of the second emotion feature is the same as that of the text encoding feature;

[0058] Adding the second emotion feature to the text encoding feature to obtain the emotion text feature.

[0059] It should be understood that when extracting the audio emotion feature using the Wav-level encoding method, due to the slight noise in the low-dimensional space, the accuracy of the extracted one-dimensional audio emotion feature may be reduced. However, the two-dimensional audio emotion feature extracted using the Frame-level encoding method can smooth the noise effect in the two-dimensional space to improve the accuracy of emotion feature extraction. Therefore, when the audio emotion feature is represented as a two-dimensional feature vector, the mean value of the audio emotion feature in the frame dimension can be fused with the text feature vector, and this method can significantly improve the accuracy and stability of emotion feature embedding.

[0060] In practical applications, feature high-dimensional mapping and feature fusion can be achieved by adding an emotion mapping module and a fusion module to the above text encoder. Specifically, the emotion mapping module is used to map the low-dimensional emotion vector represented by the first emotion feature to a high-dimensional space to align with the vector dimension of the text encoding feature extracted by the text encoder. For example, the emotion mapping module can perform high-dimensional mapping using a fully connected neural network (Fully Connected Layer) so that the mapped high-dimensional emotion vector can be better fused with the text encoding feature while retaining the expression ability of the emotion feature. The fusion module can be used to add the high-dimensional second emotion feature to the text encoding feature. For example, the information of different features can be added through the Broadcast technology: Fused emb =Phoneme emb +Emotion emb , Fusedemb Representing emotional text features, Phoneme emb Representing text encoding features, Emotion emb Representing the second emotional feature. In this way, through the superposition of feature vectors, emotional information can directly affect the speech synthesis process, thereby generating speech with specific emotions. Through this method, the emotional features of the source language dubbed audio extracted by the emotion extractor can be embedded into the text encoder of the speech synthesis model (such as VITS), constructing an emotion-embedded speech synthesis method with emotion transfer ability. In this way, not only can the target language audio with specific emotions be generated, but also cross-language emotion transfer can be achieved, that is, the emotion of the source language audio is transferred to the target language audio.

[0061] In step S15, the decoder is used to generate the target language audio based on the emotional text features, and the target language audio is used as the dubbed audio of the source video in the target language.

[0062] Among them, using the decoder to generate the target language audio based on the emotional text features is equivalent to synthesizing the audio with emotions using the text features with emotional information. It should be understood that those skilled in the art can adopt the decoder disclosed in the field of speech synthesis. For example, the decoder in the VITS model can be used to decode the emotional text features and generate the target language audio. Of course, a decoder can also be developed independently as long as it can achieve the functions that the decoder can achieve. The embodiments of the present disclosure do not limit this.

[0063] As described above, the source video can correspond to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. The emotional text features corresponding to each target language subtitle text can be obtained by using steps S12 to S14. Therefore, in step S15, the decoder can be used to generate the target language audio of each target language subtitle text based on the emotional text features corresponding to each target language subtitle text, and each target language audio carries the emotional expression in the corresponding source language dubbed audio.

[0064] In practical applications, after obtaining the target language audio of each target language subtitle text, the target language audio of each target language subtitle text can be processed such as splicing and audio-visual synchronization according to the subtitle time axis of the source language subtitle text (the subtitle time axis of the source language subtitle text is also the subtitle time axis required for the target language subtitle text) to obtain the complete dubbed audio of the source video in the target language, and the complete dubbed audio in the target language is synchronized with the video picture of the source video.

[0065] According to the speech synthesis method of the present disclosure, by extracting the audio emotion features of the source language dubbed audio and fusing them with the text encoding features of the target language subtitle text, emotion text features with the emotion information of the source language dubbed audio and the text information of the target language subtitle text are obtained. In this way, based on the emotion text features, a target language audio with the emotion expression of the source language dubbed audio can be generated, that is, it is possible to automatically and efficiently generate a high-quality target language audio with the emotion of the source language dubbed audio, without the need for professional dubbing equipment and dubbing conditions, and without consuming a large amount of manpower and financial resources, making the generation cost of the dubbed audio from the source language to the target language lower and the efficiency higher.

[0066] Figure 2 The flowchart showing another speech synthesis method according to an embodiment of the present disclosure is as Figure 2 shown, and the method includes:

[0067] Step S21, obtaining the source language dubbed audio corresponding to the source language subtitle text of the source video, and the target language subtitle text translated from the source language subtitle text, where the source language is different from the target language;

[0068] Step S22, using an emotion extractor to extract audio emotion features from the source language dubbed audio, where the audio emotion features represent the emotion expressed by the source language dubbed audio;

[0069] Step S23, converting the target language subtitle text into a phoneme sequence, and using a text encoder to encode the phoneme sequence to obtain text encoding features;

[0070] Step S24, fusing the audio emotion features with the text encoding features to obtain emotion text features;

[0071] Step S25, using a semantic extractor to extract semantic features from the phoneme sequence, where the semantic features represent the semantics of the target language subtitle text;

[0072] Step S26, fusing the emotion text features with the semantic features to obtain comprehensive features;

[0073] Step S27, using a decoder to generate a target language audio based on the comprehensive features.

[0074] Among them, the implementation manners of steps S21 to S24 can refer to the implementation manners of steps S11 to S14 in the above embodiments of the present disclosure, and will not be elaborated here.

[0075] In step S25, those skilled in the art can adopt an artificial intelligence model in the art that can extract text semantics. For example, BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder based on the Transformer) can be used as a semantic extractor to extract semantic features from the phoneme sequence. Of course, a self-developed semantic feature extraction model can also be used, and the embodiments of the present disclosure do not limit this. As described above, the source video can correspond to multiple source language subtitle texts, and multiple target language subtitle texts correspond to multiple source language subtitle texts. Therefore, the semantic extractor can be used to extract the semantic features of each target language subtitle text from the phoneme sequence of each target language subtitle text.

[0076] Steps S26 and S27 can be implemented as a way of step S15. In step S26, if the dimension of the semantic feature is the same as that of the emotion text feature, the emotion text feature and the semantic feature can be directly added to obtain a comprehensive feature; if the dimension of the semantic feature is different from that of the emotion text feature, the semantic feature can be mapped into a semantic feature with the same dimension as the emotion text feature, and then the semantic feature mapped to the same dimension and the emotion text feature are added to obtain a comprehensive feature; this comprehensive feature is equivalent to integrating emotion information, semantic information, and text information. In this way, using the comprehensive feature can make the generated target language audio have a high level in terms of tone, intonation, emotion, etc., that is, it can make the speech of the person in the target language audio more conform to the language expression in the target language while having the emotion expression in the source language dubbed audio, making the speech of the person in the target language audio more natural and real.

[0077] As described above, the source video can correspond to multiple source language subtitle texts, and multiple target language subtitle texts correspond to multiple source language subtitle texts. The emotion text features of each target language subtitle text can be obtained by using steps S22 to S24, and the semantic features of each target language subtitle text can be obtained by using step S25. Therefore, in step S26, the emotion text feature corresponding to each target language subtitle text can be fused with the corresponding semantic feature to obtain the comprehensive feature corresponding to each target language subtitle text.

[0078] In step S27, the decoder can also adopt the decoder in the VITS model, for example, to generate the target language audio based on the comprehensive features. Of course, the decoder can also be developed independently as long as it can achieve the functions that the decoder can achieve. The embodiments of the present disclosure do not limit this. As described above, the source video can correspond to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. After the comprehensive features corresponding to each target language subtitle text can be obtained by using steps S22 to S26, in step S27, the decoder can be used to generate the target language audio corresponding to each target language subtitle text based on the comprehensive features corresponding to each target language subtitle text, and the target language audio has the emotional expression in the source language dubbing audio and conforms to the language expression modes such as the tone and intonation in the target language.

[0079] In practical applications, after the target language audio of each target language subtitle text in the multiple target language subtitle texts is obtained through steps S21 to S27, the target language audio of each target language subtitle text can be processed such as splicing and audio-visual synchronization according to the subtitle time axis of the source language subtitle text, so as to obtain the complete dubbing audio of the source video in the target language, and the complete dubbing audio in the target language is synchronized with the video picture of the source video.

[0080] Based on the above steps S21 to S27, Figure 3 A schematic diagram showing a Chinese-Thai speech synthesis process is shown, as Figure 3 shown. The process includes: based on a Thai phoneme dictionary (the phoneme dictionary can be manually corrected to obtain an accurate phoneme sequence in the Thai language), converting Thai (that is, the subtitle text with the target language being Thai) into a phoneme sequence through phonemes, inputting the phoneme sequence into a text encoder to extract text encoding features, and at the same time inputting Chinese audio into an emotion recognition model (that is, an emotion extractor) to extract audio emotion features. An emotion mapping module and a fusion module are set in the text encoder to realize the fusion of the text encoding features and the audio emotion features, so as to obtain text & emotion features (that is, emotion text features). Then, the fusion module is used to fuse the semantic features extracted by a semantic extractor (not shown in the figure) with the text & emotion features to obtain text & emotion & semantic features (that is, comprehensive features), and input the text & emotion & semantic features into a decoder to obtain Thai audio.

[0081] According to the speech synthesis method of the embodiments of the present disclosure, by extracting the audio emotion features of the source language dubbed audio and fusing them with the text encoding features of the target language subtitle text, an emotional text feature with the emotional information of the source language dubbed audio and the text information of the target language subtitle text can be obtained. Then, by extracting the semantic features of the target language subtitle text and fusing the semantic features with the emotional text feature, a comprehensive feature with emotional information, text information, and semantic information can be obtained. Based on this comprehensive feature, a target language audio with the emotional expression of the source language dubbed audio and conforming to the language expressions such as tone and intonation in the target language can be generated. That is, the speech of the character in the target language audio can be made to conform more to the language expression in the target language while having the speaking emotion in the source language dubbed audio, making the speech of the character in the target language audio more natural and realistic. Without the need for professional dubbing equipment and dubbing conditions, and without consuming a large amount of manpower and financial resources, the generation cost of the dubbed audio from the source language to the target language is lower, the efficiency is higher, and the quality is higher.

[0082] Considering that in the audio production process, it is necessary to provide the source language dubbed audio as an emotion reference, but in actual production, the quality of the source language dubbed audio (such as the original dubbed audio corresponding to a Chinese video) cannot be controlled. There may even be extreme noisy audio in the source language dubbed audio, resulting in uncontrollable quality of the produced target language audio. In order to ensure the quality of the produced target language audio, in one possible implementation, after generating the target language audio based on the source language dubbed audio, that is, after the above steps S11 to step S15, or after steps S21 to step S27, as Figure 4 shown, the speech synthesis method may further include:

[0083] Step S41, using the speech recognition technology corresponding to the target language, convert the target language audio into a recognized text;

[0084] Step S42, based on the difference between the recognized text and the target language subtitle text, determine the character error rate of the recognized text, and the character error rate represents the pronunciation error rate of the target language audio;

[0085] Step S43, when the character error rate is higher than a preset threshold, screen out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library, so as to regenerate the target language audio using the audio emotion features of the target reference audio; wherein, the source language audio library includes reference audios of the source language with various emotions.

[0086] In step S41, if the target language is Thai, the Automatic Speech Recognition (ASR) technology of Thai can be used to convert the Thai audio into recognized text, that is, perform speech recognition of converting audio to text on the target language audio to obtain the recognized text. Of course, if the target language is English, the automatic speech recognition technology of English can be used to convert the English audio into recognized text. Furthermore, in step S42, the recognized text can be compared and analyzed with the original target language subtitle text to calculate the Character Error Rate (CER, a commonly used metric for speech recognition accuracy) of the recognized text compared to the target language subtitle text, so as to evaluate the pronunciation error rate (that is, pronunciation accuracy) of the target language audio, that is, evaluate the difference between the converted recognized text and the original target language subtitle text.

[0087] It should be understood that if the character error rate of the recognized text is lower than the preset threshold, it can be considered that the pronunciation error rate of the target language is lower, or the pronunciation accuracy is higher. In this case, there is no need to select a target reference audio from the source language audio library to regenerate the target language audio. If the character error rate of the recognized text is higher than the preset threshold, it can be considered that the pronunciation error rate of the target language audio is too high, or the pronunciation accuracy is lower. In this way, the target language audio with pronunciation problems can be efficiently and accurately identified. If the character error rate of the recognized text is higher than the preset threshold, it can be considered that the source language dubbing audio used to generate the target language audio with pronunciation problems is of poor quality.

[0088] Furthermore, after identifying the target language audio with pronunciation problems, a high-quality and emotion-matched target reference audio can be searched from the preset source language audio library to replace the corresponding source language dubbing audio, and the target language audio can be regenerated using the audio emotion characteristics of the target reference audio. In practical applications, the audio emotion characteristics of each reference audio in the source language audio can be extracted and stored in advance using an emotion extractor. In this way, after searching for any target reference audio from the source language audio library, the audio emotion characteristics of the target reference audio can be directly used to regenerate the target language audio, which is beneficial to ensuring consistent emotion expression and stable pronunciation of the produced target language audio.

[0089] In a possible implementation, the source language audio library may include a first audio library. The reference audios in the first audio library are annotated with emotion categories. This first audio library may be referred to as a discrete audio library. The first audio library may be formed by manually collecting high-quality source language reference audios containing discrete emotions. For example, the first audio library may include a total of 9 types of emotion reference audios (such as calm, surprised, pleasant, angry, sad, growl, afraid, excited, frustrated). Each emotion may correspond to at least one reference audio, and each reference audio may be annotated with the corresponding emotion category. Herein, the embodiments of the present disclosure do not limit the content expressed by each reference audio in the first audio library, but only require the emotion expressed by the reference audio. That is, as long as the emotion expressed by each reference audio conforms to its annotated emotion category. Based on this first audio library, in step S43 above, screening out the target reference audio that matches the emotion expressed by the source language dubbed audio from the preset source language audio library may include:

[0090] Using an emotion recognition model to perform emotion recognition on the source language dubbed audio to obtain an emotion classification result of the source language dubbed audio, where the emotion classification result represents the emotion category expressed by the source language dubbed audio;

[0091] Based on the emotion categories annotated for each reference audio in the first audio library, the reference audio with the emotion category matching the emotion classification result is used as the target reference audio.

[0092] Among them, the emotion recognition model may adopt a model for recognizing audio emotions disclosed in the art. For example, the SER model may be adopted. The embodiments of the present disclosure do not limit this. After using the emotion recognition model to identify the emotion classification result of the source language dubbed audio, the reference audio with the same emotion category as the emotion classification result can be searched from the first audio library as the target reference audio to reproduce the target language audio.

[0093] In a possible implementation, the source language audio library may include a second audio library. The reference audios in the second audio library are labeled with the audio emotion features of the reference audios extracted by an emotion extractor. Herein, the second audio library may be referred to as an emotion vector retrieval library. The second audio library may also be a high-quality source language reference audio collected manually. Different from the second audio library that only contains 9 emotions, the second audio library may contain reference audios with more emotions. Herein, the embodiments of the present disclosure do not limit the content expressed by each reference audio in the second audio library, and only the emotion expressed by the reference audio is required. The user may collect reference audios with more emotions and extract the emotion features of each reference audio to label the emotion expressed by the reference audio. Thus, the audio emotion features of each reference audio in the second audio library may be pre-extracted by the emotion extractor to characterize the emotion categories of each reference audio. Based on this, in the above step S43, screening out the target reference audio that matches the emotion expressed by the source language dubbed audio from the preset source language audio library may include:

[0094] Calculate the similarity between the audio emotion features of the source language dubbed audio and the emotion audio features labeled for each reference audio in the second audio library respectively, and use the reference audio with the highest similarity as the target reference audio. This method can be understood as obtaining the audio emotion features of the source language dubbed audio with pronunciation problems, and then obtaining the reference audio in the second audio library that is most similar to the audio emotion features of the source language dubbed audio as the new target reference audio by calculating the similarity between the audio emotion features of the source language dubbed audio and the emotion audio features labeled for each reference audio in the second audio library. Herein, the embodiments of the present disclosure do not limit the calculation method of the similarity between the two features. For example, the distance between the two features or the cosine similarity can be calculated, and the embodiments of the present disclosure do not limit this.

[0095] It should be understood that if the Figure 1 shown speech synthesis method is adopted, the process of regenerating the target language audio by using the audio emotion features of the target reference audio may include: fusing the audio emotion features of the target reference audio with the text encoding features of the target language subtitle text to obtain emotion text features, and then using the decoder to generate the target language audio based on the emotion text features. If the Figure 2 shown speech synthesis method is adopted, the process of regenerating the target language audio by using the audio emotion features of the target reference audio may include: fusing the audio emotion features of the target reference audio with the text encoding features of the target language subtitle text to obtain emotion text features, then fusing the emotion text features with the semantic features of the target language subtitle text to obtain comprehensive features, and then using the decoder to generate the target language audio based on the comprehensive features.

[0096] As described above, the source video may correspond to multiple source language subtitle texts, and the multiple source language subtitle texts correspond to multiple target language subtitle texts. Then, by using the above steps S11 to S15, or steps S21 to S27, the target language audio of each target language subtitle text in the multiple target language subtitle texts can be obtained. Therefore, for the target language audio of each target language subtitle text, the above steps S41 to S43 can be respectively executed to regenerate the target language audio of the target language subtitle text based on the reference audio in the source language audio library when the pronunciation error rate of the target language audio of any target language subtitle text is relatively high (that is, there is a problem with the pronunciation), so as to obtain high-quality target language audio for each target language subtitle text.

[0097] In practical applications, when using Figure 1 combined with Figure 4 the speech synthesis method, or Figure 2 combined with Figure 4 the speech synthesis method to obtain the target language audio of each target language subtitle text in the multiple target language subtitle texts, the target language audio of each target language subtitle text can be processed such as splicing and audio-visual synchronization according to the subtitle timeline of the source language subtitle text, so as to obtain the complete dubbed audio of the source video in the target language, and the complete dubbed audio is synchronized with the video picture of the source video.

[0098] In the embodiments of the present disclosure, by using the first audio library and the second audio library, high-quality target reference audio can be retrieved when the quality of the source language dubbed audio is low, thereby producing new high-quality target language audio and improving the quality of the dubbed audio generated by the source video in the target language.

[0099] For the emotion extractor, text encoder, and decoder used in the above speech synthesis method, the embodiments of the present disclosure also provide Figure 5 a three-stage training method shown in Figure 5 which can improve the stability of the entire emotion transfer speech synthesis model (that is, the emotion extractor, text encoder, and decoder). As

[0100] shown, the training method includes:

[0101] Step S51, in the first training stage, the initial text encoder and the initial decoder are trained using the first data set to obtain the trained first text encoder and the first decoder, where the first data set includes: multiple first target language text samples and the corresponding first target language audio samples for each first target language text sample;

[0101] Step S52: In the second training phase, use the second dataset to train the first text encoder, the first decoder, and the initial emotion extractor to obtain the trained second text encoder, the second decoder, and the first emotion extractor. The second dataset includes: a plurality of source language audio samples, the corresponding second target language text samples and the corresponding second target language audio samples for each source language audio sample.

[0102] Step S53: In the third training phase, use the third dataset to fine-tune the second text encoder, the second decoder, and the first emotion extractor to obtain the trained text encoder, decoder, and emotion extractor. The third dataset includes: a plurality of third target language texts and the corresponding third target language audio samples for each third target language text, where the third target language audio samples are obtained by collecting the audio of reading the third target language text in a standard pronunciation.

[0103] In step S51, the first training stage is equivalent to training the pronunciation base model of the entire emotion transfer speech synthesis model. Among them, a large number of publicly available corpora in the target language can be collected, and then the publicly available corpora in the target language can be processed through audio preprocessing and audio quality detection to obtain the first dataset. That is, the publicly available corpora in the target language can be processed into a high-quality first dataset. For example, when the target language is Thai, the first dataset can include a total of 35 speakers and 52 hours of Thai training samples (i.e., Thai text samples and corresponding Thai audio samples). Among them, the audio preprocessing process can include, for example: first, based on the start time and end time of each text sample in the publicly available corpus, obtain the audio sample corresponding to each text sample. At the same time, for each text sample and audio sample, use ASR technology to convert the audio sample into the corresponding recognized text, and then compare and analyze the recognized text with the text sample, and evaluate the pronunciation accuracy of the audio sample by calculating the character error rate (CER) of the recognized text. And based on the CER results of each audio sample, filter out the audio samples with CER results higher than the preset threshold. Through this CER screening process, the audio samples with pronunciation problems can be screened out; then, other sound effects such as background sounds and environmental sounds existing in the audio track of the filtered audio samples are separated. For example, known voice separation technologies in the art can be used. For example, the Demucs model can be used as a tool for voice separation to accurately extract the human voice track from the mixed audio track of the audio sample to eliminate the interference of background music or environmental sounds and ensure the purity and usability of the human voice in the generated audio sample. Through this voice separation process, the audio sample mainly contains the human voice track. Among them, the Demucs model adopts a U-Net convolutional architecture, and a bidirectional long short-term memory network (BiLSTM) is added between the encoder and the decoder, which can effectively separate the human voice and the accompaniment.

[0104] There may still be noise and low speech naturalness in some of the audio samples obtained through the above audio preprocessing process, which may affect the training of the model. Therefore, an audio quality detection process can be designed to further screen the audio samples from the perspective of audio quality. The audio quality detection process can include: by detecting the naturalness and noise value of each audio sample, removing the audio samples with naturalness and noise values lower than the preset score. For example, if the full score of naturalness and noise value is 5, the preset score can be 3 points. When any of the naturalness and noise values in any audio sample is lower than the preset score, it can be considered that the quality of this audio sample is unqualified, and it is also difficult to improve the audio quality of this part of the audio samples through methods such as speech enhancement. Then, this part of the audio samples with unqualified audio quality can be removed to ensure that only high-quality audio samples with a score greater than the preset score will be used as the audio samples in the first dataset.

[0105] In the first training stage, it is equivalent to hiding the emotion extractor in the emotion transfer speech synthesis model and first training the text encoder and decoder. That is to say, in the first training stage, emotion learning can be ignored, and the focus is mainly on training the model to learn the pronunciation and prosody of the target language. Therefore, cross-emotion transfer is not considered first, that is, the emotion extractor does not work and no parameter adjustment is made. The input and output are both in the target language, and the pronunciation base model is trained first (that is, the first text encoder and the first decoder that can convert the target language text into the target language audio are trained). Thus, in step S51, training the initial text encoder and the initial decoder using the first dataset to obtain the trained first text encoder and the first decoder can include:

[0106] Step S511, for any first target language text sample, use the initial text encoder to encode the phoneme sequence of the first target language text sample to obtain the text encoding feature of the first target language text sample;

[0107] Step S512, use the initial decoder to generate the predicted target language audio corresponding to the first target language text sample based on the text encoding feature of the first target language text sample;

[0108] Step S513, based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample, adjust the parameters of the initial text encoder and the initial decoder to obtain the trained first text encoder and the first decoder.

[0109] In step S511, the phoneme sequence of the first target language text sample can be obtained by performing phoneme conversion on the first target language text sample based on the phoneme dictionary of the target language. The embodiments of the present disclosure do not limit the types and structures of the initial text encoder and the initial decoder.

[0110] In step S513, based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample, the parameters of the initial text encoder and the initial decoder are adjusted. It can be understood that based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample, a loss is calculated. For example, the L1 loss, L2 loss, cross-entropy loss, etc. between the Mel spectrogram of the predicted target language audio and the Mel spectrogram of the first target language audio sample can be calculated. The embodiments of the present disclosure do not limit this. Then, the loss can be used to adjust the parameters of the initial text encoder and the initial decoder through backpropagation, gradient descent, etc. to obtain the trained first text encoder and the first decoder. It should be understood that in the first training stage, the initial text encoder and the initial decoder can be iteratively trained for multiple rounds. Therefore, the training batches can be divided based on the first dataset to utilize multiple batches of training samples to iteratively train the initial text encoder and the initial decoder. The embodiments of the present disclosure do not limit this.

[0111] In step S52, the second training stage is equivalent to training the emotion base model of the entire emotion transfer speech synthesis model. The model input is the target language text and the source language audio, and the output is the target language audio with the emotion of the source language audio. The training result of the first training stage is used as the training object. Among them, for example, by collecting the film and television drama resources from the source language to the target language (such as the film and television drama resources from Chinese to Thai), and then performing audio preprocessing and audio quality detection on the film and television drama resources to obtain the second dataset. That is, the film and television drama resources from the source language to the target language can be processed into a high-quality second dataset. For example, when the target language is Thai, the second dataset can include 9 speakers, a total of 7 hours of Chinese corpus (i.e., Chinese subtitles and Chinese dubbed audio) and 7 hours of Thai corpus (i.e., Thai subtitles and Thai dubbed audio). Thus, in a possible implementation manner, the construction process of the above second dataset may include:

[0112] Step S61, obtaining the bilingual subtitle file, the source language dubbed audio data, and the target language dubbed audio data corresponding to the original long video data, where the bilingual subtitle file includes the start time and end time of each bilingual subtitle text display, and the bilingual subtitle text includes the target language subtitle text and the source language subtitle text;

[0113] Step S61: Based on the start time and end time of each bilingual subtitle text in the bilingual subtitle file, segment the source-language dubbed audio data and the target-language dubbed audio data to obtain multiple segmented first audio groups. Each first audio group includes the segmented source-language audio segment and the target-language audio segment.

[0114] Step S62: Use the speech recognition technology corresponding to the target language to convert the target-language audio segment in each first audio group into recognition text, and determine the character error rate of the recognition text converted from each target-language audio segment based on the difference between the recognition text converted from each target-language audio segment and the target-language subtitle text corresponding to each target-language audio segment.

[0115] Step S63: Screen out the first audio groups where the target-language audio segments with a character error rate higher than the preset threshold are located among the multiple first audio groups to obtain multiple second audio groups.

[0116] Step S64: Separate the human voices from the source-language audio segments and the target-language audio segments in each second audio group to obtain multiple third audio groups. The source-language audio segments and the target-language audio segments in each third audio group contain human voice tracks.

[0117] Step S65: Detect the audio quality of the source-language audio segments and the target-language audio segments in each third audio group, and screen out the third audio groups with audio quality lower than the preset quality threshold among the multiple third audio groups to obtain multiple fourth audio groups, where the audio quality includes naturalness and / or noise value.

[0118] Step S66: Obtain the second dataset based on the source-language audio segments and the target-language audio segments in the multiple fourth audio groups, and the target-language subtitle text corresponding to the target-language audio segments in each fourth audio group.

[0119] Among them, the original long video data can be, for example, video resources such as movies, TV dramas, short dramas, variety shows, etc. with bilingual subtitles and bilingual dubbing. The embodiments of the present disclosure do not limit this. It should be understood that the bilingual subtitle file corresponding to the original long video data may include a subtitle timeline, and the subtitle timeline can indicate the start time start_time and end time end_time of each bilingual subtitle text display; or, if the original long video data does not have a bilingual subtitle file, it is also possible to first perform subtitle translation based on the source language subtitles and align the start_time and end_time of each subtitle in the source language subtitles and the target language subtitles based on the subtitle translation file to ensure that the subtitles from the source language to the target language are aligned with the audio. Since the dubbed audio in the target language may come from professional voice actors, and the professional voice actors may not completely dub according to the text content of the target language subtitle text, this may lead to inconsistency between the dubbed audio in the target language and the target language subtitle text. Therefore, after splitting the audio based on the start time and end time of the bilingual subtitle text display to obtain the split source language audio segment and target language audio segment, the speech recognition technology corresponding to the target language (such as ASR for Thai) can be used to convert the target language audio segment in each first audio group into a recognized text, and based on the difference between the recognized text converted from each target language audio segment and the target language subtitle text corresponding to each target language audio segment, the character error rate of the recognized text converted from each target language audio segment can be determined. Thus, the first audio group where the target language audio segment with a character error rate higher than the preset threshold is located can be screened out from multiple first audio groups, which is equivalent to screening out the target language audio with inconsistent pronunciation and subtitle and the corresponding source language audio from multiple first audio groups.

[0120] Among them, the known voice separation technology in the art, such as the Demucs model, can be used to separate the source language audio segment and the target language audio segment in each second audio group, which is equivalent to extracting the voice signals of the source language audio segment and the target language audio segment in each second audio group, so that the source language audio segment and the target language audio segment in the third audio group mainly contain pure voice tracks to eliminate the influence of background music and environmental noise.

[0121] In practical applications, noise detection methods known in the art can be adopted. For example, methods such as spectrum analysis, power spectral density measurement, signal-to-noise ratio calculation, etc. can be used to detect the noise values of the source language audio segment and the target language audio segment. Moreover, naturalness detection methods known in the art can be utilized. For example, metrics such as Mel cepstral distortion, fundamental frequency (F0) continuity, speech pause distribution, etc. can be adopted to detect audio naturalness, or a pre-trained audio quality detection model can also be used to detect the naturalness of the source language audio segment and the target language audio segment in each third audio group. The embodiments of the present disclosure do not limit this. Among them, the audio quality (noise value and naturalness) can be measured by scores respectively. The higher the score, the higher the naturalness and the lower the noise value. Therefore, the audio with an audio quality lower than the preset quality threshold (i.e., the preset score threshold) can be considered as the audio with unqualified audio quality. Furthermore, the third audio groups with audio quality lower than the preset quality threshold can be screened out from multiple third audio groups to obtain multiple fourth audio groups with higher quality. Then, based on the source language audio segments and target language audio segments in the multiple fourth audio groups, and the target language subtitle text corresponding to the target language audio segment in each fourth audio group, a second data set can be constructed. That is, each source language audio sample in the second data set can be the source language audio segment in each fourth audio group, the second target language text sample can be the target language subtitle text corresponding to the target language audio segment in each fourth audio group, and the second target language audio sample can be the target language audio segment in each fourth audio group.

[0122] It should be understood that high-quality audio samples can be obtained through the above-mentioned construction process of the second data set. In practical applications, it is also possible to continue to locate the roles corresponding to the source-language audio segments and target-language audio segments in each fourth audio group, that is, who is speaking each audio segment. For example, a speaker recognition model disclosed in the art can be used to extract the voiceprint features (i.e., timbre features) of the source-language audio segments in each fourth audio group. The speaker recognition model can be a speaker recognition model based on a densely connected delay neural network, with accurate speaker recognition effects and faster inference speeds. Specifically, the feature extraction component in the speaker recognition model can be used to separately extract the voiceprint features of the source-language audio segments in each fourth audio group and compress the voiceprint features into a vector space of the same dimension. Then, an unsupervised clustering algorithm is used to cluster the extracted voiceprint features. This clustering process can cluster the source-language audio segments in each fourth audio group into different categories according to their respective voiceprint features, and can gather audio with similar timbres (i.e., similar voiceprints) together. Each category can represent a role, and the clustering parameters can be continuously adjusted during this process to make the distinction between different categories as large as possible, while filtering out noise audio that cannot be classified. Thus, the role types corresponding to the source-language audio segments in each fourth audio group can be obtained, which is equivalent to obtaining the role types corresponding to the target-language audio segments in each fourth audio group. In this way, some or all of the fourth audio groups corresponding to the role types can be selected from the second data set to construct the second data set, so that the audio samples in the second data set can be multi-timbre (i.e., spoken by multiple people) audio. The embodiments of the present disclosure do not limit this.

[0123] Based on the above-mentioned second data set, in step S52, training the first text encoder, the first decoder, and the initial emotion extractor using the second data set to obtain the trained second text encoder, the second decoder, and the first emotion extractor may include:

[0124] Step S521, for the second target-language text sample corresponding to any source-language audio sample, encoding the phoneme sequence of the second target-language text sample using the first text encoder to obtain the text encoding features of the second target-language text sample;

[0125] Step S522, using the initial emotion extractor to extract the audio emotion features of the source-language audio sample from the source-language audio sample;

[0126] Step S523, fusing the text encoding features of the second target-language text sample with the audio emotion features of the source-language audio sample to obtain the emotion text features corresponding to the second target-language text sample;

[0127] Step S524: Use the first decoder to generate a predicted target language audio corresponding to the second target language text sample based on the emotion text features corresponding to the second target language text sample.

[0128] Step S525: Based on the difference between the predicted target language audio corresponding to the second target language text sample and the second target language audio sample corresponding to the second target language text sample, adjust the parameters of the first text encoder, the first decoder, and the initial emotion extractor to obtain the trained second text encoder, the first emotion extractor, and the second decoder.

[0129] Among them, the implementation manners of the above steps S521 to S524 can refer to the implementation manners of steps S12 to S15 in the above embodiments of the present disclosure, which will not be elaborated here.

[0130] In step S525, based on the difference between the predicted target language audio corresponding to the second target language text sample and the second target language audio sample corresponding to the second target language text sample, adjusting the parameters of the first text encoder, the first decoder, and the initial emotion extractor can be understood as calculating a loss based on the difference between the predicted target language audio corresponding to the second target language text sample and the corresponding second target language audio sample. For example, the L1 loss, L2 loss, cross-entropy loss, etc. between the Mel spectrogram of the predicted target language audio and the Mel spectrogram of the second target language audio sample can be calculated. The embodiments of the present disclosure do not limit this. Then, this loss can be used to adjust the parameters of the first text encoder, the first decoder, and the initial emotion extractor through backpropagation, gradient descent, etc., so as to obtain the trained second text encoder, the first emotion extractor, and the second decoder. It should be understood that in the second training stage, the first text encoder, the first decoder, and the initial emotion extractor can be iteratively trained for multiple rounds. Therefore, the training batches can be divided based on the second dataset to utilize multiple batches of training samples to iteratively train the first text encoder, the first decoder, and the initial emotion extractor. The embodiments of the present disclosure do not limit this.

[0131] In step S53, a more stable emotion transfer speech synthesis model can be trained using the third dataset. The model input includes the target language text and the target language audio, and the output is the target language audio. At the same time, the training result of the second training stage is used as the training object. Since the model trained in the second training stage has a particularly strong emotion performance but strongly depends on the quality of the input source language audio, in order to make the pronunciation more stable, based on the model trained in the second training stage, fine-tuning training can be performed based on multiple high-quality single-timbre third datasets to improve the stability of the model. Among them, the third target language audio samples in the third dataset can be obtained by recording the pronunciation of local people in the target language with standard pronunciation reading each of the preset third target language texts through professional equipment. This means that the pronunciation quality of the third target language audio samples in the third dataset is relatively high, so there is no need to perform preprocessing and audio quality detection on them. Among them, the third target language audio samples in the third dataset can be single-timbre, that is, each of the third target language audio samples in this third dataset can be from the same speaker; in practical applications, multiple single-timbre third datasets can be constructed for the third stage of training. For example, two single-timbre third datasets can be constructed, and the audio duration of each timbre can be 1 hour. The embodiments of the present disclosure do not limit this.

[0132] Based on the above third dataset, in step S53, the second text encoder, the second decoder, and the first emotion extractor are fine-tuned using the third dataset to obtain a trained text encoder, which may include:

[0133] Step S531: For any third target language text sample, use the second text encoder to encode the phoneme sequence of the third target language text sample to obtain the text encoding feature of the third target language text sample;

[0134] Step S532: Use the first emotion extractor to extract the audio emotion feature of the third target language audio sample from the third target language audio sample corresponding to the third target language text sample;

[0135] Step S533: Fuse the text encoding feature of the third target language text sample with the audio emotion feature of the corresponding third target language audio sample to obtain the emotion text feature of the third target language audio sample;

[0136] Step S534: Use the second decoder to generate the predicted target language audio corresponding to the third target language text sample based on the emotion text feature corresponding to the third target language audio sample;

[0137] Step S535: Based on the differences between the predicted target language audio corresponding to the third target language text sample and the third target language audio sample, adjust the parameters of the second text encoder, the second decoder, and the first emotion extractor to obtain the trained text encoder, emotion extractor, and decoder.

[0138] It should be understood that steps S531 to S534 can refer to the implementation manners of steps S12 to S15 in the above embodiments of the present disclosure, which will not be elaborated herein. The difference is that in step S523, the audio emotion features of the third target language audio sample are extracted, so that the emotion features in the third target language audio sample are used as the emotion reference to generate the target language audio.

[0139] In step S535, based on the differences between the predicted target language audio corresponding to the third target language text sample and the third target language audio sample, adjusting the parameters of the second text encoder, the second decoder, and the first emotion extractor can be understood as calculating the loss based on the differences between the predicted target language audio corresponding to the third target language text sample and the corresponding third target language audio sample. For example, the L1 loss, L2 loss, cross-entropy loss, etc. between the mel spectrogram of the predicted target language audio and the mel spectrogram of the third target language audio sample can be calculated. The embodiments of the present disclosure do not limit this. Then, the loss can be used to adjust the parameters of the second text encoder, the second decoder, and the first emotion extractor through backpropagation, gradient descent, etc., so as to obtain the trained text encoder, emotion extractor, and decoder. It should be understood that in the third training stage, the second text encoder, the second decoder, and the first emotion extractor can be iteratively trained for multiple rounds. Therefore, the training batches can be divided based on the third dataset to utilize multiple batches of training samples to iteratively train the second text encoder, the second decoder, and the first emotion extractor. The embodiments of the present disclosure do not limit this.

[0140] In practical applications, since it is found in the iterative training of the model that for pronunciation, emotion, and prosody, there are significant differences between men and women, and the model trained with separate male and female voices has better effects than the model trained with mixed male and female voices. Therefore, the first dataset, the second dataset, and the third dataset that are all male voices can be constructed respectively, and the first dataset, the second dataset, and the third dataset that are all female voices can be constructed respectively, and two emotion transfer speech synthesis models can be trained respectively according to the above three-stage training method to generate the target language audio of male and female voices respectively; or, the first dataset can be a dataset with mixed male and female voices, and the second dataset and the third dataset are datasets with separate male and female voices, to train two emotion transfer speech synthesis models. The embodiments of the present disclosure do not limit this.

[0141] Based on this, in the above Figure 1 orFigure 2 In the illustrated speech synthesis method, after obtaining any source language dubbed audio, it is also possible to first detect whether the voice line of the source language dubbed audio is male or female. If it is male, the emotion transfer speech synthesis model corresponding to the male voice is used to generate the target language audio. If it is female, the emotion transfer speech synthesis model corresponding to the female voice is used to generate the target language audio. Of course, it is also possible to use the first, second, and third datasets of mixed male and female voices to train a set of emotion transfer speech synthesis models, and use this emotion transfer speech synthesis model to generate the target language audio of male and female voices. The embodiments of the present disclosure do not limit this.

[0142] According to the three-stage training method of the embodiments of the present disclosure, in the first training stage, a large amount of target language audio corpus is used to train the base model of the speech synthesis model. In the second training stage, based on the parameters of the base model, a source-target language joint training is used to train the emotion-based speech synthesis model. In the third training stage, in order to ensure the stability of the timbre and pronunciation, fine-tuning is performed based on the single-timbre target language corpus. Through the above three-stage training method, an emotion transfer speech synthesis model with stable pronunciation and delicate emotions can be obtained.

[0143] As described above, a semantic extractor can also be used in the speech synthesis method. The semantic extractor can be a pre-trained semantic extraction model. During the above three-stage training process, the semantic extractor can be directly used to extract the semantic features of the phoneme sequences of each target language text sample, and the semantic features extracted by the semantic extractor are fused with the emotion text features corresponding to the target language audio sample to obtain comprehensive features, which are then involved in the model training process of each training stage. For example, in the second training stage, the semantic extractor can be used to extract the semantic features of the phoneme sequences of the second target language text sample, and the semantic features of the phoneme sequences of the second target language text sample are fused with the emotion text features corresponding to the second target language text sample to obtain the comprehensive features corresponding to the second target language text sample. Then, the first decoder is used to generate the predicted target language audio corresponding to the second target language text sample based on the comprehensive features corresponding to the second target language text sample, and then the model parameters are adjusted with reference to step S525.

[0144] It should be understood that the way the semantic extractor participates in the model training in the third training stage is the same as that in the second training stage described above, and will not be elaborated here; the semantic extractor may not participate in the model training in the first training stage. Of course, it may also participate in the first training stage. When the semantic extractor participates in the first training stage, the semantic features extracted by the semantic extractor can be fused with the text encoding features, and then the initial decoder is used to decode the fused features to generate the predicted target language audio corresponding to the first target language text sample, and then the model parameters are adjusted. The embodiments of the present disclosure do not limit this.

[0145] The three-stage training method according to the embodiments of the present disclosure realizes a set of data extraction and quality detection processes, accumulates a batch of high-quality joint emotion corpora, and adopts a source-target joint training method, which can provide better emotion transfer effects across languages. At the same time, the three-stage training method is adopted, which better improves the pronunciation effect of the target language audio output by the emotion transfer speech synthesis module compared with the prior art that only uses one fine-tuning training based on the pre-trained model. At the same time, a multi-class emotion recognition model is used to extract emotion features. Compared with existing speech synthesis schemes that all use the CLIP model (Contrastive Language-Image Pre-training) as the emotion feature extraction model (mainly using the multi-modal ability of the CLIP model between text and audio, but limited in the emotion extraction level), the multi-class emotion recognition model can better extract emotion features.

[0146] Based on the speech synthesis method and the three-stage training method proposed in the embodiments of the present disclosure, a cross-lingual source-target language joint emotion transfer speech synthesis framework is realized. This framework mainly involves the following technical links: 1) Construction of the joint corpus of source-target languages, that is, based on the rich audio resources of film and television dramas on the video platform, through methods such as audio preprocessing, quality detection models, and speaker clustering, a high-quality multi-voice joint TTS training corpus of source-target languages can be constructed, providing strong guarantee for training a delicate emotion speech synthesis model. 2) Emotion feature extraction, that is, with the help of a multi-class emotion recognition model trained with film and television drama emotion corpora, explicitly output the middle layer of the model as the emotion vector feature of the audio. 3) Emotion transfer speech synthesis, that is, embed the emotion feature vector output by the emotion extractor into a classic speech synthesis model, and construct a speech synthesis model with emotion transfer ability by modifying modules such as the text encoder of VITS. 4) Source-target language joint training framework, that is, in order to improve the stability of the trained emotion transfer speech synthesis model, a three-stage training method is adopted to improve the pronunciation effect of the target language audio output by the emotion transfer speech synthesis model. 5) Emotional speech synthesis stable production framework, that is, during the audio production process, due to the inability to control the quality of the source language audio, the quality of the produced target language audio is uncontrollable. In order to ensure the quality of the produced audio, the target language audio with poor pronunciation can be manually verified or automatically verified by ASR, and then a new target reference audio is matched based on the discrete audio library and the emotion vector retrieval library to replace the source language dubbing audio, so as to ensure that the produced audio is rich in emotion and stable in pronunciation.

[0147] Figure 6 A block diagram showing a speech synthesis device according to an embodiment of the present disclosure is as Figure 6 shown, and the device includes:

[0148] An acquisition module 601, configured to acquire the source language dubbed audio corresponding to the source language subtitle text of the source video, and the target language subtitle text translated from the source language subtitle text, where the source language is different from the target language;

[0149] A feature extraction module 602, configured to extract audio emotion features from the source language dubbed audio by using an emotion extractor, where the audio emotion features represent the emotion expressed by the source language dubbed audio;

[0150] A text encoding module 603, configured to convert the target language subtitle text into a phoneme sequence, and encode the phoneme sequence by using a text encoder to obtain text encoding features;

[0151] A feature fusion module 604, configured to fuse the audio emotion features and the text encoding features to obtain emotion text features;

[0152] An audio generation module 605, configured to generate target language audio by using a decoder based on the emotion text features, where the target language audio is used as the dubbed audio of the source video in the target language.

[0153] In a possible implementation manner, the audio emotion features include the emotion features of each frame of audio in the L-frame audio of the source language dubbed audio, where L represents the number of frames of the source language dubbed audio; wherein, the fusing the audio emotion features and the text encoding features to obtain emotion text features includes: calculating the mean value of the emotion features of the L-frame audio in the audio emotion features to obtain a first emotion feature; performing high-dimensional mapping on the first emotion feature to obtain a second emotion feature, where the dimension of the second emotion feature is the same as the dimension of the text encoding features; adding the second emotion feature and the text encoding features to obtain emotion text features.

[0154] In a possible implementation manner, the apparatus further includes:

[0155] A semantic extraction module, configured to extract semantic features from the phoneme sequence by using a semantic extractor, where the semantic features represent the semantics of the target language subtitle text;

[0156] The feature fusion module 604 is further configured to fuse the emotion text features and the semantic features to obtain comprehensive features;

[0157] The audio generation module 605 is further configured to generate target language audio by using a decoder based on the comprehensive features.

[0158] In a possible implementation manner, after generating the target language audio, the apparatus further includes:

[0159] A conversion module, configured to convert the target language audio into a recognized text by using the speech recognition technology corresponding to the target language;

[0160] An error rate determination module, configured to determine the character error rate of the recognized text based on the difference between the recognized text and the target language subtitle text, where the character error rate represents the pronunciation error rate of the target language audio;

[0161] A matching module, configured to, when the character error rate is higher than a preset threshold, screen out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library, so as to regenerate the target language audio by using the audio emotion feature of the target reference audio;

[0162] Wherein, the source language audio library includes reference audios of the source language in various emotions.

[0163] In a possible implementation manner, the source language audio library includes a first audio library, and the reference audios in the first audio library are labeled with emotion categories. Wherein, screening out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library includes: using an emotion recognition model to perform emotion recognition on the source language dubbed audio to obtain an emotion classification result of the source language dubbed audio, where the emotion classification result represents the emotion category expressed by the source language dubbed audio; based on the emotion categories labeled for each reference audio in the first audio library, using the reference audio whose emotion category matches the emotion classification result as the target reference audio.

[0164] In a possible implementation manner, the source language audio library includes a second audio library, and the reference audios in the second audio library are labeled with the emotion audio features of the reference audios extracted by the emotion extractor. Wherein, screening out a target reference audio that matches the emotion expressed by the source language dubbed audio from a preset source language audio library includes: respectively calculating the similarity between the audio emotion feature of the source language dubbed audio and the emotion audio features labeled for each reference audio in the second audio library, and using the reference audio with the highest similarity as the target reference audio.

[0165] In a possible implementation, the training process of the emotion extractor, text encoder, and decoder includes: in the first training stage, using a first dataset to train an initial text encoder and an initial decoder to obtain a trained first text encoder and a first decoder, where the first dataset includes: a plurality of first target language text samples and a first target language audio sample corresponding to each first target language text sample; in the second training stage, using a second dataset to train the first text encoder, the first decoder, and an initial emotion extractor to obtain a trained second text encoder, a second decoder, and a first emotion extractor; where the second dataset includes: a plurality of source language audio samples, a second target language text sample corresponding to each source language audio sample, and a corresponding second target language audio sample; in the third training stage, using a third dataset to perform fine-tuning training on the second text encoder, the second decoder, and the first emotion extractor to obtain a trained text encoder, decoder, and emotion extractor, where the third dataset includes: a plurality of third target language texts and a third target language audio sample corresponding to each third target language text, where the third target language audio sample is obtained by collecting the audio of a third target language text read aloud in a standard pronunciation.

[0166] In a possible implementation, the using the first dataset to train the initial text encoder and the initial decoder to obtain a trained first text encoder and a first decoder includes: for any first target language text sample, using the initial text encoder to encode the phoneme sequence of the first target language text sample to obtain a text encoding feature of the first target language text sample; using the initial decoder to generate a predicted target language audio corresponding to the first target language text sample based on the text encoding feature of the first target language text sample; based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample, adjusting the parameters of the initial text encoder and the initial decoder to obtain a trained first text encoder and a first decoder.

[0167] In a possible implementation, the fine-tuning training of the second text encoder, the second decoder, and the first emotion extractor using the third data set to obtain a trained text encoder includes: for any third target language text sample, encoding the phoneme sequence of the third target language text sample using the second text encoder to obtain the text encoding feature of the third target language text sample; extracting the audio emotion feature of the third target language audio sample from the third target language audio sample corresponding to the third target language text sample using the first emotion extractor; fusing the text encoding feature of the third target language text sample with the audio emotion feature of the corresponding third target language audio sample to obtain the emotion text feature corresponding to the third target language audio sample; generating a predicted target language audio corresponding to the third target language text sample using the second decoder based on the emotion text feature corresponding to the third target language audio sample; and adjusting the parameters of the second text encoder, the second decoder, and the first emotion extractor based on the difference between the predicted target language audio corresponding to the third target language text sample and the third target language audio sample to obtain a trained text encoder, emotion extractor, and decoder.

[0168] In a possible implementation, the process of constructing the second data set includes: obtaining a bilingual subtitle file corresponding to the original long video data, source language dubbed audio data, and target language dubbed audio data, where the bilingual subtitle file includes the start time and end time of each bilingual subtitle text display, and the bilingual subtitle text includes the target language subtitle text and the source language subtitle text; based on the start time and end time of each sentence of the bilingual subtitle text in the bilingual subtitle file, segmenting the source language dubbed audio data and the target language dubbed audio data to obtain multiple first audio groups after segmentation, and each first audio group includes the segmented source language audio segment and the target language audio segment; using the speech recognition technology corresponding to the target language to convert the target language audio segment in each first audio group into a recognized text, and based on the difference between the recognized text converted from each target language audio segment and the target language subtitle text corresponding to each target language audio segment, determining the character error rate of the recognized text converted from each target language audio segment; screening out the first audio groups where the target language audio segments with a character error rate higher than a preset threshold are located in the multiple first audio groups to obtain multiple second audio groups; separating the human voices from the source language audio segments and the target language audio segments in each second audio group to obtain multiple third audio groups, and the source language audio segments and the target language audio segments in each third audio group include human voice tracks; detecting the audio quality of the source language audio segments and the target language audio segments in each third audio group, and screening out the third audio groups with an audio quality lower than a preset quality threshold in the multiple third audio groups to obtain multiple fourth audio groups, where the audio quality includes naturalness and / or noise value; based on the source language audio segments and the target language audio segments in the multiple fourth audio groups, and the target language subtitle text corresponding to the target language audio segment in each fourth audio group, obtaining the second data set.

[0169] According to the speech synthesis method of the embodiments of the present disclosure, by extracting the audio emotion features of the source language dubbed audio and fusing them with the text coding features of the target language subtitle text, an emotional text feature with the emotional information of the source language dubbed audio and the text information of the target language subtitle text is obtained. In this way, based on the emotional text feature, a target language audio with the emotional expression of the source language dubbed audio can be generated, that is, it can automatically and efficiently generate a high-quality target language audio with the emotion of the source language dubbed audio, without the need for professional dubbing equipment and dubbing conditions, nor the need to consume a large amount of manpower and financial resources, making the generation cost of the dubbed audio from the source language to the target language lower and the efficiency higher.

[0170] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0171] An embodiment of the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.

[0172] An embodiment of the present disclosure also provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0173] An embodiment of the present disclosure also provides a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0174] Figure 7 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0175] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0176] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0177] A computer-readable storage medium can be a tangible device that can retain and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0178] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0179] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on a user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0180] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0181] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0182] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0183] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.

[0184] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.

Claims

1. A speech synthesis method, characterized in that: include: Obtaining source language dubbing audio corresponding to source language subtitle text of a source video, and target language subtitle text translated from the source language subtitle text, wherein the source language is different from the target language; Extracting audio emotion features from the source language dubbing audio using an emotion extractor, wherein the audio emotion features represent the emotions expressed by the source language dubbing audio; Converting the target language subtitle text into a phoneme sequence, and encoding the phoneme sequence using a text encoder to obtain text encoding features; The audio emotion feature is combined with the text encoding feature to obtain an emotion text feature; A decoder is used to generate target language audio based on the emotional text features, and the target language audio is used as dubbing audio of the source video in the target language.

2. The method according to claim 1, characterized in that The audio emotion feature comprises the emotion feature of each frame of audio in L frames of audio of the dubbing audio in the source language, where L represents the number of frames of the dubbing audio in the source language; The step of fusing the audio emotion feature with the text encoding feature to obtain the emotion text feature includes: Obtain a first emotion feature by calculating the mean of the emotion features of L frames of audio in the audio emotion feature; Performing high-dimensional mapping on the first emotion feature to obtain a second emotion feature, wherein the dimension of the second emotion feature is the same as the dimension of the text encoding feature; The second emotion feature is added to the text encoding feature to obtain an emotion text feature.

3. The method according to claim 1, characterized in that The method further comprises: Extracting semantic features from the phoneme sequence using a semantic extractor, wherein the semantic features represent the semantics of the target language subtitle text; The step of using a decoder to generate a target language audio based on the emotional text features includes: The emotional text feature is integrated with the semantic feature to obtain a comprehensive feature; The decoder is used to generate audio in the target language based on the comprehensive features.

4. The method according to any one of claims 1 to 3, characterized in that: After generating the target language audio, the method further includes: Using speech recognition technology corresponding to the target language, converting the target language audio into recognized text; Determining a character error rate of the recognized text based on a difference between the recognized text and the target language subtitle text, wherein the character error rate represents a pronunciation error rate of the target language audio; When the character error rate is higher than a preset threshold, a target reference audio matching the emotion expressed by the source language dubbing audio is selected from a preset source language audio library, so as to regenerate the target language audio by using the audio emotion features of the target reference audio; The source language audio library includes reference audios of source languages ​​with multiple emotions.

5. The method according to claim 4, characterized in that The source language audio library includes a first audio library, and the reference audio in the first audio library is annotated with an emotion category, wherein the step of selecting a target reference audio that matches the emotion expressed by the source language dubbing audio from the preset source language audio library includes: Using an emotion recognition model to perform emotion recognition on the source language dubbing audio, and obtaining an emotion classification result of the source language dubbing audio, wherein the emotion classification result represents the emotion category expressed by the source language dubbing audio; Based on the emotion categories marked on each reference audio in the first audio library, the reference audio whose emotion category matches the emotion classification result is used as the target reference audio.

6. The method according to claim 4, characterized in that The source language audio library includes a second audio library, and the reference audio in the second audio library is annotated with the emotional audio features of the reference audio extracted by the emotion extractor, wherein the step of selecting the target reference audio that matches the emotion expressed by the source language dubbing audio from the preset source language audio library includes: The similarities between the audio emotion features of the dubbing audio in the source language and the emotional audio features annotated by each reference audio in the second audio library are calculated respectively, and the reference audio with the highest similarity is used as the target reference audio.

7. The method according to claim 1, characterized in that The training process of the emotion extractor, text encoder and decoder includes: In a first training stage, an initial text encoder and an initial decoder are trained using a first data set to obtain a trained first text encoder and a first decoder, wherein the first data set includes: a plurality of first target language text samples and a first target language audio sample corresponding to each first target language text sample; In the second training stage, the first text encoder, the first decoder and the initial emotion extractor are trained using a second data set to obtain a trained second text encoder, a second decoder and a first emotion extractor; wherein the second data set includes: a plurality of source language audio samples, a second target language text sample corresponding to each source language audio sample and a corresponding second target language audio sample; In the third training stage, the second text encoder, the second decoder and the first emotion extractor are fine-tuned using a third data set to obtain trained text encoders, decoders and emotion extractors, wherein the third data set includes: multiple third target language texts and third target language audio samples corresponding to each third target language text, wherein the third target language audio samples are obtained by collecting audio of the third target language text read aloud in standard pronunciation.

8. The method according to claim 7, characterized in that The method of using the first data set to train an initial text encoder and an initial decoder to obtain a trained first text encoder and a first decoder includes: For any first target language text sample, using an initial text encoder to encode the phoneme sequence of the first target language text sample to obtain a text encoding feature of the first target language text sample; generating, using the initial decoder based on the text encoding features of the first target language text sample, a predicted target language audio corresponding to the first target language text sample; Based on the difference between the predicted target language audio corresponding to the first target language text sample and the first target language audio sample corresponding to the first target language text sample, the parameters of the initial text encoder and the initial decoder are adjusted to obtain a trained first text encoder and a first decoder.

9. The method according to claim 7, characterized in that: The method of fine-tuning the second text encoder, the second decoder, and the first emotion extractor using the third data set to obtain a trained text encoder includes: For any third target language text sample, using the second text encoder to encode the phoneme sequence of the third target language text sample to obtain a text encoding feature of the third target language text sample; Extracting audio emotion features of the third target language audio sample from the third target language audio sample corresponding to the third target language text sample using the first emotion extractor; fusing the text encoding features of the third target language text sample with the audio emotion features of the corresponding third target language audio sample to obtain the emotion text features corresponding to the third target language audio sample; Using the second decoder to generate a predicted target language audio corresponding to the third target language text sample based on the emotional text features corresponding to the third target language audio sample; Based on the difference between the predicted target language audio corresponding to the third target language text sample and the third target language audio sample, the parameters of the second text encoder, the second decoder and the first emotion extractor are adjusted to obtain a trained text encoder, emotion extractor and decoder.

10. The method according to claim 7, characterized in that The construction process of the second data set includes: Obtaining a bilingual subtitle file, source language dubbing audio data, and target language dubbing audio data corresponding to the original long video data, wherein the bilingual subtitle file includes the start time and end time of each bilingual subtitle text display, and the bilingual subtitle text includes the target language subtitle text and the source language subtitle text; Based on the start time and end time of displaying each bilingual subtitle text in the bilingual subtitle file, the source language dubbing audio data and the target language dubbing audio data are segmented to obtain a plurality of segmented first audio groups, each of which includes a segmented source language audio segment and a segmented target language audio segment; Using the speech recognition technology corresponding to the target language, convert the target language audio segment in each first audio group into a recognized text, and determine the character error rate of the recognized text converted from each target language audio segment based on the difference between the recognized text converted from each target language audio segment and the target language subtitle text corresponding to each target language audio segment; The first audio groups including the target language audio segments having a character error rate higher than a preset threshold are screened out from the plurality of first audio groups to obtain a plurality of second audio groups; By performing vocal separation on the source language audio segment and the target language audio segment in each second audio group, a plurality of third audio groups are obtained, wherein the source language audio segment and the target language audio segment in each third audio group contain a vocal track; Detecting the audio quality of the source language audio segment and the target language audio segment in each third audio group, and filtering out the third audio groups whose audio quality is lower than a preset quality threshold from the plurality of third audio groups, to obtain a plurality of fourth audio groups, wherein the audio quality includes naturalness and / or noise value; The second data set is obtained based on the source language audio segments and the target language audio segments in the plurality of fourth audio groups, and the target language subtitle text corresponding to the target language audio segment in each fourth audio group.

11. A speech synthesis device, characterized in that: include: An acquisition module, used to acquire source language dubbing audio corresponding to source language subtitle text of a source video, and target language subtitle text translated from the source language subtitle text, wherein the source language is different from the target language; A feature extraction module, used to extract audio emotion features from the source language dubbing audio using an emotion extractor, wherein the audio emotion features represent the emotions expressed by the source language dubbing audio; A text encoding module, used for converting the target language subtitle text into a phoneme sequence, and encoding the phoneme sequence using a text encoder to obtain text encoding features; A feature fusion module, used to fuse the audio emotion feature with the text encoding feature to obtain an emotion text feature; The audio generation module is used to generate target language audio based on the emotional text features using a decoder, and the target language audio is used as the dubbing audio of the source video in the target language.

12. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.

13. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer program product, comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Voice synthesis method and device, storage medium and electronic device

    CN111653265A

  • Speech synthesis model training method, speech synthesis method and speech synthesis device

    CN114783409A

  • Confrontation and meta-learning method based on speaker emotion speech synthesis model

    CN115359778A

  • Automatic voice data verification method for voice synthesis

    CN116524899A

  • Voice data synthesis method and device, electronic equipment and storage medium

    CN116863910A

Cited By

  • Emotion recognition method and device in film and television play dubbing

    CN120319273A

  • A method and device for emotion recognition in dubbing of a movie or a TV series

    CN120319273B

  • Speech synthesis method and system with emotion recognition capability

    CN121214907A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN121565138A

  • Video translation method and device

    CN121814981A