Speech synthesis method, speech synthesis device, electronic equipment and storage medium
By extracting and encoding speech features, combining style encoder and attention encoding technology, the problem of insufficient timbre consistency and naturalness in existing speech synthesis technologies is solved, and a more efficient speech synthesis effect is achieved.
Patent Information
- Application Number
- CN202510354047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing speech synthesis technology has shortcomings in timbre consistency and naturalness, resulting in poor naturalness of synthetic speech.
By obtaining the source voice data of the source speaker and the target voice data of the target speaker, the target timbre features, the source language content features and the source initial style features are extracted, and the style encoder is used for style encoding and attention encoding, the purpose-coded voice features are generated, and the target synthetic voice data is finally obtained through speech decoding.
It improves the timbre consistency and nature of speech synthesis, and enhances the independence and stability of the tone processing of the target speaker and the style processing of the source speaker.
Smart Images

Figure CN120220640A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and is applicable to the fields of fintech and healthcare, and particularly relates to a speech synthesis method, a speech synthesis device, an electronic device, and a storage medium. Background Art
[0002] Currently, the voice conversion technology in speech synthesis aims to transfer the timbre of the source speaker's voice to the target speaker while preserving the language content of the source speaker's voice. For example, in the scenarios of audiobooks and short video dubbing in the fintech field, the timbre in the book narration voice or video narration voice of dubber A can be transferred to dubber B, and the language content in the book narration voice or video narration voice (such as: This is a short video about how to brush your teeth correctly) is preserved. Another example is that in the intelligent medical guidance scenario in the healthcare field, the timbre in the treatment guidance voice pre-recorded by medical staff can be transferred to a robot, and the language content in the treatment guidance voice (such as: Please ask patient XX to go to consulting room 1 for treatment) is preserved.
[0003] However, the related technologies still have many limitations. For example, the related speech synthesis methods mainly focus on the transfer of timbre and have insufficient control ability over language styles (such as intonation, speech rate, and emotion), lacking style control, resulting in poor naturalness of the synthesized speech. Another example is that although the related speech synthesis methods attempt style conversion, they usually confuse timbre with style, that is, when processing the style characteristics of the source speaker, the timbre characteristics of the source speaker are often mixed in, resulting in the failure to achieve ideal effects in terms of timbre consistency and naturalness in speech synthesis.
[0004] Therefore, the related technologies have problems of poor timbre consistency and poor naturalness in speech synthesis. Summary of the Invention
[0005] The main objective of the embodiments of the present application is to provide a speech synthesis method, a speech synthesis device, an electronic device, and a storage medium, which can improve the timbre consistency of speech synthesis and improve the naturalness of speech synthesis.
[0006] To achieve the above objective, a first aspect of the embodiments of the present application provides a speech synthesis method, and the method includes:
[0007] Obtain source speech data of a source speaker and target speech data of a target speaker; wherein, the source speaker and the target speaker are different;
[0008] Extract the timbre feature from the target speech data to obtain a target timbre feature;
[0009] Extract the feature of the source language content from the source speech data to obtain a source language content feature;
[0010] Extract the style of the source speech data to obtain the initial source style features;
[0011] Perform style encoding on the initial source style features through a preset style encoder to obtain enhanced source style features;
[0012] Perform attention encoding on the enhanced source style features, the source language content features, and the target voice characteristics through the style encoder to obtain the target encoded speech features;
[0013] Perform speech decoding on the enhanced source style features, the target encoded speech features, and the target voice characteristics to obtain the target synthesized speech data; wherein, the voice of the target synthesized speech data is from the target speaker, and the style and language content of the target synthesized speech are from the source speaker.
[0014] Optionally, the performing attention encoding on the enhanced source style features, the source language content features, and the target voice characteristics through the style encoder to obtain the target encoded speech features includes:
[0015] Perform attention processing on the source language content features through the style encoder to obtain a language content attention vector;
[0016] Perform attention processing on the enhanced source style features through the style encoder to obtain a style attention vector;
[0017] Perform attention processing on the target voice characteristics through the style encoder to obtain a voice characteristic attention vector;
[0018] Perform style enhancement on the style attention vector and a preset first encoding parameter through the style encoder to obtain a style-enhanced attention vector;
[0019] Perform voice characteristic enhancement on the voice characteristic attention vector and a preset second encoding parameter through the style encoder to obtain a voice characteristic-enhanced attention vector;
[0020] Perform vector fusion on the language content attention vector, the style-enhanced attention vector, and the voice characteristic-enhanced attention vector through the style encoder to obtain the target encoded speech features.
[0021] Optionally, the performing style enhancement on the style attention vector and a preset first encoding parameter through the style encoder to obtain a style-enhanced attention vector includes: performing non-linear activation on the first encoding parameter through a predetermined activation function to obtain a first weight; multiplying each element of the style attention vector by the first weight to obtain a style-enhanced attention vector;
[0022] Enhancing the timbre attention vector and a preset second coding parameter through the style encoder to obtain a timbre-enhanced attention vector, including: non-linearly activating the second coding parameter through the predetermined activation function to obtain a second weight; multiplying each element of the timbre attention vector by the second weight to obtain a timbre-enhanced attention vector.
[0023] Optionally, before encoding the source-enhanced style feature, the source language content feature, and the target timbre feature through the style encoder to obtain a target-encoded speech feature, the method further includes:
[0024] Pre-training the style encoder, specifically including:
[0025] Obtaining sample source speech data of a sample source speaker and sample target speech data of a sample target speaker, extracting a sample source style feature, a sample source language content feature from the sample source speech data, and extracting a sample target timbre feature from the sample target speech data; wherein, the sample source speaker and the sample target speaker are different;
[0026] Performing style encoding on the sample source style feature through a preset initial style encoder to obtain a sample-enhanced style feature;
[0027] Performing attention processing on the sample source language content feature through the initial style encoder to obtain a sample language content attention vector;
[0028] Performing attention processing on the sample-enhanced style feature through the initial style encoder to obtain a sample style attention vector;
[0029] Performing attention processing on the sample target timbre feature through the initial style encoder to obtain a sample timbre attention vector;
[0030] Enhancing the sample style attention vector and a preset first coding parameter through the initial style encoder to obtain a sample style-enhanced attention vector;
[0031] Enhancing the sample timbre attention vector and a preset second coding parameter through the initial style encoder to obtain a sample timbre-enhanced attention vector;
[0032] Performing vector fusion on the sample language content attention vector, the sample style-enhanced attention vector, and the sample timbre-enhanced attention vector through the initial style encoder to obtain a sample speech coding feature;
[0033] Decode the sample speech coding features to obtain sample synthesized speech data; wherein, the timbre of the sample synthesized speech data is from the sample target speaker, and the style and language content of the sample synthesized speech are from the sample source speaker;
[0034] Calculate the loss based on the sample synthesized speech data and the preset labeled speech data to obtain the target loss function;
[0035] Update the first coding parameter and the second coding parameter according to the target loss function, and adjust the parameters of the initial style encoder according to the target loss function to obtain the style encoder.
[0036] Optionally, the calculating the loss based on the sample synthesized speech data and the preset labeled speech data to obtain the target loss function includes:
[0037] Calculate the speech similarity between the sample synthesized speech data and the preset labeled speech data to obtain the speech loss function;
[0038] Extract the style from the labeled speech data to obtain the labeled style features;
[0039] Calculate the style similarity between the sample speech coding features and the labeled style features to obtain the first style loss function;
[0040] Calculate the style similarity between the sample enhanced style features and the labeled style features to obtain the second style loss function;
[0041] Fuse the losses according to the speech loss function, the first style loss function and the second style loss function to obtain the target loss function.
[0042] Optionally, the obtaining the source enhanced style features by performing style encoding on the source initial style features through a preset style encoder includes:
[0043] Randomly divide the source initial style features to obtain at least two source initial style sub-features;
[0044] Perform attention processing on each of the at least two source initial style sub-features through the style encoder to obtain style attention sub-vectors;
[0045] Generate attention sub-weights according to each source initial style sub-feature;
[0046] Fuse the features according to the attention sub-weights and the style attention sub-vectors of each source initial style sub-feature to obtain the source enhanced style features.
[0047] Optionally, generating the attention sub-weights according to each of the source initial style sub-features includes:
[0048] Performing parameter mapping according to elements in the source initial style sub-features to obtain attention parameters;
[0049] Performing non-linear activation on the attention parameters through a predetermined activation function to obtain the attention sub-weights.
[0050] To achieve the above object, a second aspect of the embodiments of the present application proposes a speech synthesis device, the device includes:
[0051] A speech acquisition module, configured to acquire source speech data of a source speaker and target speech data of a target speaker;
[0052] A timbre extraction module, configured to extract the timbre from the target speech data to obtain target timbre features;
[0053] A feature extraction module, configured to extract features from the source speech data to obtain source language content features;
[0054] A style extraction module, configured to extract the style from the source speech data to obtain source initial style features;
[0055] A style encoding module, configured to perform style encoding on the source initial style features through a preset style encoder to obtain source enhanced style features;
[0056] An attention encoding module, configured to perform attention encoding on the source enhanced style features, the source language content features, and the target timbre features through the style encoder to obtain target encoded speech features;
[0057] A language decoding module, configured to perform speech decoding on the source enhanced style features, the target encoded speech features, and the target timbre features to obtain target synthesized speech data; wherein, the timbre of the target synthesized speech data comes from the target speaker, and the style and language content of the target synthesized speech come from the source speaker.
[0058] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the speech synthesis method described in the first aspect above is implemented.
[0059] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the speech synthesis method described in the first aspect above is implemented.
[0060] For the speech synthesis method, speech synthesis device, electronic device and storage medium provided by the present application, first obtain the source speech data of the source speaker and the target speech data of the target speaker. The source speech data is used to provide language content and style, and the target speech data is used to provide timbre. Further, the following steps can be executed in parallel: extract the timbre features from the target speech data to obtain target timbre features; extract the feature of the source language content from the source speech data to obtain source language content features; extract the initial style features from the source speech data to obtain source initial style features. Due to the complexity of speech, in addition to style information, the source initial style features also contain some timbre information. Further, the source initial style features are encoded by a style encoder to obtain source enhanced style features. In this way, the style encoder can effectively highlight the style information in the source initial style features and ignore the timbre information, and obtain source enhanced style features with richer style information. Further, the source enhanced style features, source language content features and target timbre features are encoded by the style encoder through an attention mechanism to obtain target encoded speech features. In this way, the attention mechanism of the style encoder can enhance the independence and stability of the timbre processing of the target speaker and the style processing of the source speaker. Finally, the source enhanced style features, target encoded speech features and target speaker timbre features are decoded to obtain target synthesized speech data. In this way, since the target encoded speech features carry language content information, style information and timbre information, and the source enhanced style features and target timbre features are additionally added during decoding, the influence of the timbre of the source speaker on the style can be further reduced during speech decoding, and the independence and stability between the style features of the source speaker and the timbre features of the target speaker can be improved. In summary, the present application can improve the timbre consistency of speech synthesis and the naturalness of speech synthesis.
[0061] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flowchart of the speech synthesis method provided by the embodiments of the present application;
[0063] Figure 2 is a flowchart of the speech synthesis method provided by another embodiment of the present application;
[0064] Figure 3 isFigure 2 The flowchart of step 207 in
[0065] Figure 4 is Figure 1 The flowchart of step 105 in
[0066] Figure 5 is Figure 1 The flowchart of step 106 in
[0067] Figure 6 The module structure block diagram of the speech synthesis device provided by the embodiment of the present application;
[0068] Figure 7 The schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0069] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0070] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the description and claims and the above drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0072] First, several nouns involved in the present application are analyzed:
[0073] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. AI is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. AI can simulate the information process of human consciousness and thinking. It is also a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0074] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of AI and an interdisciplinary field of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information image processing, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, AI research, and linguistic research related to language computing, etc.
[0075] Zero-Shot Text-to-Speech (Zero-Shot TTS) is a technology for synthesizing speech. It can generate the speech of any speaker without specific speech data. This technology utilizes deep learning and neural networks. By training on a dataset of multiple speakers, it can generate an output of the speech style of a specific speaker after inputting text. Zero-Shot TTS does not require collecting independent speech data for each speaker, thus greatly improving the synthesis efficiency and flexibility.
[0076] When performing speech synthesis, related technologies usually have at least one of the following defects: (1) Lack of style control: For example, speech synthesis methods mainly focus on the transmission of timbre and have insufficient control over language styles (such as intonation, speech rate, and emotion), resulting in poor naturalness of the synthesized speech. Another example is that although speech synthesis methods attempt style conversion, they usually confuse timbre with style, that is, when processing the style characteristics of the source speaker, the timbre characteristics of the source speaker are often mixed in, and the style cannot be effectively controlled. (2) Slow inference speed: For example, zero-shot speech conversion technologies based on speech models or diffusion models have low inference efficiency due to relying on autoregressive generation or multi-step sampling. Another example is that although language model methods can generate high-quality speech, their inference time increases with the increase of the input length, affecting practical use. Diffusion models require multiple reverse sampling steps, resulting in high computational overhead. (3) Sound quality and similarity problems: There is still much room for improvement in the naturalness and sound quality of the speech generated by related speech synthesis methods, especially in terms of timbre and the consistency with the target speaker, and the ideal effect has not been achieved.
[0077] The above defects make it difficult for existing speech synthesis methods to meet the actual application scenarios in the requirements of efficient, flexible, and high-quality speech conversion, such as audiobook production scenarios, multimedia content creation scenarios (including short video dubbing scenarios), and multilingual learning tool scenarios (including language exchange scenarios).
[0078] Based on this, the embodiments of the present application propose a speech synthesis method, a speech synthesis device, an electronic device, and a computer-readable storage medium, which can improve the timbre consistency of speech synthesis, improve the naturalness of speech synthesis, and also improve the synthesis efficiency.
[0079] The speech synthesis method provided by the embodiments of the present application can be applied to terminals and server sides, and can also be software running on the server side. The server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the speech synthesis method, etc., but is not limited to the above forms.
[0080] This application can be used in numerous general or specific computer system environments or configurations. For example: server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0081] Embodiments of this application provide a speech synthesis method, a speech synthesis device, an electronic device, and a computer-readable storage medium, which will be specifically described through the following embodiments. First, the speech synthesis method in the embodiments of this application will be described.
[0082] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as the user's speech data, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards.
[0083] Referring to Figure 1 , Figure 1 is an optional flowchart of the speech synthesis method provided by the embodiments of this application, which may include but is not limited to steps 101 to 107.
[0084] Step 101, obtain the source speech data of the source speaker and the target speech data of the target speaker; wherein, the source speaker and the target speaker are different;
[0085] Step 102, extract the timbre of the target speech data to obtain the target timbre feature;
[0086] Step 103, extract the features of the source speech data to obtain the source language content feature;
[0087] Step 104, extract the style of the source speech data to obtain the source initial style feature;
[0088] Step 105, perform style encoding on the source initial style feature through a preset style encoder to obtain the source enhanced style feature;
[0089] Step 106, perform attention encoding on the source enhanced style feature, the source language content feature, and the target timbre feature through the style encoder to obtain the target encoded speech feature;
[0090] Step 107: Perform speech decoding on the source enhanced style feature, the target encoded speech feature, and the target voice color feature to obtain target synthesized speech data; wherein, the voice color of the target synthesized speech data comes from the target speaker, and the style and language content of the target synthesized speech come from the source speaker.
[0091] Steps 101 to 107 illustrated in the embodiments of the present application first obtain the source speech data of the source speaker and the target speech data of the target speaker. The source speech data is used to provide language content and style, and the target speech data is used to provide voice color. Further, the following steps can be executed in parallel: Extract the voice color feature from the target speech data to obtain the target voice color feature; extract the feature of the source language content from the source speech data to obtain the source language content feature; extract the style from the source speech data to obtain the source initial style feature. Due to the complexity of speech, in addition to style information, the source initial style feature also contains some voice color information. Further, perform style encoding on the source initial style feature through a style encoder to obtain a source enhanced style feature. In this way, the style encoder can effectively highlight the style information in the source initial style feature and ignore the voice color information, obtaining a source enhanced style feature with richer style information. Further, perform attention encoding on the source enhanced style feature, the source language content feature, and the target voice color feature through the style encoder to obtain the target encoded speech feature. In this way, the attention mechanism of the style encoder can enhance the independence and stability of the voice color processing of the target speaker and the style processing of the source speaker. Finally, perform speech decoding on the source enhanced style feature, the target encoded speech feature, and the target speaker's voice color feature to obtain the target synthesized speech data. In this way, since the target encoded speech feature carries language content information, style information, and voice color information, and the source enhanced style feature and the target voice color feature are additionally added during decoding, the influence of the voice color of the source speaker on the style can be further reduced during speech decoding, and the independence and stability between the style feature of the source speaker and the voice color feature of the target speaker can be improved. In summary, the present application can improve the voice color consistency of speech synthesis and improve the naturalness of speech synthesis.
[0092] In step 101 of some embodiments, obtain the source speech data of the source speaker and the target speech data of the target speaker. Among them, the source speaker and the target speaker are different. The source speaker refers to a speaker with a certain specific dialect, accent, or language ability. The target speaker refers to a speaker with a certain specific dialect, accent, or language ability.
[0093] In the scenario of a multilingual learning tool in the fintech field, the source speaker is student S1 who has a first native language (such as Chinese), and the target speaker is student S2 who has a second native language (such as English). For another example, in the scenario of short video dubbing in the fintech field, the source speaker is dubbing staff A, and the target speaker is dubbing staff B. In the scenario of intelligent medical guidance in the medical and health field, the source speaker is medical staff, and the target speaker is a robot.
[0094] Source speech data refers to the speech emitted by the source speaker according to specific language content. Target speech data refers to the speech emitted by the target speaker according to random language content.
[0095] For example, in the scenario of short video dubbing, for a short video with the content of how to brush teeth correctly, the language content in the source speech data can be set as "This is a short video about how to brush teeth correctly", and the language content in the target speech data is "I am dubbing staff XX". For another example, in the scenario of intelligent medical guidance, for the treatment guidance speech used to guide patients to seek medical treatment, the language content in the source speech data can be set as "Please ask patient XX to go to consulting room 1 for treatment", and the language content in the target speech data is "I am intelligent assistant XX". It should be noted that the main purpose of this application is to extract the timbre characteristics of the target speaker from the target speech data, so there is no need to pay attention to whether the language content in the target speech data is related to the language content in the source speech data.
[0096] In an example, when the speech synthesis method is applied to a terminal, the source speech data and the target speech data can be obtained by recording, Bluetooth transmission, wired transmission, or downloading. When the source speech data and the target speech data are obtained by recording, the terminal is correspondingly configured with a microphone, and audio acquisition is performed through the microphone to achieve the recording of the source speech data and the target speech data. When the speech synthesis method is applied to the server side, the source speech data and the target speech data can be uploaded to the server side by the terminal, or can be downloaded by the server side from other server sides or databases.
[0097] In step 102 of some embodiments, timbre extraction is performed on the target speech data to obtain target timbre characteristics. Timbre characteristics refer to the unique attributes of sound, enabling different sounds to be distinguishable even when the pitch is the same or similar. The target timbre characteristics are used to indicate the voice or identity of the target speaker.
[0098] In an example, a pre-trained timbre extractor can be used to perform timbre extraction on the target speech data to obtain target timbre characteristics. A timbre extractor is a deep neural network model used to analyze and extract timbre characteristics from speech data, such as a convolutional neural network, a recurrent neural network, etc.
[0099] In step 103 of some embodiments, feature extraction is performed on the source speech data to obtain source language content features. In speech processing, features representing text content or language content are usually referred to as language content features, which emphasize more on the specific language information and its meaning expressed by the speech. The source language content features are used to indicate the features corresponding to the language content in the source speech data. For example, if the language content in the source speech data combined with the above text is set to "This is a short video about how to brush teeth correctly", then the source language content features specifically refer to the features corresponding to the language content of "This is a short video about how to brush teeth correctly".
[0100] In one example, the source speech data can be subjected to feature extraction through a pre-trained content feature extractor to obtain source language content features. The content feature extractor is a deep neural network model for analyzing and extracting language content features in speech data, such as convolutional neural networks, recurrent neural networks, etc.
[0101] In step 104 of some embodiments, style extraction is performed on the source speech data to obtain source initial style features. Style features refer to the features related to language style (such as intonation, speech rate, and emotion) in speech.
[0102] In one example, the source speech data can be subjected to style extraction through a pre-trained style extractor to obtain source initial style features. The style extractor is a deep neural network model for analyzing and extracting language style features in speech data, such as convolutional neural networks, recurrent neural networks, etc.
[0103] Before step 105, it is necessary to pre-train the style encoder. In one embodiment, referring to Figure 2 , the speech synthesis method may further include:
[0104] Step 201, obtain the sample source speech data of the sample source speaker and the sample target speech data of the sample target speaker, extract the sample source style features and sample source language content features from the sample source speech data, and extract the sample target timbre features from the sample target speech data; wherein, the sample source speaker and the sample target speaker are different;
[0105] Step 202, perform style encoding on the sample source style features through a preset initial style encoder to obtain sample enhanced style features;
[0106] Step 203, perform attention processing on the sample source language content features through the initial style encoder to obtain a sample language content attention vector, perform attention processing on the sample enhanced style features through the initial style encoder to obtain a sample style attention vector, and perform attention processing on the sample target timbre features through the initial style encoder to obtain a sample timbre attention vector;
[0107] Step 204: Enhance the style of the sample style attention vector and the preset first encoding parameter through the initial style encoder to obtain the sample style-enhanced attention vector, and enhance the timbre of the sample timbre attention vector and the preset second encoding parameter through the initial style encoder to obtain the sample timbre-enhanced attention vector;
[0108] Step 205: Perform vector fusion on the sample language content attention vector, the sample style-enhanced attention vector, and the sample timbre-enhanced attention vector through the initial style encoder to obtain the sample speech coding feature;
[0109] Step 206: Decode the sample speech coding feature to obtain the sample synthesized speech data; wherein, the timbre of the sample synthesized speech data comes from the sample target speaker, and the style and language content of the sample synthesized speech come from the sample source speaker;
[0110] Step 207: Calculate the loss according to the sample synthesized speech data and the preset labeled speech data to obtain the target loss function;
[0111] Step 208: Update the first encoding parameter and the second encoding parameter according to the target loss function, and adjust the parameters of the initial style encoder according to the target loss function to obtain the style encoder.
[0112] In step 201, the introduction of the sample source speaker, the sample source speech data, the sample target speaker, and the sample target speech data is the same as the explanations of steps 101 to 104 in the above text, except that one is applied to the training stage and the other is applied to the actual application stage.
[0113] In step 202, the initial style encoder is composed of transformer blocks. The attention mechanism used in each block is the multi-attention mechanism, and the accurate modeling of style and timbre is realized through the adaptive gating mechanism.
[0114] In one embodiment, step 202 may include: randomly dividing the sample source style feature to obtain at least two sample source style sub-features; performing attention processing on each of the at least two sample source style sub-features through the initial style encoder to obtain the sample style attention sub-vectors; performing parameter mapping according to the elements in the sample source style sub-features to obtain the sample attention parameters; performing non-linear activation on the sample attention parameters through a predetermined activation function to obtain the sample attention sub-weights; performing feature fusion on the sample attention sub-weights and the sample style attention sub-vectors of each sample source style sub-feature to obtain the sample enhanced style feature. In this way, the style information in the sample source style feature can be effectively highlighted and the timbre information can be ignored, and the sample enhanced style feature with richer style information can be obtained.
[0115] In step 203, the process of performing attention processing on the sample source language content features is shown in the following formula:
[0116] Among them, L1 refers to the sample language content attention vector, softmax is the activation function, c q refers to the query vector corresponding to the sample source language content features, c k refers to the key vector corresponding to the sample source language content features, c v refers to the value vector corresponding to the sample source language content features, and d refers to the dimension of the sample source language content features.
[0117] The process of performing attention processing on the sample enhanced style features is shown in the following formula:
[0118] Among them, L2 refers to the sample style attention vector, softmax is the activation function, c q refers to the query vector corresponding to the sample source language content features, s k refers to the key vector corresponding to the sample enhanced style features, s v refers to the value vector corresponding to the sample enhanced style features, and d refers to the dimension of the sample enhanced style features.
[0119] The process of performing attention processing on the sample target tone features is shown in the following formula:
[0120] Among them, L3 refers to the sample tone attention vector, softmax is the activation function, c q refers to the query vector corresponding to the sample source language content features, p k refers to the key vector corresponding to the sample target tone features, p v refers to the value vector corresponding to the sample target tone features, and d refers to the dimension of the sample enhanced style features.
[0121] In step 204, the process of performing style enhancement on the sample style attention vector and the preset first coding parameter is shown in the following formula:
[0122] Among them, L1 refers to the sample style enhancement attention vector, tanh is the hyperbolic tangent function, and a refers to the trainable first coding parameter.
[0123] The process of performing tone enhancement on the sample tone attention vector and the preset second coding parameter is shown in the following formula:
[0124] Among them, L2 refers to the sample timbre enhancement attention vector, tanh is the hyperbolic tangent function, and b refers to the trainable second coding parameter.
[0125] In step 205, the process of vector fusion can be expressed as: L = L1 + L1 + L2, where L refers to the sample speech coding feature.
[0126] In step 206, the sample speech coding feature is decoded by a preset speech decoder to obtain the sample synthesized speech data. The timbre of the sample synthesized speech data comes from the sample target speaker, and the style and language content of the sample synthesized speech come from the sample source speaker. An efficient module is used in the speech decoder part. For example, a conditional flow matching module is used to construct the decoder. Compared with the traditional diffusion model, it can significantly reduce the sampling steps, improve the inference speed, and at the same time ensure the clarity and stability of the generated speech. The speech decoder in this embodiment adopts a non-autoregressive framework, abandoning the information compression and delay problems in the autoregressive generation process, not only improving the model inference efficiency, but also maintaining the consistency in the long speech sequence.
[0127] In step 207, the target loss function is used to indicate the gap between the sample synthesized speech data and the labeled speech data, thereby reflecting the error of the initial style encoder in speech synthesis.
[0128] In one embodiment, referring to Figure 3 , step 207 may include:
[0129] Step 301, calculate the speech similarity between the sample synthesized speech data and the preset labeled speech data to obtain the speech loss function;
[0130] Step 302, extract the style of the labeled speech data to obtain the labeled style feature;
[0131] Step 303, calculate the style similarity between the sample speech coding feature and the labeled style feature to obtain the first style loss function;
[0132] Step 304, calculate the style similarity between the sample enhanced style feature and the labeled style feature to obtain the second style loss function;
[0133] Step 305, perform loss fusion according to the speech loss function, the first style loss function, and the second style loss function to obtain the target loss function.
[0134] In step 301, the speech loss function can be calculated using a mean square error function, a cosine similarity function, etc.
[0135] In step 302, the source initial style features can be extracted from the source speech data through a pre-trained style extractor. The style extractor is a deep neural network model used to analyze and extract the language style features in speech data, such as convolutional neural network, recurrent neural network, etc.
[0136] In steps 303 to 304, the first style loss function and the second style loss function can be calculated using mean square error function, cosine similarity function, etc.
[0137] In step 305, the speech loss function, the first style loss function and the second style loss function can be weighted and summed to obtain the target loss function.
[0138] The benefits of the embodiments of the above steps 301 to 305 are that, on the basis of calculating the speech loss function, the introduction of the first style loss function and the second style loss function greatly improves the control ability of the style encoder over the language style of the source speaker, and further reduces the influence of the timbre of the sample source speaker on the language style, thereby contributing to improving the naturalness of speech synthesis.
[0139] In step 208, the first coding parameter and the second coding parameter are updated according to the target loss function, and the parameters of the initial style encoder are adjusted according to the target loss function to obtain the style encoder. Specifically, the gradients of the first coding parameter, the second coding parameter and the parameters of the initial style encoder are calculated according to the target loss function; according to the calculated gradients and the learning rate (a hyperparameter that controls the adjustment step), the first coding parameter, the second coding parameter and the parameters of the initial style encoder are updated. Repeat steps 202 to 208 until the stop condition is met (such as the loss is lower than a certain threshold, the maximum number of iterations is reached or the loss function converges), and then the style encoder is obtained.
[0140] In step 105 of some embodiments, the source initial style features are style-encoded through a preset style encoder to obtain source enhanced style features.
[0141] In one embodiment, referring to Figure 4 , step 105 may include:
[0142] Step 401, randomly dividing the source initial style features to obtain at least two source initial style sub-features;
[0143] Step 402, performing attention processing on each of the at least two source initial style sub-features through the style encoder to obtain style attention sub-vectors;
[0144] Step 403, generating attention sub-weights according to each source initial style sub-feature;
[0145] Step 404: Perform feature fusion on the attention sub-weights of each source initial style sub-feature and the style attention sub-vector to obtain the source enhanced style feature.
[0146] In step 401, for example, if the number of features of the source initial style feature is M, after random partitioning, the source initial style sub-feature T1 with the number of features M1 and the source initial style sub-feature T2 with the number of features M2 can be obtained, where M1 + M2 = M.
[0147] In step 402, the process of performing attention processing on each source initial style sub-feature may include: performing non-linear processing on the source initial style sub-feature to obtain the source initial style query sub-vector, the source initial style key sub-vector, and the source initial style value sub-vector; performing attention calculation on the source initial style query sub-vector, the source initial style key sub-vector, and the source initial style value sub-vector to obtain the style attention sub-vector. Among them, the attention calculation adopts the self-attention mechanism, and the specific calculation process will not be elaborated.
[0148] In step 403, the attention sub-weight is related to the elements in the source initial style sub-feature.
[0149] In one embodiment, step 403 may include: performing parameter mapping according to the elements in the source initial style sub-feature to obtain the attention parameter; performing non-linear activation on the attention parameter through a predetermined activation function to obtain the attention sub-weight. For example, performing weighted summation on all elements in the source initial style sub-feature to obtain the attention parameter. The predetermined activation function selects the hyperbolic tangent function, and performing non-linear activation on the attention parameter through the hyperbolic tangent function can obtain the attention sub-weight.
[0150] Finally, in step 404, the attention sub-weights of each source initial style sub-feature and the style attention sub-vector can be feature-fused by weighted summation to obtain the source enhanced style feature.
[0151] The advantages of the embodiments of the above steps 401 to 404 are that using the pre-trained style encoder to highlight the style information in the source initial style feature can reduce the influence of the source speaker's voice tone on the style feature, thereby improving the naturalness of speech synthesis.
[0152] In step 106 of some embodiments, the source enhanced style feature, the source language content feature, and the target voice tone feature are subjected to attention encoding through the style encoder to obtain the target encoded speech feature. The target encoded speech feature contains the language content information and language style information of the source speaker, and also contains the voice tone information of the target speaker.
[0153] In one embodiment, referring to Figure 5 , step 106 may include:
[0154] Step 501: Perform attention processing on the source language content features through a style encoder to obtain a language content attention vector;
[0155] Step 502: Perform attention processing on the source enhanced style features through a style encoder to obtain a style attention vector;
[0156] Step 503: Perform attention processing on the target timbre features through a style encoder to obtain a timbre attention vector;
[0157] Step 504: Perform style enhancement on the style attention vector and a preset first encoding parameter through a style encoder to obtain a style-enhanced attention vector;
[0158] Step 505: Perform timbre enhancement on the timbre attention vector and a preset second encoding parameter through a style encoder to obtain a timbre-enhanced attention vector;
[0159] Step 506: Perform vector fusion on the language content attention vector, the style-enhanced attention vector, and the timbre-enhanced attention vector through a style encoder to obtain the target encoded speech features.
[0160] The specific processes of Steps 501 to 506 are basically the same as those of Steps 203 to 205 in the above text, so they will not be elaborated here. Among them, Step 504 may include: performing non-linear activation on the first encoding parameter through a predetermined activation function to obtain a first weight; multiplying each element of the style attention vector by the first weight to obtain the style-enhanced attention vector. Step 505 may include: performing non-linear activation on the second encoding parameter through a predetermined activation function to obtain a second weight; multiplying each element of the timbre attention vector by the second weight to obtain the timbre-enhanced attention vector.
[0161] The benefits of the above embodiments of Steps 501 to 506 are that the attention mechanism of the style encoder can enhance the independence and stability of the timbre processing of the target speaker and the style processing of the source speaker.
[0162] In Step 107 of some embodiments, the source enhanced style features, the target encoded speech features, and the target timbre features may be decoded through a speech decoder to obtain the target synthesized speech data. Among them, the timbre of the target synthesized speech data comes from the target speaker, and the style and language content of the target synthesized speech come from the source speaker. The introduction of the speech decoder is described in the description of Step 206 in the above text, so it will not be elaborated here.
[0163] Based on the above embodiments, the present application can at least achieve the following beneficial effects: (1) Improved sound quality and similarity: Through efficient feature separation and multiple attention mechanisms, the naturalness of the generated speech and the similarity of timbre can be significantly improved, especially the timbre and style conversion effects in the zero-shot scenario are more prominent. (2) Efficient inference: The model proposed in the present application uses a non-autoregressive framework, which can significantly improve the inference speed and support real-time speech conversion requirements. (3) Flexible style control: The independent modeling of timbre and style enables users to freely combine the timbres and styles of different speakers to meet diverse speech synthesis needs, such as emotional speech generation or multi-role audio production. (4) Wide application prospects: The technical solution proposed in the present application can be applied to fields such as audiobooks, short video dubbing, and multilingual learning tools, effectively improving the efficiency of content creators and the user experience.
[0164] Please refer to Figure 6 , the embodiments of the present application also provide a speech synthesis device, which can implement the above speech synthesis method. Figure 6 FIG. is a block diagram of the module structure of the speech synthesis device provided by the embodiment of the present application. The device includes:
[0165] A speech acquisition module 601, configured to acquire the source speech data of the source speaker and the target speech data of the target.
[0166] A timbre extraction module 602, configured to extract the timbre of the target speech data to obtain the target timbre feature.
[0167] A feature extraction module 603, configured to extract features from the source speech data to obtain the source language content feature.
[0168] A style extraction module 604, configured to extract the style from the source speech data to obtain the source initial style feature.
[0169] A style encoding module 605, configured to perform style encoding on the source initial style feature through a preset style encoder to obtain the source enhanced style feature.
[0170] An attention encoding module 606, configured to perform attention encoding on the source enhanced style feature, the source language content feature, and the target timbre feature through the style encoder to obtain the target encoded speech feature.
[0171] A language decoding module 607, configured to perform speech decoding on the source enhanced style feature, the target encoded speech feature, and the target timbre feature to obtain the target synthesized speech data; wherein, the timbre of the target synthesized speech data comes from the target speaker, and the style and language content of the target synthesized speech come from the source speaker.
[0172] In an embodiment, the speech synthesis device further includes: a model training module, configured to pre-train the style encoder.
[0173] It should be noted that the specific implementation manner of this voice synthesis device is basically the same as the specific embodiments of the above voice synthesis method, and will not be elaborated here.
[0174] The embodiments of the present application also provide an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the program is executed by the processor, the above voice synthesis method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0175] Please refer to Figure 7 , Figure 7 which shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0176] A processor 701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0177] A memory 702, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702, and the processor 701 is called to execute the voice synthesis method of the embodiments of the present application;
[0178] An input / output interface 703, which is used to implement information input and output;
[0179] A communication interface 704, which is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0180] A bus 705, which transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);
[0181] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.
[0182] The embodiments of the present application also provide a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above voice synthesis method.
[0183] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0185] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0188] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0189] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0190] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0191] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0193] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0194] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A speech synthesis method, characterized in that: The method comprises: Acquire source speech data of a source speaker and target speech data of a target speaker; wherein the source speaker and the target speaker are different; Performing timbre extraction on the target speech data to obtain target timbre features; Extracting features from the source speech data to obtain source language content features; Performing style extraction on the source speech data to obtain source initial style features; Performing style encoding on the source initial style feature through a preset style encoder to obtain a source enhanced style feature; Attention encoding is performed on the source enhanced style feature, the source language content feature and the target timbre feature by the style encoder to obtain a target encoded speech feature; The source enhanced style features, the target encoded speech features and the target timbre features are speech decoded to obtain target synthesized speech data; wherein the timbre of the target synthesized speech data is derived from the target speaker, and the style and language content of the target synthesized speech are derived from the source speaker.
2. The method according to claim 1, characterized in that The method of performing attention encoding on the source enhanced style feature, the source language content feature and the target timbre feature by the style encoder to obtain a target encoded speech feature includes: Performing attention processing on the source language content features by the style encoder to obtain a language content attention vector; Performing attention processing on the source enhanced style feature through the style encoder to obtain a style attention vector; Performing attention processing on the target timbre feature by the style encoder to obtain a timbre attention vector; Performing style enhancement on the style attention vector and the preset first encoding parameter by the style encoder to obtain a style enhanced attention vector; Performing timbre enhancement on the timbre attention vector and the preset second encoding parameter by the style encoder to obtain a timbre enhanced attention vector; The style encoder performs vector fusion on the language content attention vector, the style enhancement attention vector and the timbre enhancement attention vector to obtain the target encoded speech feature.
3. The method according to claim 2, characterized in that The step of performing vector enhancement on the style attention vector and the preset first encoding parameter by the style encoder to obtain the style enhanced attention vector includes: Performing nonlinear activation on the first encoding parameter by using a predetermined activation function to obtain a first weight; Multiplying each element of the style attention vector by the first weight to obtain a style enhanced attention vector; The step of performing vector enhancement on the timbre attention vector and the preset second encoding parameter by the style encoder to obtain the timbre enhanced attention vector includes: Performing nonlinear activation on the second encoding parameter by using the predetermined activation function to obtain a second weight; Multiply each element of the timbre attention vector by the second weight to obtain a timbre enhancement attention vector.
4. The method according to claim 3, characterized in that Before performing attention encoding on the source enhanced style feature, the source language content feature and the target timbre feature by the style encoder to obtain the target encoded speech feature, the method further includes: Pre-training the style encoder includes: Acquire sample source speech data of a sample source speaker and sample target speech data of a sample target speaker, extract sample source style features and sample source language content features from the sample source speech data, and extract sample target timbre features from the sample target speech data; wherein the sample source speaker and the sample target speaker are different; Performing style encoding on the sample source style feature through a preset initial style encoder to obtain a sample enhanced style feature; Performing attention processing on the sample source language content features through the initial style encoder to obtain a sample language content attention vector; Performing attention processing on the sample enhanced style feature through the initial style encoder to obtain a sample style attention vector; Performing attention processing on the sample target timbre feature through the initial style encoder to obtain a sample timbre attention vector; Performing style enhancement on the sample style attention vector and a preset first encoding parameter by the initial style encoder to obtain a sample style enhanced attention vector; Performing timbre enhancement on the sample timbre attention vector and a preset second encoding parameter by using the initial style encoder to obtain a sample timbre enhanced attention vector; The initial style encoder performs vector fusion according to the sample language content attention vector, the sample style enhancement attention vector and the sample timbre enhancement attention vector to obtain a sample speech coding feature; Decoding the sample speech coding features to obtain sample synthesized speech data; wherein the timbre of the sample synthesized speech data is derived from the sample target speaker, and the style and language content of the sample synthesized speech are derived from the sample source speaker; Performing loss calculation based on the sample synthesized speech data and the preset labeled speech data to obtain a target loss function; The first encoding parameter and the second encoding parameter are updated according to the target loss function, and the parameters of the initial style encoder are adjusted according to the target loss function to obtain the style encoder.
5. The method according to claim 4, characterized in that The loss calculation is performed according to the sample synthesized speech data and the preset label speech data to obtain a target loss function, including: Calculating speech similarity between the sample synthesized speech data and the preset labeled speech data to obtain a speech loss function; Performing style extraction on the label voice data to obtain label style features; Calculating the style similarity between the sample speech coding feature and the label style feature to obtain a first style loss function; Calculating the style similarity between the sample enhanced style feature and the label style feature to obtain a second style loss function; Loss fusion is performed according to the speech loss function, the first style loss function and the second style loss function to obtain the target loss function.
6. The method according to any one of claims 1 to 5, characterized in that: The method of encoding the source initial style feature by a preset style encoder to obtain the source enhanced style feature includes: Randomly dividing the source initial style feature to obtain at least two source initial style sub-features; Performing attention processing on each of the at least two source initial style sub-features by the style encoder to obtain a style attention sub-vector; generating an attention sub-weight according to each of the source initial style sub-features; The source enhanced style feature is obtained by performing feature fusion according to the attention sub-weight of each of the source initial style sub-features and the style attention sub-vector.
7. The method according to claim 6, characterized in that Generating an attention sub-weight according to each of the source initial style sub-features includes: Perform parameter mapping according to the elements in the source initial style sub-feature to obtain attention parameters; The attention parameter is nonlinearly activated by a predetermined activation function to obtain the attention sub-weight.
8. A speech synthesis device, characterized in that: The device comprises: A speech acquisition module, used to acquire source speech data of a source speaker and target speech data of a target speaker; A timbre extraction module, used to extract the timbre of the target speech data to obtain target timbre features; A feature extraction module, used to extract features from the source speech data to obtain source language content features; A style extraction module, used to extract the style of the source speech data to obtain source initial style features; A style encoding module, used for performing style encoding on the source initial style feature through a preset style encoder to obtain a source enhanced style feature; An attention encoding module, configured to perform attention encoding on the source enhanced style feature, the source language content feature and the target timbre feature through the style encoder to obtain a target encoded speech feature; A language decoding module is used to perform speech decoding on the source enhanced style features, the target encoded speech features and the target timbre features to obtain target synthesized speech data; wherein the timbre of the target synthesized speech data comes from the target speaker, and the style and language content of the target synthesized speech come from the source speaker.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the speech synthesis method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 7 is implemented.