Speech synthesis method and device based on latent diffusion model, server and medium
Patent Information
- Application Number
- CN202410854549.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-06-27
AI Technical Summary
现有的语音合成方法合成的语音不够自然,质量较差,用户体验不好。
采用基于潜在扩散模型的语音合成方法,通过获取目标文本,确定语音情感特征和目标时长信息,利用神经网络音频编码器、残差向量量化器、潜在扩散模型和神经网络音频解码器进行编码、量化、融合和解码处理,生成自然度高、质量好的合成语音。
提高了合成语音的自然度和质量,增强了用户体验。
Smart Images

Figure CN118629390B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech synthesis, and in particular to a speech synthesis method, apparatus, server and storage medium based on a latent diffusion model. Background Technology
[0002] Text-to-Speech (TTS) technology is widely used in various fields. For example, in the financial or insurance sector, text can be converted into speech suitable for financial or insurance customer service scenarios, thereby helping financial institutions or insurance companies provide services such as intelligent outbound calling, intelligent response, and intelligent debt collection, improving customer satisfaction and operational efficiency. However, current speech synthesis methods still suffer from problems such as unnatural synthesized speech, poor quality, and a poor user experience. Summary of the Invention
[0003] This application provides a speech synthesis method, apparatus, server, and storage medium based on a latent diffusion model, aiming to improve the naturalness and quality of synthesized speech.
[0004] In a first aspect, embodiments of this application provide a speech synthesis method based on a latent diffusion model, comprising:
[0005] Obtain the target text to be synthesized;
[0006] Based on the target text, determine the emotional features and target duration information of the synthesized speech;
[0007] The target text is encoded using a neural network audio encoder in a preset target speech synthesis model to obtain a first hidden vector.
[0008] The first latent vector is quantized and compressed using the residual vector quantizer in the target speech synthesis model to obtain the second latent vector.
[0009] The vector representations of the speech emotion features, the vector representations of the target duration information, and the second latent vector are fused to obtain a fused latent vector;
[0010] The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model.
[0011] The target latent vector is decoded using the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
[0012] Secondly, embodiments of this application also provide a speech synthesis device based on a latent diffusion model, the speech synthesis device comprising:
[0013] The acquisition module is used to acquire the target text to be synthesized;
[0014] The determination module is used to determine the emotional features and target duration information of the synthesized speech based on the target text.
[0015] The encoding module is used to encode the target text using a neural network audio encoder in a preset target speech synthesis model to obtain a first latent vector;
[0016] The compression module is used to quantize and compress the first latent vector through the residual vector quantizer in the target speech synthesis model to obtain the second latent vector.
[0017] The fusion module is used to fuse the vector representation of the speech emotion features, the vector representation of the target duration information, and the second latent vector to obtain a fused latent vector;
[0018] The diffusion module is used to perform reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model to obtain the target latent vector;
[0019] The decoding module is used to decode the target latent vector through the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
[0020] Thirdly, embodiments of this application also provide a server, the server including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the speech synthesis method based on the latent diffusion model as described in the first aspect.
[0021] Fourthly, embodiments of this application also provide a storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the speech synthesis method based on the latent diffusion model as described in the first aspect.
[0022] This application provides a speech synthesis method, apparatus, server, and storage medium based on a latent diffusion model. The speech synthesis method, using the target text to be synthesized, can predict the emotional features and target duration information of the synthesized speech. It fuses the vector representations of the emotional features, the vector representations of the target duration information, and the latent vectors of the target text obtained through quantization and compression processing by a residual vector quantizer. This results in the fused latent vectors describing not only the semantics of the synthesized speech but also its emotional features and duration. By decoding the fused latent vectors using a neural network audio decoder, highly natural and high-quality synthesized speech can be obtained, effectively improving the naturalness and quality of the synthesized speech. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of a speech synthesis method provided in an embodiment of this application;
[0025] Figure 2 This is a schematic diagram of the network structure of the target speech synthesis model in the embodiments of this application;
[0026] Figure 3 This is a flowchart illustrating another speech synthesis method provided in an embodiment of this application;
[0027] Figure 4 This is a schematic block diagram of a speech synthesis device provided in an embodiment of this application;
[0028] Figure 5 This is a schematic block diagram of the structure of a server provided in an embodiment of this application.
[0029] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and the accompanying drawings. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0032] Speech synthesis technology is widely used in various fields. For example, in the financial or insurance sector, text can be converted into speech suitable for financial or insurance customer service scenarios, thereby helping financial institutions or insurance companies provide services such as intelligent outbound calling, intelligent response, and intelligent debt collection, improving customer satisfaction and operational efficiency. However, current speech synthesis methods still suffer from problems such as unnatural synthesized speech, poor quality, and a poor user experience.
[0033] To address the aforementioned issues, this application provides a speech synthesis method, apparatus, server, and storage medium based on a latent diffusion model. This speech synthesis method, using the target text to be synthesized, can predict the emotional features and target duration information of the synthesized speech. It fuses the vector representations of the emotional features, the vector representation of the target duration information, and the latent vectors of the target text obtained through quantization and compression processing using a residual vector quantizer. This results in latent vectors that describe not only the semantics of the synthesized speech but also its emotional features and duration. By decoding these latent vectors using a neural network audio decoder, highly natural and high-quality synthesized speech can be obtained, effectively improving the naturalness and quality of the synthesized speech.
[0034] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. This speech synthesis method can be applied to a server. This server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0035] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0036] Please see Figure 1 , Figure 1This is a flowchart illustrating the steps of a speech synthesis method based on a latent diffusion model provided in an embodiment of this application.
[0037] like Figure 1 As shown, the speech synthesis method includes steps S101 to S107.
[0038] Step S101: Obtain the target text to be synthesized.
[0039] In this embodiment, the target text to be synthesized can be pre-stored text suitable for a specific scenario, or text manually entered by the user. The specific scenario can be customer service scenarios in various fields, such as voice-based intelligent customer service scenarios in the financial or insurance sectors. In such scenarios, the target text to be synthesized is any text that needs to be read aloud to the user. As another example, in a smart reading scenario, the target text to be synthesized can be text scanned by a smart reading device.
[0040] Step S102: Determine the speech emotion features and target duration information of the synthesized speech based on the target text.
[0041] In this embodiment, the speech emotion features and target duration information of the synthesized speech corresponding to the target text can be predicted. The speech emotion features include at least one of prosodic features, cepstral features, phonological features, and zero-crossing peak amplitude features. The cepstral features include Linear Prediction Cepstrum Coefficients (LPCC) and Mel-Frequency Cepstrum Coefficients (MFCC), etc.
[0042] In some embodiments, prosodic features include fundamental frequency-related features, duration-related features, and energy-related features. Fundamental frequency-related features include fundamental frequency, mean fundamental frequency, fundamental frequency range, fundamental frequency rate of change, and fundamental frequency mean square deviation. Duration-related features include speech rate and short-time average zero-crossing rate. Energy-related features include short-time average energy, short-time energy rate of change, and short-time average amplitude.
[0043] In some embodiments, phonological features are used to describe the properties of the glottal excitation signal, including the speaker's tone, breathiness, vibrato, and choking, to measure speech purity, intelligibility, and recognizability. Specifically, phonological features include formant frequencies, bandwidth, frequency perturbations, amplitude perturbations, harmonic noise ratio, flashbacks, and glottal parameters.
[0044] In some embodiments, determining the speech emotion features and target duration information of the synthesized speech based on the target text may include: performing speech emotion feature prediction processing on the target text using a preset speech emotion feature prediction model to obtain the speech emotion features of the synthesized speech; and performing duration prediction processing on the target text using a preset duration prediction model to obtain the target duration information of the synthesized speech. The target duration information includes the duration of each phoneme in the synthesized speech corresponding to the target text.
[0045] In some embodiments, the speech emotion feature prediction model is pre-trained on a first neural network model using multiple first samples. The first samples include sample text and labeled speech emotion features, where the labeled speech emotion features include the speech emotion features of the speech segment corresponding to the sample text. The duration prediction model is pre-trained on a second neural network model using multiple second samples. The second samples include sample text and labeled duration information, where the labeled duration information includes the duration of each phoneme in the speech segment corresponding to the sample text. The duration of each phoneme in the speech segment corresponding to the sample text can be determined using a speech alignment tool (Montreal Forced Aligner, MFA).
[0046] In some embodiments, training a first neural network model based on multiple first samples may include: selecting one first sample from the multiple first samples as a training sample; inputting the sample text from the training sample into a preset first neural network model for speech emotion feature prediction processing to obtain predicted speech emotion features; determining a first model loss value based on the predicted speech emotion features and the speech emotion features labeled in the training sample; when the first model loss value is greater than a first set loss value, updating the model parameters of the first neural network model based on the first model loss value, and returning to the step of selecting one first sample from the multiple first samples as a training sample; when the first model loss value is less than or equal to the first set loss value, stopping the training of the first neural network model to obtain a speech emotion feature prediction model.
[0047] In some embodiments, training a second neural network model based on multiple second samples may include: selecting one second sample from the multiple second samples as a training sample; inputting the sample text from the training sample into a preset second neural network model for duration prediction processing to obtain predicted duration information; determining a second model loss value based on the predicted duration information and the duration information labeled in the training sample; when the second model loss value is greater than a second set loss value, updating the model parameters of the second neural network model based on the second model loss value, and returning to the step of selecting one second sample from the multiple second samples as a training sample; when the second model loss value is less than or equal to the second set loss value, stopping the training of the second neural network model to obtain a duration prediction model.
[0048] Step S103: Encode the target text using the neural network audio encoder in the preset target speech synthesis model to obtain the first hidden vector.
[0049] In this embodiment, as Figure 2 As shown, the target speech synthesis model 10 includes a neural network audio encoder 11, a residual vector quantization (RVQ) 12, a latent diffusion model (LDM) 13, and a neural network audio decoder 14. The neural network audio encoder 11 is connected to the residual vector quantization 12, the residual vector quantization 12 is connected to the latent diffusion model 13, and the latent diffusion model 13 is connected to the neural network audio decoder 14.
[0050] The neural network audio encoder 11 is an audio encoder implemented based on a neural network, and the neural network audio decoder 14 is an audio decoder implemented based on a neural network. For example, the neural network audio encoder 11 has the same or similar structure as the encoder in the SoundStream model, and the neural network audio decoder 14 has the same or similar structure as the decoder in the SoundStream model.
[0051] Step S104: The first latent vector is quantized and compressed using the residual vector quantizer in the target speech synthesis model to obtain the second latent vector.
[0052] In this embodiment, the residual vector quantizer can quantize and compress the first hidden vector output by the neural network audio encoder to a second hidden vector of a set byte length, which can effectively reduce the amount of computation.
[0053] Step S105: The vector representations of speech emotion features, target duration information, and the second latent vector are fused to obtain the fused latent vector.
[0054] In this embodiment, the vector representation of speech emotion features, the vector representation of target duration information, and the second latent vector can be fused based on the cross-attention mechanism to obtain a fused latent vector. This fused latent vector can not only describe the semantics of the synthesized speech, but also describe the emotional features and duration of the synthesized speech.
[0055] In some embodiments, fusing the vector representations of speech emotion features, the vector representation of target duration information, and the second latent vector to obtain a fused latent vector may include: concatenating the vector representations of speech emotion features and the vector representations of target duration information to obtain a concatenated latent vector; and fusing the concatenated latent vector with the second latent vector to obtain the fused latent vector. The fused latent vector can be obtained by fusing the concatenated latent vector with the second latent vector based on a cross-attention mechanism.
[0056] Step S106: Using the latent diffusion model in the target speech synthesis model, perform reverse diffusion processing on the fused latent vector to obtain the target latent vector.
[0057] In this embodiment, the latent diffusion model can perform reverse diffusion processing on the fused latent vector in the latent space, thereby gradually removing noise from the fused latent vector to recover the target latent vector.
[0058] In some embodiments, obtaining the target latent vector by performing inverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model may include: obtaining the prior conditions of the latent diffusion model, wherein the prior conditions include at least one of pitch, timbre, and euphemism; and performing inverse diffusion processing on the fused latent vector through the latent diffusion model and the prior conditions in the target speech synthesis model to obtain the target latent vector. This embodiment performs inverse diffusion processing on the fused latent vector through the latent diffusion model and the prior conditions, so that the recovered target latent vector can not only describe the semantic, emotional features, and duration of the synthesized speech, but also describe the pitch, timbre, and / or euphemism of the synthesized speech, making the subsequently synthesized speech more natural and of better quality, further improving the naturalness and quality of the synthesized speech.
[0059] In some embodiments, obtaining the prior conditions of the latent diffusion model may include: determining the scene label of the scene to which the target text belongs; querying the pre-stored mapping relationship between scene labels and prior conditions to obtain the prior conditions corresponding to the scene label of the scene to which the target text belongs, and determining the queried prior conditions as the prior conditions of the latent diffusion model. Different scene labels correspond to different prior conditions, and the mapping relationship between scene labels and prior conditions can be set based on actual conditions; this embodiment of the invention does not specifically limit this. This embodiment can adaptively match the corresponding prior conditions based on the scene label of the scene to which the target text belongs, so that the pitch, timbre, and / or euphemism of the subsequently generated synthesized speech can meet the needs of the scene to which the target text belongs, resulting in a better user experience.
[0060] In some embodiments, the process of performing reverse diffusion processing on the fused latent vector using the latent diffusion model and prior conditions in the target speech synthesis model to obtain the target latent vector may include: concatenating the vector representation corresponding to the prior conditions with the fused latent vector to obtain the concatenated fused latent vector; and performing reverse diffusion processing on the concatenated fused latent vector using the latent diffusion model in the target speech synthesis model to obtain the target latent vector.
[0061] Step S107: The target latent vector is decoded by the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
[0062] In this embodiment, the target latent vector is decoded using a neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text. Post-processing can then be performed on the synthesized speech to improve its clarity and stability. For example, denoising, enhancement, and / or compression can be applied to the synthesized speech corresponding to the target text to obtain the target synthesized speech.
[0063] In some embodiments, such as Figure 3 As shown, before step S101, the following steps are also included:
[0064] Step S108: Obtain the training sample set and the speech synthesis model to be trained.
[0065] In this embodiment, the training samples in the training sample set include training text, the transcribed speech corresponding to the training text, and the first training latent vector of the transcribed speech. The first training latent vector is obtained by concatenating the vector representation of the speech emotion features of the transcribed speech and the vector representation of the duration of the transcribed speech. The network structure of the speech synthesis model to be trained is the same as that of the target speech synthesis model, but the parameters of the speech synthesis model to be trained are different from those of the target speech synthesis model.
[0066] In some embodiments, the training sample set includes multiple training sample subsets, each corresponding to a different speech emotion type. Each training sample subset includes multiple training samples, which include training text, transcribed speech corresponding to the training text, and a first training latent vector of the transcribed speech. This embodiment trains the speech synthesis model using training sample subsets containing different speech emotion types, enabling the trained target speech synthesis model to accurately synthesize synthesized speech of different speech emotion types as needed, thereby improving the emotional diversity of synthesized speech.
[0067] Step S109: Select a training sample from the training sample set.
[0068] In this embodiment, a training sample can be randomly selected from the training sample set, or a training sample can be selected sequentially from the training sample set according to a set order. This embodiment of the invention does not make specific limitations on this.
[0069] Step S110: The training text in the selected training samples is encoded by the neural network audio encoder in the speech synthesis model to obtain the third hidden vector. The third hidden vector is quantized and compressed by the residual vector quantizer in the speech synthesis model to obtain the fourth hidden vector. The first training hidden vector and the fourth hidden vector in the selected training samples are fused to obtain the fifth hidden vector.
[0070] In this embodiment, fusing the first and fourth hidden vectors from the selected training samples to obtain the fifth hidden vector may include: concatenating the first and fourth hidden vectors from the selected training samples to obtain the fifth hidden vector. Alternatively, it may involve fusing the first and fourth hidden vectors from the selected training samples based on a cross-attention mechanism to obtain the fifth hidden vector.
[0071] Step S111: The fifth hidden vector is reverse-diffused using the latent diffusion model in the speech synthesis model to obtain the sixth hidden vector. The sixth hidden vector is then decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
[0072] In this embodiment, since the fifth hidden vector is obtained by fusing the first training hidden vector and the fourth hidden vector in the training samples, the latent diffusion model performs reverse diffusion processing on the fifth hidden vector to obtain a sixth hidden vector that can not only describe the semantics of the synthesized speech, but also describe the emotional features and duration information of the synthesized speech. This allows the neural network audio decoder to obtain more natural synthesized speech after decoding the sixth hidden vector.
[0073] In some embodiments, the training samples further include a second training latent vector of the transcribed speech, which is obtained by encoding the transcribed speech. Decoding the sixth latent vector using a neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text may include: fusing the second and sixth latent vectors from the selected training samples to obtain a seventh latent vector; and decoding the seventh latent vector using a neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text. This embodiment fuses the vector representation of the transcribed speech with the latent vector obtained through inverse diffusion processing using a latent diffusion model, making the fused latent vector more accurate and facilitating the subsequent output of higher-quality and more natural synthesized speech.
[0074] Step S112: Determine the model loss value based on the transcribed speech and the synthesized speech corresponding to the training text in the selected training samples.
[0075] In this embodiment, the difference between the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text can be determined, and the model loss value can be determined based on this difference.
[0076] Step S113: When the model loss value is greater than the preset loss value, update the parameters of the speech synthesis model.
[0077] In this embodiment, the preset loss value can be set based on actual conditions, and this embodiment of the invention does not impose specific limitations on it. Specifically, after updating the parameters of the speech synthesis model, the process returns to step S109 until the model loss value is less than or equal to the preset loss value or each training sample in the training sample set is selected once.
[0078] In some embodiments, when the model loss value is greater than a preset loss value, the parameters of the residual vector quantizer and the latent diffusion model in the speech synthesis model can be updated based on the model loss value and the backpropagation algorithm, without updating the parameters of the neural network audio encoder and the neural network audio decoder. Alternatively, the parameters of the neural network audio encoder, the residual vector quantizer, the latent diffusion model, and the neural network audio decoder in the speech synthesis model can be updated together.
[0079] Step S114: When the model loss value is less than or equal to the preset loss value, stop training the speech synthesis model to obtain the target speech synthesis model.
[0080] This embodiment iteratively trains the speech synthesis model using training samples containing training text, the corresponding transcribed speech, and the first training latent vector of the transcribed speech. This enables the final converged target speech synthesis model to synthesize more natural and higher-quality synthesized speech, effectively improving the speech synthesis effect of the target speech synthesis model.
[0081] Please see Figure 4 , Figure 4 This is a schematic block diagram of a speech synthesis device based on a latent diffusion model provided in an embodiment of this application.
[0082] like Figure 4 As shown, the speech synthesis device 100 based on the latent diffusion model includes:
[0083] Module 110 is used to acquire the target text to be synthesized;
[0084] The determining module 120 is used to determine the speech emotion features and target duration information of the synthesized speech based on the target text;
[0085] The encoding module 130 is used to encode the target text using a neural network audio encoder in a preset target speech synthesis model to obtain a first latent vector;
[0086] The quantization compression module 140 is used to quantize and compress the first latent vector through the residual vector quantizer in the target speech synthesis model to obtain the second latent vector.
[0087] The fusion module 150 is used to fuse the vector representation of the speech emotion features, the vector representation of the target duration information, and the second latent vector to obtain a fused latent vector;
[0088] The diffusion module 160 is used to perform reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model to obtain the target latent vector;
[0089] The decoding module 170 is used to decode the target latent vector through the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
[0090] In some embodiments, the diffusion module 160 is further configured to:
[0091] Obtain the prior conditions of the potential diffusion model, wherein the prior conditions include at least one of pitch, timbre, and euphemism;
[0092] The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector using the latent diffusion model in the target speech synthesis model and the prior conditions.
[0093] In some embodiments, the fusion module 150 is further configured to:
[0094] The vector representation of the speech emotion feature and the vector representation of the target duration information are concatenated to obtain the concatenated latent vector;
[0095] The concatenated latent vector is fused with the second latent vector to obtain the fused latent vector.
[0096] In some embodiments, the determining module 120 includes:
[0097] The first prediction unit is used to perform speech emotion feature prediction processing on the target text through a preset speech emotion feature prediction model to obtain the speech emotion features of the synthesized speech. The speech emotion feature prediction model is obtained by training a first neural network model in advance based on multiple first samples. The first samples include sample text and labeled speech emotion features.
[0098] The second prediction unit is used to perform duration prediction processing on the target text through a preset duration prediction model to obtain the target duration information of the synthesized speech. The duration prediction model is obtained by training a second neural network model in advance based on multiple second samples. The second samples include sample text and labeled duration information.
[0099] In some embodiments, the speech synthesis device 100 based on a latent diffusion model further includes a model training module, the model training module being configured to:
[0100] Obtain a training sample set and a speech synthesis model to be trained. The training samples in the training sample set include training text, transcribed speech corresponding to the training text, and a first training latent vector of the transcribed speech. The first training latent vector is obtained by concatenating the vector representation of the speech emotion features of the transcribed speech and the vector representation of the duration of the transcribed speech.
[0101] Select a training sample from the training sample set;
[0102] The neural network audio encoder in the speech synthesis model encodes the training text in the selected training samples to obtain the third hidden vector.
[0103] The third hidden vector is quantized and compressed using the residual vector quantizer in the speech synthesis model to obtain the fourth hidden vector.
[0104] The first training hidden vector and the fourth hidden vector in the selected training samples are fused to obtain the fifth hidden vector;
[0105] The fifth hidden vector is reverse-diffusioned using the latent diffusion model in the speech synthesis model to obtain the sixth hidden vector.
[0106] The sixth hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
[0107] The model loss value is determined based on the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text.
[0108] When the model loss value is greater than the preset loss value, the parameters of the speech synthesis model are updated, and the process returns to the step of selecting a training sample from the training sample set.
[0109] When the model loss value is less than or equal to a preset loss value, training of the speech synthesis model is stopped, and the target speech synthesis model is obtained.
[0110] In some embodiments, the training samples further include a second training latent vector of the transcribed speech, the second training latent vector being obtained by encoding the transcribed speech, and the model training module is further configured to:
[0111] The second training hidden vector and the sixth hidden vector in the selected training samples are fused to obtain the seventh hidden vector;
[0112] The seventh hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
[0113] In some embodiments, the model training module is further configured to:
[0114] The difference between the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text is determined, and the model loss value is determined based on the difference.
[0115] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the above-described device and its modules and units can be referred to the corresponding processes in the aforementioned embodiments of the speech synthesis method based on the latent diffusion model, and will not be repeated here.
[0116] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 5 It is running on the server shown.
[0117] Please see Figure 5 , Figure 5 This is a schematic block diagram of the structure of a server provided in an embodiment of this application.
[0118] like Figure 5 As shown, the server includes a processor, memory, and network interface connected via a system bus, wherein the memory may include storage media and internal memory.
[0119] The storage medium may store the operating system and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any speech synthesis method.
[0120] The processor provides computing and control capabilities to support the operation of the entire server.
[0121] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the solution of this application and does not constitute a limitation on the server to which the solution of this application is applied. A specific server may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0122] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0123] In one embodiment, the processor is used to run a computer program stored in memory to perform the following steps:
[0124] Obtain the target text to be synthesized;
[0125] Based on the target text, determine the emotional features and target duration information of the synthesized speech;
[0126] The target text is encoded using a neural network audio encoder in a preset target speech synthesis model to obtain a first hidden vector.
[0127] The first latent vector is quantized and compressed using the residual vector quantizer in the target speech synthesis model to obtain the second latent vector.
[0128] The vector representations of the speech emotion features, the vector representations of the target duration information, and the second latent vector are fused to obtain a fused latent vector;
[0129] The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model.
[0130] The target latent vector is decoded using the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
[0131] In some embodiments, when the processor performs inverse diffusion processing on the fused latent vector using the latent diffusion model in the target speech synthesis model to obtain the target latent vector, it is configured to:
[0132] Obtain the prior conditions of the potential diffusion model, wherein the prior conditions include at least one of pitch, timbre, and euphemism;
[0133] The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector using the latent diffusion model in the target speech synthesis model and the prior conditions.
[0134] In some embodiments, when the processor performs fusion processing on the vector representation of the speech emotion features, the vector representation of the target duration information, and the second latent vector to obtain a fused latent vector, it is configured to:
[0135] The vector representation of the speech emotion feature and the vector representation of the target duration information are concatenated to obtain the concatenated latent vector;
[0136] The concatenated latent vector is fused with the second latent vector to obtain the fused latent vector.
[0137] In some embodiments, when the processor determines the speech emotion features and target duration information of the synthesized speech based on the target text, it is configured to:
[0138] The target text is processed by a preset speech emotion feature prediction model to obtain the speech emotion features of the synthesized speech. The speech emotion feature prediction model is obtained by training a first neural network model with multiple first samples. The first samples include sample text and labeled speech emotion features.
[0139] The target text is processed by a preset duration prediction model to obtain the target duration information of the synthesized speech. The duration prediction model is obtained by training a second neural network model in advance based on multiple second samples. The second samples include sample text and labeled duration information.
[0140] In some embodiments, the processor is further configured to perform the following before acquiring the target text to be synthesized:
[0141] Obtain a training sample set and a speech synthesis model to be trained. The training samples in the training sample set include training text, transcribed speech corresponding to the training text, and a first training latent vector of the transcribed speech. The first training latent vector is obtained by concatenating the vector representation of the speech emotion features of the transcribed speech and the vector representation of the duration of the transcribed speech.
[0142] Select a training sample from the training sample set;
[0143] The neural network audio encoder in the speech synthesis model encodes the training text in the selected training samples to obtain the third hidden vector.
[0144] The third hidden vector is quantized and compressed using the residual vector quantizer in the speech synthesis model to obtain the fourth hidden vector.
[0145] The first training hidden vector and the fourth hidden vector in the selected training samples are fused to obtain the fifth hidden vector;
[0146] The fifth hidden vector is reverse-diffusioned using the latent diffusion model in the speech synthesis model to obtain the sixth hidden vector.
[0147] The sixth hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
[0148] The model loss value is determined based on the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text.
[0149] When the model loss value is greater than the preset loss value, the parameters of the speech synthesis model are updated, and the process returns to the step of selecting a training sample from the training sample set.
[0150] When the model loss value is less than or equal to a preset loss value, training of the speech synthesis model is stopped, and the target speech synthesis model is obtained.
[0151] In some embodiments, the training samples further include a second training latent vector of the transcribed speech, the second training latent vector being encoded from the transcribed speech. When the processor decodes the sixth latent vector using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text, it is configured to:
[0152] The second training hidden vector and the sixth hidden vector in the selected training samples are fused to obtain the seventh hidden vector;
[0153] The seventh hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
[0154] In some embodiments, when the processor determines the model loss value based on the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text, it is configured to:
[0155] The difference between the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text is determined, and the model loss value is determined based on the difference.
[0156] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the server described above can be referred to the corresponding process in the aforementioned speech synthesis method embodiments, and will not be repeated here.
[0157] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a server (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0158] This application also provides a storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can be found in various embodiments of the speech synthesis method of this application.
[0159] The storage medium can be volatile or non-volatile. It can be an internal storage unit of the server as described in the foregoing embodiments, such as the server's hard drive or memory. Alternatively, it can be an external storage device of the server, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card.
[0160] Furthermore, the storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store the operating system, applications required for at least one function, etc.; and the data storage area may store data created based on the use of blockchain nodes, etc.
[0161] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0162] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0163] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0164] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific implementations of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A speech synthesis method based on a latent diffusion model, characterized in that, include: Obtain the target text to be synthesized; Based on the target text, determine the emotional features and target duration information of the synthesized speech; The target text is encoded using a neural network audio encoder in a preset target speech synthesis model to obtain a first hidden vector. The first latent vector is quantized and compressed using the residual vector quantizer in the target speech synthesis model to obtain the second latent vector. The vector representations of the speech emotion features, the vector representations of the target duration information, and the second latent vector are fused to obtain a fused latent vector; The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model. The target latent vector is decoded using the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
2. The speech synthesis method according to claim 1, characterized in that, The step of performing reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model to obtain the target latent vector includes: Obtain the prior conditions of the potential diffusion model, wherein the prior conditions include at least one of pitch, timbre, and euphemism; The target latent vector is obtained by performing reverse diffusion processing on the fused latent vector using the latent diffusion model in the target speech synthesis model and the prior conditions.
3. The speech synthesis method according to claim 1, characterized in that, The process of fusing the vector representation of the speech emotion features, the vector representation of the target duration information, and the second latent vector to obtain the fused latent vector includes: The vector representation of the speech emotion feature and the vector representation of the target duration information are concatenated to obtain the concatenated latent vector; The concatenated latent vector is fused with the second latent vector to obtain the fused latent vector.
4. The speech synthesis method according to claim 1, characterized in that, The step of determining the speech emotion features and target duration information of the synthesized speech based on the target text includes: The target text is processed by a preset speech emotion feature prediction model to obtain the speech emotion features of the synthesized speech. The speech emotion feature prediction model is obtained by training a first neural network model with multiple first samples. The first samples include sample text and labeled speech emotion features. The target text is processed by a preset duration prediction model to obtain the target duration information of the synthesized speech. The duration prediction model is obtained by training a second neural network model in advance based on multiple second samples. The second samples include sample text and labeled duration information.
5. The speech synthesis method according to any one of claims 1-4, characterized in that, Before obtaining the target text to be synthesized, the process also includes: Obtain a training sample set and a speech synthesis model to be trained. The training samples in the training sample set include training text, transcribed speech corresponding to the training text, and a first training latent vector of the transcribed speech. The first training latent vector is obtained by concatenating the vector representation of the speech emotion features of the transcribed speech and the vector representation of the duration of the transcribed speech. Select a training sample from the training sample set; The neural network audio encoder in the speech synthesis model encodes the training text in the selected training samples to obtain the third hidden vector. The third hidden vector is quantized and compressed using the residual vector quantizer in the speech synthesis model to obtain the fourth hidden vector. The first training hidden vector and the fourth hidden vector in the selected training samples are fused to obtain the fifth hidden vector; The fifth hidden vector is reverse-diffusioned using the latent diffusion model in the speech synthesis model to obtain the sixth hidden vector. The sixth hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text. The model loss value is determined based on the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text. When the model loss value is greater than the preset loss value, the parameters of the speech synthesis model are updated, and the process returns to the step of selecting a training sample from the training sample set. When the model loss value is less than or equal to a preset loss value, training of the speech synthesis model is stopped, and the target speech synthesis model is obtained.
6. The speech synthesis method according to claim 5, characterized in that, The training samples also include a second training latent vector of the transcribed speech, which is obtained by encoding the transcribed speech. The step of decoding the sixth latent vector using a neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text includes: The second training hidden vector and the sixth hidden vector in the selected training samples are fused to obtain the seventh hidden vector; The seventh hidden vector is decoded using the neural network audio decoder in the speech synthesis model to obtain the synthesized speech corresponding to the training text.
7. The speech synthesis method according to claim 5, characterized in that, The step of determining the model loss value based on the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text includes: The difference between the transcribed speech in the selected training samples and the synthesized speech corresponding to the training text is determined, and the model loss value is determined based on the difference.
8. A speech synthesis device based on a latent diffusion model, characterized in that, The speech synthesis device includes: The acquisition module is used to acquire the target text to be synthesized; The determination module is used to determine the emotional features and target duration information of the synthesized speech based on the target text. The encoding module is used to encode the target text using a neural network audio encoder in a preset target speech synthesis model to obtain a first latent vector; The quantization compression module is used to quantize and compress the first latent vector through the residual vector quantizer in the target speech synthesis model to obtain the second latent vector. The fusion module is used to fuse the vector representation of the speech emotion features, the vector representation of the target duration information, and the second latent vector to obtain a fused latent vector; The diffusion module is used to perform reverse diffusion processing on the fused latent vector through the latent diffusion model in the target speech synthesis model to obtain the target latent vector; The decoding module is used to decode the target latent vector through the neural network audio decoder in the target speech synthesis model to obtain the synthesized speech corresponding to the target text.
9. A server, characterized in that, The server includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the speech synthesis method based on the latent diffusion model as described in any one of claims 1 to 7.
10. A storage medium for computer-readable storage, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the speech synthesis method based on a latent diffusion model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
End-to-end speech synthesis method and system based on fusion of acoustic features and text emotion features
CN113506562A
End-to-end speech synthesis method, device, equipment and medium
CN116469375A