Speech synthesis method and device, electronic equipment and readable storage medium

Through the trained speech synthesis model, the text features of the target text are directly converted into synthetic speech, solving the problem of long and heavy calculation burden of speech synthesis in the prior art, and achieving efficient real-time speech synthesis.

CN120108381AInactive Publication Date: 2025-06-06BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Patent Information

Application Number
CN202510593144.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing speech synthesis methods take a long time to generate natural speech signals and are burdened with heavy calculations, which cannot meet the needs of real-time fast speech synthesis.

Method used

Through the trained speech synthesis model, the potential variables of the audio characteristics corresponding to the text characteristics of the target text are directly converted into the target synthetic speech that matches the target text.

Benefits of technology

It improves the synthesis efficiency of speech synthesis, reduces the time-consuming of the inference process, and is suitable for real-time speech synthesis scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108381A_ABST
    Figure CN120108381A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech synthesis, in particular to a speech synthesis method and device, electronic equipment and a readable storage medium, and the method comprises the steps: obtaining a target text of a to-be-synthesized speech and a random sampling noise; inputting the target text into a text feature extraction module of a trained target speech synthesis model, and determining text features of the target text; the target speech synthesis model further comprises a text feature coding module and a speech generation module; inputting the text features and the random sampling noise into a text feature coding module, and determining potential variables of audio features corresponding to the text features; and inputting the potential variable into a voice generation module, generating an audio signal corresponding to the potential variable, and determining the audio signal as a target synthetic voice matched with the target text. Thus, through the trained speech synthesis model, the potential variable of the audio feature corresponding to the text feature of the target text can be directly converted into the target synthesis speech matched with the target text, and the synthesis efficiency of speech synthesis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech synthesis method, device, electronic equipment and readable storage medium. Background Art

[0002] Text-to-speech (TTS) is an artificial intelligence technology that can convert any text into speech that is as natural as possible. Among them, neural network TTS based on deep learning technology can already achieve speech quality close to that of humans in designed tasks.

[0003] Existing speech synthesis methods, such as the speech synthesis method using codec combined with diffusion model, require multiple iterations of denoising process to generate natural speech signals, and the reasoning process takes a long time; while the speech synthesis method using codec combined with autoregressive and non-autoregressive models requires accurate modeling of audio sequences in large-scale feature space, resulting in increased computational burden. These methods all rely on large-scale data and model parameters, resulting in a long time for text-to-speech synthesis, which cannot meet the needs of scenarios such as conversation scenarios that require real-time and fast speech synthesis. Therefore, the synthesis efficiency of speech synthesis is low. Summary of the invention

[0004] In view of this, the embodiments of the present application at least provide a speech synthesis method, device, electronic device and readable storage medium. Through a trained speech synthesis model, the latent variables of the audio features corresponding to the text features of the target text can be directly converted into a target synthesized speech that matches the target text, thereby improving the synthesis efficiency of speech synthesis.

[0005] This application mainly includes the following aspects: In a first aspect, an embodiment of the present application provides a speech synthesis method, the method comprising: Obtain the target text of the speech to be synthesized and a randomly sampled noise; Input the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; Inputting the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature; The latent variables are input into the speech generation module, an audio signal corresponding to the latent variables is generated, and the audio signal is determined as a target synthesized speech matching the target text.

[0006] In a second aspect, an embodiment of the present application further provides a speech synthesis device, the speech synthesis device comprising: An acquisition module, used to acquire a target text of speech to be synthesized and a random sampling noise; A first determination module is used to input the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; A second determination module is used to input the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature; The third determination module is used to input the latent variables into the speech generation module, generate an audio signal corresponding to the latent variables, and determine it as a target synthesized speech matching the target text.

[0007] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the speech synthesis method as described above.

[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the speech synthesis method as described above are executed.

[0009] The embodiment of the present application provides a method, device, electronic device and readable storage medium for speech synthesis, wherein the method comprises: obtaining a target text of speech to be synthesized and a random sampling noise; inputting the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also comprises a text feature encoding module and a speech generation module; inputting the text features and the random sampling noise into the text feature encoding module to determine the latent variables of the audio features corresponding to the text features; inputting the latent variables into the speech generation module to generate an audio signal corresponding to the latent variables, and determining it as a target synthesized speech matching the target text. In this way, through the trained speech synthesis model, the latent variables of the audio features corresponding to the text features of the target text can be directly converted into the target synthesized speech matching the target text, thereby improving the synthesis efficiency of speech synthesis.

[0010] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0012] Figure 1 A flow chart of a speech synthesis method provided by an embodiment of the present application is shown; Figure 2 One of the schematic diagrams showing the reasoning process of the target speech synthesis model in the embodiment of the present application; Figure 3 A second schematic diagram showing the reasoning process of the target speech synthesis model in an embodiment of the present application; Figure 4 A schematic diagram showing the training process of the initial speech synthesis model in an embodiment of the present application is shown; Figure 5 One of the functional module diagrams of a speech synthesis device provided in an embodiment of the present application is shown; Figure 6 A second functional module diagram of a speech synthesis device provided in an embodiment of the present application is shown; Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0014] To facilitate the understanding of the present application, the technical solution provided by the present application is described in detail below in conjunction with specific embodiments.

[0015] The embodiments of the present application provide a speech synthesis method, device, electronic device and readable storage medium. Through a trained speech synthesis model, the latent variables of the audio features corresponding to the text features of the target text can be directly converted into a target synthesized speech that matches the target text, thereby improving the synthesis efficiency of speech synthesis.

[0016] See also Figure 1 , Figure 1 This is a flow chart of a speech synthesis method provided by an embodiment of the present application. Figure 1 As shown, the speech synthesis method provided in the embodiment of the present application includes the following steps: S101, obtaining a target text of speech to be synthesized and a randomly sampled noise.

[0017] In this step, first obtain the target text of the speech to be synthesized and a random sampling noise. Specifically, the target text of the speech to be synthesized is the original text data that the user wants to convert into speech. It can be any type of written language, such as a sentence, a paragraph, or an entire document. The random sampling noise is usually a Gaussian distributed random sample noise, which is used to simulate the natural changes in real speech, such as differences in speakers, fluctuations in emotions, etc.

[0018] S102, inputting the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module.

[0019] In this step, see Figure 2 , Figure 2 FIG. 1 is a schematic diagram of the inference process of the target speech synthesis model in the embodiment of the present application. Figure 2 As shown, the target speech synthesis model includes a text feature extraction module 210, a text feature encoding module 220 and a speech generation module 230. The target text to be synthesized is input into the text feature extraction module 210 to obtain the text features of the target text. Specifically, the text features include information such as the phoneme sequence, tone information, and prosodic features of the target text, which can represent the phonetic, semantic, and grammatical properties of the text.

[0020] S103, inputting the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature.

[0021] In this step, the text features and random sampling noise are input into the text feature encoding module 220 to obtain the latent variables (Latent Variables) of the audio features corresponding to the text features. The text feature encoding module 220 learns how to compress high-dimensional text information into a compact latent representation through deep learning-based model training, and converts the text features into the latent variables of the corresponding audio features. The latent variables are usually located in a continuous latent space, which maps the semantic and grammatical information of the text to acoustic parameters such as pitch, timbre, intensity, etc. At the same time, in order to increase the diversity of the generated speech, random sampling noise is introduced in the text encoding process.

[0022] S104, inputting the latent variables into the speech generation module, generating an audio signal corresponding to the latent variables, and determining the audio signal as a target synthesized speech matching the target text.

[0023] In this step, the latent variables of the audio features corresponding to the text features are input into the speech generation module 230, and the speech generation module converts the latent variables into audio signals, that is, the target synthesized speech that matches the target text. In the embodiment of the present application, the speech generation module 230 uses a trained Encodec decoder, which can directly decode the latent variables to generate the target synthesized speech. Specifically, the Encodec model reconstructs the latent variables through the Encodec decoder on the basis of efficiently compressing the original audio signal to generate the latent variables of low-dimensional potential representation, thereby reducing the processing complexity of high-dimensional audio signals in the decoding stage. Compared with the traditional method of training a large decoder from scratch, the Encodec decoder reduces the scale of decoding parameters, improves the efficiency of speech synthesis, and also improves the training efficiency and reasoning performance of the entire model. When the model parameters are reduced, the required computing resources will also decrease accordingly, which is conducive to model miniaturization.

[0024] At the same time, by introducing the Encodec model for audio compression and decoding, more personalized audio details can be retained. Encodec's efficient coding not only reduces information loss in data compression, but also allows for more delicate audio reconstruction, thereby better reflecting personalized features such as speaking speed, pitch, timbre, etc. Compared with traditional methods, this method can generate high-fidelity personalized speech through low-dimensional potential representation, ensuring that individual differences can be fully expressed in the generation process.

[0025] Furthermore, the text feature extraction module includes a semantic feature extraction module and a text encoding module, and the step of inputting the target text into the text feature extraction module of the trained target speech synthesis model to determine the text features of the target text includes: Step a1: input the target text into the semantic feature extraction module to determine the semantic features of the target text.

[0026] Specifically, see Figure 3 , Figure 3 FIG2 is a second schematic diagram of the inference process of the target speech synthesis model in the embodiment of the present application. Figure 3 As shown, the text feature extraction module 210 includes a semantic feature extraction module 211 and a text encoding module 212. In the embodiment of the present application, the semantic feature extraction module 211 uses a pre-trained BERT (Bidirectional Encoder Representations from Transformers) language model, inputs the target text into the pre-trained Bert language model, and determines the semantic features of the target text. . The benefit of using the BERT language model in speech synthesis is that it can capture contextual information, rich semantic and syntactic features, and improve the naturalness and fluency of speech. The pre-trained knowledge of the BERT language model can help the model generate speech that is more in line with human expression habits, especially when dealing with polysemous words, emotional intonation, and complex sentences. In addition, the multilingual characteristics of the BERT language model and its ability to process rare vocabulary also enhance the adaptability of speech synthesis while reducing dependence on large-scale annotated data. This application uses the BERT language model to improve the contextual consistency and personalization capabilities of the target text-generated speech, making it more adaptable to specific speech synthesis needs.

[0027] Step a2: input the semantic feature into the text encoding module to determine the text feature of the target text that matches the semantic feature.

[0028] In this step, the text encoding module 212 uses a text encoder to encode the semantic features of the target text. Convert to text features of target text that match semantic features .

[0029] Furthermore, the text feature encoding module includes a mapping layer, a random duration prediction module, a data alignment module and a data conversion module; the text feature and the random sampling noise are input into the text feature encoding module to determine the latent variables of the audio feature corresponding to the text feature, including: Step b1, inputting the text features into the mapping layer, and determining the text distribution parameters corresponding to the text features.

[0030] Specifically, if Figure 3As shown, the text feature encoding module 220 includes a mapping layer 221, a random duration prediction module 222, a data alignment module 223 and a data conversion module 224. Input mapping layer 221 to obtain text features The corresponding text distribution parameters. Among them, the text distribution parameters include text features The mean of the corresponding text distribution and standard deviation , which is used to define the latent distribution of text features.

[0031] Step b2: input the text features and random sampling noise into the random duration prediction module to generate an alignment matrix of phoneme durations corresponding to the text features.

[0032] In this step, the random duration prediction module 222 is a trained random duration predictor that converts the text feature and random sampling noise are input into the random duration prediction module 222 to generate text features predicted by the random duration predictor The corresponding duration of each phoneme in . However, since the predicted value is a discrete value, it is rounded up to obtain the text feature. The frame length of each phoneme in the text is obtained Alignment matrix of corresponding phoneme durations .

[0033] Step b3, inputting the text distribution parameters and alignment matrix into the data alignment module to determine the temporal text features after the text features are fused with the corresponding phoneme durations.

[0034] In this step, the text distribution parameters , and the alignment matrix The input data alignment module 223 aligns the text features according to the frame length of each phoneme in the alignment matrix. The text distribution is assigned, that is, the continuous frame length of the phoneme in the alignment matrix is ​​assigned to the text feature Each phoneme in constitutes the temporal text feature after integrating the corresponding phoneme duration .

[0035] Step b4: input the time-series text features into the data conversion module to determine the latent variables of the audio features corresponding to the text features.

[0036] In this step, the time-series text features after the corresponding phoneme duration are fused Input data conversion module 224, where data conversion module 224 is a standardized flow model. Since the text distribution is relatively simple, the time series text features are converted into The latent variables of audio features converted to text features , which can make the synthesized audio closer to the real human voice.

[0037] Furthermore, before obtaining the target text of the speech to be synthesized and a randomly sampled noise, the method further includes: Step c1, training an initial speech synthesis model based on each text sample and an audio sample corresponding to each text sample; the initial speech synthesis model also includes an audio encoding module.

[0038] In this step, before obtaining the target text and a randomly sampled noise for the synthesized speech, the initial speech synthesis model needs to be trained to obtain a trained target speech synthesis model. For details, please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of the training process of the initial speech synthesis model in the embodiment of the present application. Figure 4 As shown, the initial speech synthesis model of the training process, compared with the target speech synthesis model of the inference process, further includes an audio encoding module 240. Specifically, the audio encoding module 240 is an Encodec encoder, which is only used in the model training stage and not in the model inference stage.

[0039] Step c2: If the KL divergence loss function, reconstruction loss function, quantization embedding loss function, random duration prediction loss function and adversarial training loss function of the initial speech synthesis model all meet the corresponding preset conditions, the training is terminated to obtain the target speech synthesis model.

[0040] In this step, if the KL divergence loss function (Kullback-LeiblerDivergence Loss), reconstruction loss function (Reconstruction Loss), quantization embedding loss function (Quantization Embedding Loss), random duration prediction loss function (Random Duration PredictionLoss) and adversarial training loss function (Adversarial Training Loss Function) of the initial speech synthesis model all meet the corresponding preset conditions, it means that the model's ability to generate speech from text meets the requirements of the scene use, and the training is terminated to obtain the target speech synthesis model. Among them, the corresponding preset conditions satisfied by each loss function are specifically set according to the actual situation and are not limited here.

[0041] Furthermore, in the process of training the initial speech synthesis model, the KL divergence loss function and the reconstruction loss function of the initial speech synthesis model are determined according to the following steps: Step d1, during the process of training the initial speech synthesis model, obtaining any text sample and an audio sample corresponding to the text sample.

[0042] In this step, during the training of the initial speech synthesis model, firstly, any text sample and an audio sample corresponding to the text sample are obtained from a training data set that includes various text samples and pre-labeled audio samples corresponding to each text sample.

[0043] Specifically, the initial speech synthesis model adopts the model structure of the variational autoencoder (VAE). The powerful generation ability of VAE benefits from the smooth latent space, which requires the use of a priori encoder and a posterior encoder to provide the corresponding prior distribution and posterior distribution respectively. In the embodiment of the present application, the audio encoding module 240, that is, the Encodec encoder, is used as the part of the VAE posterior encoder, and the text feature encoding module 220 is used as the part of the VAE a priori encoder. The core idea of ​​VAE is to map the original data to a known potential distribution (usually a Gaussian distribution), and then randomly sample from this potential distribution to obtain new data that is similar to the original data. This method is essentially a fitting of the data distribution, and then sampling from the distribution to obtain new data.

[0044] In the initial speech synthesis model of the embodiment of the present application, the KL divergence is used to quantify the difference between the latent space distribution generated by the audio encoding module 240 and the prior distribution, that is, to quantify the difference between the posterior distribution and the prior distribution, and the smaller the difference, the better. When training the initial speech synthesis model, we expect the audio encoding module 240 to encode the input data into a distribution in the latent space, and make the distribution as close as possible to the prior distribution (usually a Gaussian distribution), which is the key to smoothing the latent space.

[0045] Step d2, inputting the text sample into the text feature extraction module of the initial speech synthesis model, and determining the latent variables of the audio features corresponding to the text features through the text feature encoding module.

[0046] In this step, the text sample is input into the text feature extraction module 210 of the initial speech synthesis model, and the latent variables of the audio features corresponding to the text features are determined through the text feature encoding module 220, that is, the prior distribution in the VAE.

[0047] Step d3: input the audio sample into the audio encoding module to generate latent variables of audio features corresponding to the audio sample.

[0048] In this step, the audio sample is input into the audio encoding module 240 of the initial speech synthesis model to generate latent variables of the audio features corresponding to the audio sample, that is, the posterior distribution in the VAE.

[0049] Step d4, determining the KL divergence loss function of the initial speech synthesis model according to the difference between the latent variables of the audio features corresponding to the audio sample and the latent variables of the audio features corresponding to the text features.

[0050] In this step, the KL divergence loss function of the initial speech synthesis model is determined by comparing the difference between the latent variables of the audio features corresponding to the audio samples and the latent variables of the audio features corresponding to the text features, that is, the difference between the posterior distribution and the prior distribution. Specifically, the KL divergence loss function can be expressed by the following formula: ; in, The latent variable for the output The posterior distribution of (Here Refers to the data of the audio sample), this distribution is encoded by the Encodec encoder, which describes how to generate the latent variables of the audio features corresponding to the audio sample from the audio sample data; The latent variable for the output The prior distribution of represents the semantic features of the given target text. and the alignment matrix When the latent variable This distribution is generated by the text encoding module 212, which describes how the data of the text sample affects the latent variable distribution of; is a latent variable The logarithm of the posterior distribution of , indicating that the latent variable is obtained from the Encodec encoder probability; is a latent variable The logarithm of the posterior distribution of , indicating that when given the text semantic features and alignment information, the latent variable probability.

[0051] ; Here we explicitly state that the latent variable is the distribution output from the Encodec encoder The distribution is a Gaussian distribution with a mean of , the standard deviation is .

[0052] Furthermore, by minimizing the KL divergence loss function , the initial speech synthesis model is encouraged to make the potential distribution of the Encodec encoder output close to the prior distribution, which helps the model learn more meaningful potential representations and maintain the continuity and smoothness of the latent space.

[0053] Step d5, inputting the latent variables of the audio features corresponding to the audio sample into the speech generation module of the initial speech synthesis model to generate synthesized speech matching the audio sample.

[0054] In this step, the latent variables of the audio features corresponding to the audio samples are The data is input into the speech generation module 230 of the initial speech synthesis model, that is, the Encodec decoder, to generate synthesized speech matching the audio sample.

[0055] Step d6, determining a reconstruction loss function of the initial speech synthesis model according to the difference between the synthesized speech matching the audio sample and the audio sample.

[0056] Specifically, in VAE, the reconstruction loss function is The linear spectrum obtained by converting the input audio sample and the linear spectrum obtained by converting the output synthetic audio are used for calculation. The present application proposes to use Encodec encoder to replace the posterior encoder in VAE to directly process the audio sample instead of the linear spectrum of the audio sample. However, the method of converting the input audio into a linear spectrum to calculate the reconstruction loss function is the same as that in VAE. It is also necessary to convert the audio into a spectrum to improve the generation ability. Reconstruction loss function The calculation formula is as follows: ; Among them, the reconstruction loss function Mel-spectrogram vector representing the input audio sample The Mel-spectrogram vector of the synthesized audio output The L1 norm (also known as the Manhattan distance) between the two vectors, that is, the sum of the absolute differences between the two vectors.

[0057] Furthermore, we reconstruct the loss function The difference between the mel-spectrogram of the audio sample and the mel-spectrogram of the reconstructed synthetic audio is measured. By minimizing this loss function, the initial speech synthesis model is encouraged to generate reconstructed synthetic audio that is as similar to the audio sample as possible.

[0058] Furthermore, in the process of training the initial speech synthesis model, the quantized embedding loss function of the initial speech synthesis model is determined according to the following steps: Step e1, after inputting the audio sample into the audio encoding module to generate a latent variable of the audio feature corresponding to the audio sample, the latent variable is discretely quantized into a plurality of codebook vectors based on residual vector quantization.

[0059] In the embodiment of the present application, after the audio sample is input into the audio encoding module 240, that is, the Encodec encoder, the Encodec encoder discretely quantizes the latent variable into a plurality of codebook vectors through residual vector quantization (RVQ).

[0060] Specifically, the audio sample is passed through the Encodec encoder to obtain the latent variable , It is calculated according to the downsampling rate for each frame of audio. For example, the total downsampling rate of 16KHz audio samples is 200, that is, each frame corresponds to a 12.5 ms speech segment. That is, 16k / 200=80, 1 / 80=0.0125. Discrete quantization is performed using RVQ, where It is a frame in an audio feature; It is one of the layers of the multi-layer codebook. For example, if a total of 8 layers of codebooks are used for quantization, there should be 8 They respectively represent the quantization results of codebooks at different layers, that is, multiple codebook vectors.

[0061] Step e2, performing vector concatenation on the multiple codebook vectors to obtain quantized and embedded latent variables.

[0062] In the embodiment of the present application, the encoder outputs a discretely distributed latent variable after residual vector quantization. , and the model requires a continuous distribution, so on the basis of RVQ, all discrete codebook vectors are concatenated to get the continuous vector we need again, that is, the latent variable after quantization embedding Specifically, taking the above 8-layer codebook as an example, the quantization results obtained from different codebooks are concatenated. , and obtain the continuous features of a certain frame; finally, concatenate the quantization results of all frames together. , completing the discrete to continuous conversion of the entire audio segment.

[0063] Step e3, determining the quantization embedding loss function of the initial speech synthesis model according to the difference between the latent variables after quantization embedding and the latent variables before quantization embedding.

[0064] In the embodiment of the present application, compared with directly outputting a continuous distribution, the method of first discretizing and then continuousizing has the following two advantages: first, when training other parts of the initial speech synthesis model, there is no need to store continuous vectors that occupy more memory, but only to store discrete codebook vectors that occupy less memory, and then splice them into a continuous distribution when used, which can improve the efficiency of model training; second, when predicting continuous vectors, based on these quantization marks, additional regularized quantization embedding loss functions can be added to discrete classifications ,in, It is a cross entropy loss based on RVQ. It quantizes the residual layer by layer, calculates the difference between the latent variables after quantization embedding and the latent variables before quantization embedding, and converts it into probability distribution through L2 distance (also known as Euclidean distance) to determine the quantization embedding loss function of the initial speech synthesis model.

[0065] Furthermore, the quantization embedding loss function is used to measure the model's ability to generate correct quantization embeddings, and the cross entropy loss of all quantizers is averaged. This loss term is used to optimize the model to ensure that it can effectively select the correct quantizer, thereby improving the accuracy and robustness of quantization embeddings.

[0066] Furthermore, in the process of training the initial speech synthesis model, the random duration prediction loss function of the initial speech synthesis model is determined according to the following steps: Step f1, after the KL divergence loss function, reconstruction loss function, quantization embedding loss function and adversarial training loss function of the initial speech synthesis model all meet corresponding preset conditions, obtain any text sample and the audio sample corresponding to the text sample.

[0067] In the embodiment of the present application, in order to obtain the latent variable of the audio feature corresponding to the text feature through the text feature of the target text, it is necessary to determine the frame length of each phoneme in the text feature. In the training phase, the text features can be obtained by using the monotonic alignment search (MAS) through the text sample and the corresponding audio sample. Alignment matrix corresponding to phoneme duration However, in the inference process of speech synthesis, there is no speech sample for reference, so it is necessary to train an additional phoneme duration predictor to predict the duration of each phoneme and use it in the generation of the alignment matrix for inference. The duration of each phoneme is calculated by summing up all columns of each row, but this method cannot show the diversity of rhythmic changes. Therefore, in order to reflect the diversity of real speech rhythm, the present application uses a random duration predictor. It is a generative model based on the flow model, which is trained in a way that maximizes likelihood estimation. Therefore, after the KL divergence loss function, reconstruction loss function, quantization embedding loss function and adversarial training loss function of the initial speech synthesis model all meet the corresponding preset conditions, that is, after the other parts of the model are trained, any text sample and the audio sample corresponding to the text sample are obtained for training the random duration prediction module 222, that is, the random duration predictor.

[0068] Step f2: input the text sample into the text feature extraction module of the initial speech synthesis model, and determine the text distribution parameters corresponding to the text features of the text sample through a mapping layer.

[0069] In this step, the text sample is first input into the text feature extraction module 210 of the initial speech synthesis model to obtain the text feature, and then the text feature is generated through the mapping layer 221 Corresponding text distribution parameters , , used to solve the text features later Alignment matrix of corresponding phoneme durations .

[0070] Step f3, inputting the audio sample into the audio encoding module, and reversely converting the latent variables of the audio features corresponding to the audio sample through the data conversion module to determine the temporal text features corresponding to the audio sample.

[0071] In this step, the audio sample is input into the audio encoding module 240 to obtain the latent variable of the audio feature corresponding to the audio sample, that is, the posterior distribution. The latent variable is then reversed through the data conversion module 224 to reversely convert the more complex audio distribution into a simpler text distribution to determine the time-series text feature corresponding to the audio sample. .

[0072] Specifically, in order to generate more realistic synthesized speech, it is important to improve the expressive power of the prior distribution because the prior distribution has a weak expressive power. Especially in speech synthesis tasks, the audio distribution has many sampling points, the data distribution is more complex, and the text distribution is relatively simple. When aligning the phoneme duration, the audio distribution needs to be converted into a simpler text distribution to match the phonemes, and the Encodec decoder needs latent variables to restore high-quality audio. It has sufficient richness to convert simple text distribution into complex audio distribution during inference. Therefore, the standardized flow model is introduced as a data conversion module to perform reversible changes between the simpler text distribution generated by the mapping layer and the more complex audio distribution generated by the encoder decoder. The conversion formula from the prior distribution to the posterior distribution using the flow model is shown as follows: ; ; in, is the absolute value of the Jacobian determinant in the Flow model, indicating arrive The scale change of the transformation; is a condition, containing the semantic features of the text And get the corresponding duration information of each phoneme in the monotone search phase (i.e. how many frames each phoneme lasts).

[0073] Step f4, determining the alignment matrix corresponding to the text features of the text sample according to the text distribution parameters corresponding to the text features of the text sample, the temporal text features corresponding to the audio sample, and the monotonic alignment search.

[0074] In the embodiment of the present application, the purpose of monotonic search alignment is to find the best alignment matrix , the specific formula is as follows: ; ; So that under the given text condition and the alignment matrix Under this condition, the probability distribution of generated speech is maximized. The most direct goal of monotone alignment search is to find the corresponding duration for each phoneme in the speech synthesis process, ensuring that the exact position of the phoneme in the time dimension is strictly aligned with the input text. Specifically, this method assigns a clear time period to the phoneme so that each phoneme in the text has a corresponding moment in the speech, thereby ensuring the consistency of the generated speech in time and content. The core of this process is to find the optimal alignment matrix , for precise alignment between text and speech.

[0075] Step f5, inputting the text features of the text sample into the random duration prediction module to generate a prediction alignment matrix corresponding to the text features of the text sample.

[0076] Step f6, determining the random duration prediction loss function of the initial speech synthesis model according to the difference between the predicted alignment matrix corresponding to the text features of the text sample and the alignment matrix corresponding to the text features of the text sample.

[0077] In the embodiment of the present application, there are two problems in directly training the random duration prediction module 222. First, the phoneme duration is a discrete integer and needs to be dequantized using a continuous standard flow model; second, the phoneme duration is a scalar and it is difficult to achieve high-dimensional transformation. Therefore, two strategies, variational dequantization and variational data regularization, are used. is a duration sequence. In the variational dequantization method, the random variable ,Will The range of is limited to [0, 1), so It becomes a positive real number sequence. In the variational data regularization method, the random variable ,Will and Combined together, they form a higher-dimensional potential expression. Since the goal of the random duration predictor is to increase the randomness of the phoneme duration to be close to the rhythm of the real speech, the final output is noise (because of its reversibility, the randomness of the phoneme duration can be guaranteed during inference). The final random duration prediction loss function is The formula is as follows: ; in, is the approximate posterior distribution; Given the text condition The predicted phoneme duration and random variables The joint distribution of Text features ; is a sequence of given durations and text conditions Under the condition that and The approximate posterior distribution of The expected value of is used to estimate the variational lower bound of the log-likelihood of the phoneme duration. Random duration prediction loss function The model parameters are optimized by maximizing the likelihood estimate, thereby improving the accuracy of the model in predicting phoneme duration.

[0078] Furthermore, the gradient backpropagation of the random duration prediction module will be disconnected during training to prevent the gradient generated during the training of this part from affecting other modules.

[0079] Furthermore, if Figure 4 As shown, the initial speech synthesis model also includes a discriminator module 250. The discriminator module 250 and the speech generation module 230 serve as a discriminator and a generator respectively to form an adversarial training model. The discriminator module 250 is used to distinguish between real audio samples. and the fake audio samples generated by the speech generation module 230 The purpose of the speech generation module 230 is to generate sufficiently realistic audio samples so that the discriminator module 250 cannot distinguish between true and false. In the process of training the initial speech synthesis model, the adversarial training loss function of the initial speech synthesis model is simultaneously optimized. The adversarial training loss function includes the discriminator loss function of the discriminator module 250 , the generator loss function of the speech generation module 230 and the feature matching loss function of the speech generation module 230 During the training process, the loss functions of the discriminator module 250 and the speech generation module 230 are optimized alternately. First, the speech generation module 230 is fixed and the discriminator loss function of the discriminator module 250 is optimized. ; Then, fix the discriminator module 250 and optimize the generator loss function of the speech generation module 230 And feature matching loss function .in, and is the least squares loss function, and the specific formula is as follows: ; In the formula, is the mathematical expectation; It is expected that the output of the discriminator for the real sample is close to 1, that is, the discriminator can correctly identify the real sample; It is expected that the output of the discriminator for the generated samples is close to 0, that is, the discriminator can correctly identify the generated samples as false.

[0080] ; In the formula, is the mathematical expectation; It is expected that the discriminator will judge the output of the sample it generates as 1, thereby deceiving the discriminator.

[0081] Feature matching loss function It is used to constrain the representation of the audio samples generated by the generator in each layer of the discriminator to match the real samples. The specific formula is as follows: ; In the formula, Indicates the number of layers of the discriminator; It is Output feature map of the layer discriminator; It is The number of feature maps of the layer; It is L1 distance between the feature representations of real samples and generated samples; feature matching loss function By minimizing the difference between the feature representations of real samples and generated samples at each layer of the discriminator, the generator is forced to learn to generate more realistic samples.

[0082] The model parameters are updated using the gradient descent method, where the discriminator tries to maximize its loss function and the generator tries to minimize its adversarial loss and feature matching loss.

[0083] In an embodiment of the present application, the initial speech synthesis model is jointly optimized under a unified framework through end-to-end training, which effectively avoids the parameter redundancy and transitional dependency problems caused by staged training. Unlike traditional staged training, end-to-end training can simultaneously optimize the parameters of the text encoding, duration prediction, and speech generation modules, thereby reducing the redundancy of intermediate representations and ensuring that the feature representations of each module are collaboratively optimized under a unified goal. In addition, end-to-end training can share gradient information, making parameter updates between modules more consistent, avoiding the optimization inconsistency problem in staged training. This joint optimization method not only improves the generation quality of the model, but also significantly improves the naturalness and efficiency of speech synthesis. The total loss function of end-to-end training It can be expressed as: .

[0084] The embodiment of the present application provides a speech synthesis method, including: obtaining a target text of a speech to be synthesized and a random sampling noise; inputting the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; inputting the text features and the random sampling noise into the text feature encoding module to determine the latent variables of the audio features corresponding to the text features; inputting the latent variables into the speech generation module to generate an audio signal corresponding to the latent variables, and determining it as a target synthesized speech matching the target text. In this way, through the trained speech synthesis model, the latent variables of the audio features corresponding to the text features of the target text can be directly converted into the target synthesized speech matching the target text, thereby improving the synthesis efficiency of speech synthesis.

[0085] Based on the same application concept, the embodiments of the present application also provide a speech synthesis device corresponding to the speech synthesis method provided in the above embodiments. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the speech synthesis method in the above embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0086] See also Figure 5 , Figure 5 This is one of the functional module diagrams of a speech synthesis device provided in an embodiment of the present application. Figure 5 As shown, the speech synthesis device 500 includes: The acquisition module 510 is used to acquire the target text of the speech to be synthesized and a randomly sampled noise.

[0087] The first determination module 520 is used to input the target text into the text feature extraction module of the trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module.

[0088] The second determination module 530 is used to input the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature.

[0089] The third determination module 540 is used to input the latent variables into the speech generation module, generate an audio signal corresponding to the latent variables, and determine it as a target synthesized speech matching the target text.

[0090] Furthermore, the text feature extraction module includes a semantic feature extraction module and a text encoding module. When the first determination module 520 is used to input the target text into the text feature extraction module of the trained target speech synthesis model and determine the text features of the target text, the first determination module 520 is specifically used to: Inputting the target text into the semantic feature extraction module to determine the semantic features of the target text; The semantic feature is input into the text encoding module to determine the text feature of the target text that matches the semantic feature.

[0091] Further, the text feature encoding module includes a mapping layer, a random duration prediction module, a data alignment module and a data conversion module; when the second determination module 530 is used to input the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature, the second determination module 530 is specifically used to: Inputting the text feature into the mapping layer to determine the text distribution parameter corresponding to the text feature; Inputting the text features and random sampling noise into the random duration prediction module to generate an alignment matrix of phoneme durations corresponding to the text features; Inputting the text distribution parameter and the alignment matrix into the data alignment module to determine the temporal text features after the text features are fused with the corresponding phoneme durations; The time-series text features are input into the data conversion module to determine the latent variables of the audio features corresponding to the text features.

[0092] For further information, see Figure 6 , Figure 6 This is a second functional module diagram of a speech synthesis device provided in an embodiment of the present application. Figure 6 As shown, the speech synthesis device 500 also includes: The model training module 550 is used to train the initial speech synthesis model based on each text sample and the audio sample corresponding to each text sample; the initial speech synthesis model also includes an audio encoding module.

[0093] The model determination module 560 is used to terminate the training and obtain the target speech synthesis model if the KL divergence loss function, reconstruction loss function, quantization embedding loss function, random duration prediction loss function and adversarial training loss function of the initial speech synthesis model all meet the corresponding preset conditions.

[0094] Furthermore, in the process of training the initial speech synthesis model, the model training module 550 determines the KL divergence loss function and the reconstruction loss function of the initial speech synthesis model according to the following steps: In the process of training the initial speech synthesis model, obtaining any text sample and an audio sample corresponding to the text sample; Inputting the text sample into the text feature extraction module of the initial speech synthesis model, and determining the latent variables of the audio features corresponding to the text features through the text feature encoding module; Inputting the audio sample into the audio encoding module to generate a latent variable of the audio feature corresponding to the audio sample; Determining a KL divergence loss function of the initial speech synthesis model according to a difference between a latent variable of the audio feature corresponding to the audio sample and a latent variable of the audio feature corresponding to the text feature; Inputting the latent variables of the audio features corresponding to the audio sample into the speech generation module of the initial speech synthesis model to generate synthesized speech matching the audio sample; A reconstruction loss function of the initial speech synthesis model is determined according to a difference between the synthesized speech matching the audio sample and the audio sample.

[0095] Furthermore, in the process of training the initial speech synthesis model, the model training module 550 determines the quantized embedding loss function of the initial speech synthesis model according to the following steps: After the audio sample is input into the audio encoding module to generate a latent variable of the audio feature corresponding to the audio sample, the latent variable is discretely quantized into a plurality of codebook vectors based on residual vector quantization; Performing vector concatenation on the multiple codebook vectors to obtain quantized and embedded latent variables; According to the difference between the latent variables after quantization embedding and the latent variables before quantization embedding, the quantization embedding loss function of the initial speech synthesis model is determined.

[0096] Furthermore, in the process of training the initial speech synthesis model, the model training module 550 determines the random duration prediction loss function of the initial speech synthesis model according to the following steps: After the KL divergence loss function, the reconstruction loss function, the quantization embedding loss function, and the adversarial training loss function of the initial speech synthesis model all meet corresponding preset conditions, obtaining any text sample and an audio sample corresponding to the text sample; Inputting the text sample into the text feature extraction module of the initial speech synthesis model, and determining the text distribution parameters corresponding to the text features of the text sample through a mapping layer; The audio sample is input into the audio encoding module, and the latent variables of the audio features corresponding to the audio sample are reversely converted through the data conversion module to determine the temporal text features corresponding to the audio sample; Determine an alignment matrix corresponding to the text feature of the text sample according to a text distribution parameter corresponding to the text feature of the text sample, a temporal text feature corresponding to the audio sample, and a monotonic alignment search; Inputting the text features of the text sample into the random duration prediction module to generate a prediction alignment matrix corresponding to the text features of the text sample; The random duration prediction loss function of the initial speech synthesis model is determined according to the difference between the predicted alignment matrix corresponding to the text features of the text sample and the alignment matrix corresponding to the text features of the text sample.

[0097] The embodiment of the present application provides a speech synthesis device, including: an acquisition module, used to acquire a target text of the speech to be synthesized and a random sampling noise; a first determination module, used to input the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; a second determination module, used to input the text features and the random sampling noise into the text feature encoding module to determine the latent variables of the audio features corresponding to the text features; a third determination module, used to input the latent variables into the speech generation module, generate an audio signal corresponding to the latent variables, and determine it as a target synthesized speech matching the target text. In this way, through the trained speech synthesis model, the latent variables of the audio features corresponding to the text features of the target text can be directly converted into the target synthesized speech matching the target text, thereby improving the synthesis efficiency of speech synthesis.

[0098] Based on the same application idea, please refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 700 includes: a processor 710 , a memory 720 and a bus 730 .

[0099] The memory 720 stores machine-readable instructions executable by the processor 710. When the electronic device 700 is running, the processor 710 communicates with the memory 720 through the bus 730. The machine-readable instructions are executed by the processor 710 to perform the steps of the speech synthesis method provided in the above embodiment. The specific implementation method can be found in the method embodiment, which will not be repeated here.

[0100] Based on the same application concept, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech synthesis method provided in the above embodiment are executed. The specific implementation method can be found in the method embodiment, which will not be repeated here.

[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0102] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0105] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0106] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.

[0107] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Obtain the target text of the speech to be synthesized and a randomly sampled noise; Input the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; Inputting the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature; The latent variables are input into the speech generation module, an audio signal corresponding to the latent variables is generated, and the audio signal is determined as a target synthesized speech matching the target text.

2. The speech synthesis method according to claim 1, characterized in that: The text feature extraction module includes a semantic feature extraction module and a text encoding module. The text feature extraction module inputs the target text into the trained target speech synthesis model to determine the text features of the target text, including: Inputting the target text into the semantic feature extraction module to determine the semantic features of the target text; The semantic feature is input into the text encoding module to determine the text feature of the target text that matches the semantic feature.

3. The speech synthesis method according to claim 1, characterized in that: The text feature encoding module includes a mapping layer, a random duration prediction module, a data alignment module and a data conversion module; the text feature and the random sampling noise are input into the text feature encoding module to determine the latent variables of the audio feature corresponding to the text feature, including: Inputting the text feature into the mapping layer, and determining the text distribution parameter corresponding to the text feature; Inputting the text features and random sampling noise into the random duration prediction module to generate an alignment matrix of phoneme durations corresponding to the text features; Inputting the text distribution parameter and the alignment matrix into the data alignment module to determine the temporal text features after the text features are fused with the corresponding phoneme durations; The time-series text features are input into the data conversion module to determine the latent variables of the audio features corresponding to the text features.

4. The speech synthesis method according to claim 3, characterized in that: Before obtaining the target text of the speech to be synthesized and a randomly sampled noise, the method further includes: Training an initial speech synthesis model based on each text sample and an audio sample corresponding to each text sample; the initial speech synthesis model also includes an audio encoding module; If the KL divergence loss function, reconstruction loss function, quantization embedding loss function, random duration prediction loss function and adversarial training loss function of the initial speech synthesis model all meet the corresponding preset conditions, the training is terminated and the target speech synthesis model is obtained.

5. The speech synthesis method according to claim 4, characterized in that: In the process of training the initial speech synthesis model, the KL divergence loss function and the reconstruction loss function of the initial speech synthesis model are determined according to the following steps: In the process of training the initial speech synthesis model, obtaining any text sample and an audio sample corresponding to the text sample; Inputting the text sample into the text feature extraction module of the initial speech synthesis model, and determining the latent variables of the audio features corresponding to the text features through the text feature encoding module; Inputting the audio sample into the audio encoding module to generate a latent variable of the audio feature corresponding to the audio sample; Determining a KL divergence loss function of the initial speech synthesis model according to a difference between a latent variable of the audio feature corresponding to the audio sample and a latent variable of the audio feature corresponding to the text feature; Inputting the latent variables of the audio features corresponding to the audio sample into the speech generation module of the initial speech synthesis model to generate synthesized speech matching the audio sample; A reconstruction loss function of the initial speech synthesis model is determined according to a difference between the synthesized speech matching the audio sample and the audio sample.

6. The speech synthesis method according to claim 5, characterized in that: In the process of training the initial speech synthesis model, the quantized embedding loss function of the initial speech synthesis model is determined according to the following steps: After the audio sample is input into the audio encoding module to generate a latent variable of the audio feature corresponding to the audio sample, the latent variable is discretely quantized into a plurality of codebook vectors based on residual vector quantization; Performing vector concatenation on the multiple codebook vectors to obtain quantized and embedded latent variables; According to the difference between the latent variables after quantization embedding and the latent variables before quantization embedding, the quantization embedding loss function of the initial speech synthesis model is determined.

7. The speech synthesis method according to claim 4, characterized in that: In the process of training the initial speech synthesis model, the random duration prediction loss function of the initial speech synthesis model is determined according to the following steps: After the KL divergence loss function, the reconstruction loss function, the quantization embedding loss function, and the adversarial training loss function of the initial speech synthesis model all meet corresponding preset conditions, obtaining any text sample and an audio sample corresponding to the text sample; Inputting the text sample into the text feature extraction module of the initial speech synthesis model, and determining the text distribution parameters corresponding to the text features of the text sample through a mapping layer; The audio sample is input into the audio encoding module, and the latent variables of the audio features corresponding to the audio sample are reversely converted through the data conversion module to determine the temporal text features corresponding to the audio sample; Determine an alignment matrix corresponding to the text feature of the text sample according to a text distribution parameter corresponding to the text feature of the text sample, a temporal text feature corresponding to the audio sample, and a monotonic alignment search; Inputting the text features of the text sample into the random duration prediction module to generate a prediction alignment matrix corresponding to the text features of the text sample; The random duration prediction loss function of the initial speech synthesis model is determined according to the difference between the predicted alignment matrix corresponding to the text features of the text sample and the alignment matrix corresponding to the text features of the text sample.

8. A speech synthesis device, characterized in that: The speech synthesis device comprises: An acquisition module, used to acquire a target text of speech to be synthesized and a random sampling noise; A first determination module is used to input the target text into a text feature extraction module of a trained target speech synthesis model to determine the text features of the target text; the target speech synthesis model also includes a text feature encoding module and a speech generation module; A second determination module is used to input the text feature and the random sampling noise into the text feature encoding module to determine the latent variable of the audio feature corresponding to the text feature; The third determination module is used to input the latent variables into the speech generation module, generate an audio signal corresponding to the latent variables, and determine it as a target synthesized speech matching the target text.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the speech synthesis method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the speech synthesis method according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN117711371A

  • Speech synthesis method and device, computer equipment and storage medium

    CN119580700A

  • Synthetic audio output method and apparatus, storage medium, and electronic device

    US12051400B1

  • Method of training speech recognition model, electronic device and storage medium

    US20230386448A1

Cited By

  • Discrete audio feature generation method and device and audio data word segmentation device training method and device

    CN120748434A

  • Sound effect generation method and device and electronic equipment

    CN120853546A