Model training, speech synthesis method and device, electronic equipment and storage medium

By optimizing the generative model through adversarial training and periodic feature extraction, the problem of naturalness and realism in timbre reconstruction in speech synthesis is solved, achieving high-quality personalized timbre synthesis and reducing storage space requirements.

CN119763544BActive Publication Date: 2025-11-18ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411842981.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-11-18
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing speech synthesis technologies struggle to accurately reconstruct and mimic personalized timbres, resulting in a lack of naturalness and realism in the timbres. Furthermore, traditional methods require significant storage space and complex signal processing, which negatively impacts the quality of synthesized speech.

Method used

By extracting the periodic features of the reconstructed acoustic features output by the generative model and inputting them into the discriminative model for authenticity determination, adversarial training is conducted between the generative and discriminative models to explicitly optimize timbre features. Multiple periodic components based on prime numbers are used for feature extraction and discrimination, and adversarial training is performed in conjunction with real speech features.

Benefits of technology

It improves the naturalness and realism of the timbre in speech synthesis, reduces storage space requirements, and enhances the generalization of timbre reconstruction and the quality of synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763544B_ABST
    Figure CN119763544B_ABST
Patent Text Reader

Abstract

The application provides a model training method, a speech synthesis method, a device, electronic equipment and a storage medium. The model training method comprises: obtaining reconstructed acoustic features output by a generation model, and extracting a reconstructed periodic feature of the reconstructed acoustic features; inputting the reconstructed periodic feature into a discriminative model to obtain a true or false discrimination result of the reconstructed periodic feature output by the discriminative model; performing adversarial training on the generation model and the discriminative model based on the true or false discrimination result, and taking the trained generation model as an acoustic model. The method, device, electronic equipment and storage medium provided by the application explicitly take the periodic feature capable of reflecting timbre as an optimization target, so that the adversarial training makes the generation model have higher naturalness and higher fidelity of synthesized speech when applied to speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model training, speech synthesis method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of artificial intelligence technology, speech synthesis has evolved from the initial simple text-to-speech (TTS) conversion to a synthesis technology that focuses on complex emotional expression and personalized timbre generation. How to accurately reconstruct and imitate a specific timbre has become a major challenge for speech synthesis technology.

[0003] Because everyone's pronunciation habits and speech characteristics are different, it is difficult to find a universal method to simulate all timbres. Even if a specific timbre can be simulated, it is difficult to guarantee its naturalness and realism. Current speech synthesis solutions usually mix timbre, prosody, and content information together for constrained training, which also leads to the timbre reconstruction not being as natural as expected. Summary of the Invention

[0004] This invention provides a model training, speech synthesis method, apparatus, electronic device, and storage medium to address the shortcomings of speech synthesis in related technologies where the naturalness of the timbre does not meet expectations.

[0005] This invention provides a model training method, comprising:

[0006] Obtain the reconstructed acoustic features output by the generative model, and extract the reconstruction periodicity features of the reconstructed acoustic features;

[0007] The reconstructed periodic features are input into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model;

[0008] Based on the true / false discrimination results, the generative model and the discrimination model are subjected to adversarial training, and the resulting generative model is used as the acoustic model.

[0009] According to a model training method provided by the present invention, the extraction of reconstructed periodic features of the reconstructed acoustic features includes:

[0010] Extract the reconstructed periodic features of the reconstructed acoustic features under multiple periodic components;

[0011] The step of inputting the reconstructed periodic features into the discrimination model to obtain the true / false discrimination result of the reconstructed periodic features output by the discrimination model includes:

[0012] The reconstructed periodic features under each periodic component are input into the discrimination model corresponding to each periodic component, and the true and false discrimination results output by the discrimination model corresponding to each periodic component are obtained.

[0013] According to a model training method provided by the present invention, the plurality of periodic components are a plurality of periodic components based on prime numbers.

[0014] According to a model training method provided by the present invention, the step of extracting the reconstructed periodic features of the reconstructed acoustic features under multiple periodic components includes:

[0015] The reconstructed acoustic features are converted into a linear amplitude spectrum.

[0016] Reconstructed periodic features under multiple periodic components are extracted from the linear amplitude spectrum.

[0017] According to a model training method provided by the present invention, the step of performing adversarial training on the generative model and the discriminative model based on the true / false discrimination result includes:

[0018] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, the generative model and the discriminative model are subjected to adversarial training.

[0019] The authenticity determination result of the real periodic feature is output by the discrimination model based on the real periodic feature, which is extracted from the real acoustic features.

[0020] According to a model training method provided by the present invention, the true / false discrimination results based on the real periodic features and the true / false discrimination results based on the reconstructed periodic features are used to perform adversarial training on the generative model and the discriminative model, including:

[0021] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, a discrimination loss is determined, and the discrimination model is iterated based on the discrimination loss.

[0022] Based on the true / false discrimination results of the reconstructed periodic features, the generation loss is determined, and the generation model is iterated based on the generation loss.

[0023] The present invention also provides a speech synthesis method, comprising:

[0024] Text is input into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model;

[0025] The speech synthesis model includes an acoustic model trained using a model training method.

[0026] The present invention also provides a model training apparatus, comprising:

[0027] The feature extraction unit is used to obtain the reconstructed acoustic features output by the generation model and extract the reconstruction periodicity features of the reconstructed acoustic features.

[0028] The authenticity discrimination unit is used to input the reconstructed periodic features into the discrimination model and obtain the authenticity discrimination result of the reconstructed periodic features output by the discrimination model;

[0029] The adversarial training unit is used to perform adversarial training on the generative model and the discriminative model based on the true / false discrimination results, and to use the trained generative model as the acoustic model.

[0030] The present invention also provides a speech synthesis device, comprising:

[0031] A speech synthesis unit is used to input text into a speech synthesis model and obtain synthesized speech output by the speech synthesis model.

[0032] The speech synthesis model includes an acoustic model trained using a model training method.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the model training methods or speech synthesis methods described above.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model training method or speech synthesis method as described above.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the model training methods or speech synthesis methods described above.

[0036] The model training, speech synthesis method, device, electronic device, and storage medium provided by this invention input the reconstructed periodic features of the reconstructed acoustic features output by the generator model into the discriminant model for true / false discrimination. This allows for adversarial training between the generator model and the discriminant model. In this process, the periodic features that reflect timbre are explicitly used as optimization targets. As a result, when the generator model is applied to speech synthesis, the synthesized speech has a higher degree of naturalness and realism. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is one of the flowcharts illustrating the model training method provided by the present invention.

[0039] Figure 2 This is the acoustic feature spectrum provided by the present invention.

[0040] Figure 3 This is a schematic diagram of the extraction of periodic features provided by the present invention.

[0041] Figure 4 This is the second flowchart of the model training method provided by the present invention.

[0042] Figure 5 This is a flowchart illustrating the speech synthesis method provided by the present invention.

[0043] Figure 6 This is a schematic diagram of the model training device provided by the present invention.

[0044] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0046] Speech synthesis is the technology that converts text information into audible speech signals, enabling machines to "speak" like humans. With the development of artificial intelligence, speech synthesis has evolved from simple text-to-speech conversion to a synthesis technology that focuses on complex emotional expression and personalized timbre generation. Timbre, as a crucial attribute of speech, determines the individuality and recognizability of synthesized speech. Therefore, accurately reconstructing and imitating specific timbres has become a key challenge in speech synthesis technology.

[0047] In the field of speech synthesis technology, traditional speech synthesis methods rely on a large number of pre-recorded speech samples and complex signal processing techniques to simulate and reconstruct a specific timbre. The inherent limitations of this method restrict its widespread adoption in practical applications.

[0048] First, traditional speech synthesis methods require a large amount of storage space to store pre-recorded speech samples. As the number of speech samples increases, the required storage space also increases dramatically, which is undoubtedly a huge challenge for resource-constrained devices or application scenarios.

[0049] Secondly, because everyone's pronunciation habits and speech characteristics are unique, it is difficult to find a universal method to simulate all timbres. Even if it is possible to simulate the timbres of a specific individual or group, it is difficult to guarantee the naturalness and realism of the generated speech, which directly affects the overall effect of speech synthesis.

[0050] To address these issues, speech synthesis methods based on machine learning and deep learning have emerged. These methods train models to learn the features of speech signals and attempt to reconstruct signals similar to the original speech. However, while these methods have achieved significant results in the field of speech synthesis, during model training, they typically mix information such as timbre, prosody, and content together as constraints, without explicitly designing optimization targets for timbre. This results in the trained speech synthesis models not achieving the expected naturalness in timbre reconstruction.

[0051] Furthermore, such models typically use acoustic features like Mel-Frequency Cepstral Coefficients (MFCCs) as modeling targets, constructing a loss through the MFCC to constrain the reconstruction process. Speech, as a strongly periodic signal, has formants in its spectrogram that contribute significantly to timbre, representing a strong periodicity. However, the MFCC compresses some high-frequency information based on human hearing characteristics, simplifying the reconstruction task but also disrupting the periodicity of the audio spectrogram. This directly leads to the loss of features contributing to timbre, thus affecting the timbre reconstruction effect.

[0052] To address the above problems, this invention provides a model training method. Figure 1 This is one of the flowcharts illustrating the model training method provided by this invention, such as... Figure 1 As shown, the method includes:

[0053] Step 110: Obtain the reconstructed acoustic features output by the generated model, and extract the reconstruction periodicity features of the reconstructed acoustic features.

[0054] Here, the generative model is the model to be trained, specifically the initial model of the acoustic model to be trained.

[0055] For example, when the generative model is an initial model of an acoustic model for speech synthesis, text can be input into the generative model, which then performs speech synthesis based on the input text, thereby outputting reconstructed acoustic features. It is understood that the reconstructed acoustic features here are the acoustic features output by the generative model, and these reconstructed acoustic features reflect the acoustic features of the synthesized speech generated based on the generative model. Specifically, these acoustic features can be F-bank (FilterBank) features, Mel spectrograms, perceptual linear predictive (PLP) features, etc., but this embodiment of the invention does not specifically limit them.

[0056] After obtaining the reconstructed acoustic features, periodic features can be extracted from them. In this embodiment of the invention, the periodic features extracted from the reconstructed acoustic features are denoted as reconstructed periodic features.

[0057] Understandably, the periodicity of speech has a significant impact on the perception of timbre. Periodic features, as a type of feature reflecting the periodicity of speech, can characterize the periodic information in the speech frequency domain, especially the formant information that determines timbre, and thus have a significant influence and contribution to timbre. By extracting reconstructed periodic features from reconstructed acoustic features, we can provide a basis for subsequent evaluation of the reconstruction effect of reconstructed acoustic features on timbre.

[0058] There are several ways to extract periodic features here. For example, for a Mel spectrum, the Mel spectrum can be inversely transformed into a linear amplitude spectrum, and then periodic features can be extracted from the linear amplitude spectrum; another example is that a model for extracting periodic features from acoustic features can be pre-trained, and the extraction of periodic features can be achieved based on the trained model. This embodiment of the invention does not specifically limit this approach.

[0059] Step 120: Input the reconstructed periodic features into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model.

[0060] Here, the discriminant model is used to determine the authenticity of input periodic features. It can be understood that a true periodic feature means the corresponding acoustic feature originates from real speech, while a false periodic feature means the corresponding acoustic feature originates from synthesized fake speech. In other words, determining the authenticity of periodic features can also be understood as determining whether the speech from which the acoustic feature originates is real speech or synthesized fake speech.

[0061] After obtaining the reconstructed periodic features, these features can be input into a discriminant model. The model then performs a true / false judgment on the input reconstructed periodic features, thus obtaining the true / false judgment result of the reconstructed periodic features output by the model. This true / false judgment result can be from real speech, from synthesized fake speech, or it can be the probability of either real or synthesized fake speech. This embodiment of the invention does not specifically limit this.

[0062] Step 130: Based on the true / false discrimination result, perform adversarial training on the generative model and the discrimination model, and use the trained generative model as the acoustic model.

[0063] Specifically, after obtaining the true / false discrimination results, the aforementioned generative model and discriminative model can be trained adversarially based on these results.

[0064] In adversarial training, the generative model can be regarded as a generator and the discriminative model as a discriminator. The generator and the discriminator compete with each other. The generator aims to make the reconstructed periodic features of the output acoustic features as close as possible to the real periodic features, so that the discriminator can hardly distinguish between the reconstructed periodic features and the real periodic features. The discriminator aims to make the output true or false discrimination result consistent with the actual situation of the input periodic features, so as to achieve a more accurate and reliable true or false discrimination effect.

[0065] During adversarial training, the true / false discrimination results can reflect the difference between the reconstructed periodic features of the reconstructed acoustic features output by the generative model and the true periodic features. This allows us to determine the generation loss for the generative model, guiding it to update and iterate in the direction of outputting reconstructed acoustic features with periodic features that are closer to the true periodic features. Furthermore, the true / false discrimination results can also reflect the accuracy of the discriminative model in distinguishing the true / false nature of the input periodic features. This allows us to determine the discriminative loss for the discriminative model, guiding it to update and iterate in the direction of outputting true / false discrimination results that are closer to the real situation.

[0066] After completing adversarial training on the generative and discriminative models, the trained generative and discriminative models are obtained. The trained generative model not only possesses the ability to map text features to acoustic features, but also exhibits periodic features in the reconstructed acoustic features output by the generative model that are very close to the real periodic features. In other words, the timbre of the synthesized speech based on the reconstructed acoustic features output by the generative model is very close to the real timbre, ensuring the naturalness and realism of the synthesized speech timbre.

[0067] In the method provided in this embodiment of the invention, the reconstructed periodic features of the reconstructed acoustic features output by the generative model are input into the discriminative model for true / false discrimination. Adversarial training is then performed on the generative model and the discriminative model. In this process, the periodic features that can reflect timbre are explicitly used as optimization targets. As a result, when the generative model is applied to speech synthesis, the timbre of the synthesized speech is more natural and realistic.

[0068] Furthermore, the training method described above involves two types of models: generative models and discriminative models. It has a simple structure, is easy to port, and can be applied to any acoustic model that targets Mel spectrograms or linear amplitude spectrograms to improve the timbre reconstruction capability of such acoustic models.

[0069] Based on the above embodiments, the discrimination model consists of one or more two-dimensional convolutional layers.

[0070] Specifically, in the discriminative model, two-dimensional convolutional layers can effectively extract local features of periodic features in the form of the input image, including the edges and energy changes of the periodic feature image.

[0071] Furthermore, two-dimensional convolutional layers can slide their kernels across periodic feature images to perform operations on local regions of the periodic feature images. Due to the characteristics of convolution operations, the discriminative model can possess position invariance, meaning that the discriminative model can correctly identify any feature used for true / false discrimination appearing at any location in the periodic feature image.

[0072] Furthermore, the application of multiple two-dimensional convolutional layers in the discriminative model enables hierarchical feature extraction. That is, the shallow two-dimensional convolutional layers extract low-level features, while subsequent two-dimensional convolutional layers gradually extract higher-level features. This allows the discriminative model to gradually build complex feature layers under multiple two-dimensional convolutions, and finally achieve true / false discrimination for periodic features based on hierarchical features.

[0073] Based on any of the above embodiments, step 110, extracting the reconstructed periodic features of the reconstructed acoustic features, includes:

[0074] Extract the reconstructed periodic features of the reconstructed acoustic features under multiple periodic components.

[0075] Specifically, different people have different timbres, which correspond to different formants in speech, and thus the periodic components of harmonics in speech are also different. In order to enable the reconstructed periodic features to cover more timbre-related information and thus bring better timbre generalization, when extracting periodic features for reconstructed acoustic features, periodic features can be extracted separately for multiple periodic components.

[0076] This allows us to obtain the reconstruction periodicity of the reconstructed acoustic features under multiple periodic components. That is, the reconstructed acoustic features can have multiple reconstruction periodic features, each corresponding to a periodic component. For example, we can extract the reconstruction periodicity of the reconstructed acoustic features under periodic components such as 2, 3, 5, 7, and 11.

[0077] It is understandable that the reconstructed periodic features under multiple periodic components can relatively completely and comprehensively cover the timbre-related information in the speech corresponding to the reconstructed acoustic features.

[0078] Accordingly, in step 120, the step of inputting the reconstructed periodic features into the discrimination model to obtain the true / false discrimination result of the reconstructed periodic features output by the discrimination model includes:

[0079] The reconstructed periodic features under each periodic component are input into the discrimination model corresponding to each periodic component, and the true and false discrimination results output by the discrimination model corresponding to each periodic component are obtained.

[0080] Specifically, for reconstructing periodic features under multiple periodic components, a discriminant model corresponding to each periodic component can be pre-trained. It can be understood that for each periodic component, there is a corresponding discriminant model used to determine the authenticity of the periodic features under that component. That is, the number of periodic components is consistent with the number of discriminant models.

[0081] After obtaining the reconstructed periodic features of the acoustic features under multiple periodic components, the reconstructed periodic features under each periodic component can be input into the discrimination model under the corresponding periodic component, thereby obtaining the true / false discrimination results output by the discrimination model under each periodic component. It can be understood that each discrimination model outputs a true / false discrimination result for the reconstructed periodic features under that periodic component, and the number of true / false discrimination results obtained is consistent with the number of periodic components.

[0082] In subsequent adversarial training combining the generative and discriminative models, the true / false discrimination results under each periodic component can be used as the basis for guiding the parameter iteration of the generative and corresponding discriminative models. That is, the parameter iteration of the generative model can refer to the true / false discrimination results under all periodic components, while the parameter iteration of each discriminative model refers to the true / false discrimination results under the periodic component corresponding to that discriminative model.

[0083] In the method provided in the embodiments of the present invention, by extracting periodic features from multiple periodic components and performing true and false discrimination respectively, the coverage of periodic features in the periodic harmonic range can be guaranteed, thereby ensuring the generalization of model training in timbre reconstruction optimization.

[0084] Based on any of the above embodiments, the plurality of periodic components are a plurality of periodic components based on prime numbers.

[0085] Here, a prime number is a natural number greater than 1 that is not divisible by any other natural number except 1 and itself, such as 2, 3, 5, 7, etc. The unique divisibility property of prime numbers makes the distribution of periodic components based on prime numbers more uniform across the spectrum, reducing overlap and interference between different periodic components. This facilitates maintaining the independence and clarity of periodic characteristics under different periodic components.

[0086] Furthermore, multiple periodic components based on prime numbers can cover a wider harmonic range, allowing the periodic characteristics obtained from these multiple periodic components to more comprehensively reflect relevant timbre information.

[0087] Based on any of the above embodiments, step 110, which involves extracting the reconstructed periodic features of the reconstructed acoustic features under multiple periodic components, includes:

[0088] The reconstructed acoustic features are converted into a linear amplitude spectrum.

[0089] Reconstructed periodic features under multiple periodic components are extracted from the linear amplitude spectrum.

[0090] Specifically, for acoustic features, such as reconstructing acoustic features, the method of periodic feature extraction can be represented by the following steps:

[0091] First, the reconstructed acoustic features can be converted into a linear amplitude spectrum. Understandably, this is different from a Mel spectrum. Figure 1 For reconstructing acoustic features, linear amplitude spectrograms can retain most of the periodic features of speech in the frequency domain, including formant information that determines timbre. Therefore, linear amplitude spectrograms are more suitable for periodic feature extraction than reconstructed acoustic features.

[0092] Taking the reconstructed acoustic features as a Mel spectrum as an example, generally, speech needs to be converted into a frequency domain signal through a short-time Fourier transform, thus obtaining a linear amplitude spectrum. Since the human ear's perception of frequency is non-linear, the frequency axis of the linear scale in the linear amplitude spectrum needs to be converted to a Mel frequency scale. The Mel scale simulates how the human ear perceives different frequencies; that is, the human ear is more sensitive to low-frequency components and relatively less sensitive to high-frequency components. A Mel filter bank is used to weight and sum the spectrum, calculating the energy or amplitude of each Mel frequency band, thus forming the Mel spectrum. In practice, the conversion from a linear amplitude spectrum to a Mel spectrum is usually achieved by multiplying the Mel matrix by the linear amplitude spectrum. Therefore, after obtaining the reconstructed acoustic features in the form of a Mel spectrum, the Mel spectrum can be multiplied by the pseudo-inverse of the Mel matrix to achieve the inverse transformation, obtaining a rough linear amplitude spectrum, which is the linear amplitude spectrum obtained in this embodiment of the invention.

[0093] For example, Figure 2 This is the acoustic feature spectrum provided by the present invention. Figure 2 In the diagram, (a) is the original linear amplitude spectrum, (b) is the Mel spectrum obtained by transforming the linear amplitude spectrum, and (c) is a coarse linear amplitude spectrum obtained by inverse transformation of the Mel spectrum. Comparing (a) and (c), it can be seen that the coarse linear amplitude spectrum loses some detail compared to the original linear amplitude spectrum, but the periodic features, especially the formant information that determines timbre, are still preserved. Therefore, periodic features can be extracted from the linear amplitude spectrum obtained by transforming the Mel spectrum.

[0094] After obtaining the linear amplitude spectrum, reconstructed periodic features can be extracted from multiple periodic components of the linear amplitude spectrum. For example, for a single frame of linear amplitude spectrum, reconstructed periodic features can be constructed based on periodic components such as 2, 3, 5, 7, and 11. Figure 3 This is a schematic diagram of the extraction of periodic features provided by the present invention. Figure 3 In the image, (a) is a linear amplitude spectrum, and (b), (c), and (d) are feature maps of periodic features with periods of 2, 3, and 5 extracted from (a), respectively.

[0095] Based on any of the above embodiments, step 130, which involves performing adversarial training on the generative model and the discriminative model based on the authenticity discrimination result, includes:

[0096] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, the generative model and the discriminative model are subjected to adversarial training.

[0097] The authenticity determination result of the real periodic feature is output by the discrimination model based on the real periodic feature, which is extracted from the real acoustic features.

[0098] Specifically, in adversarial training that combines generative and discriminative models, in addition to applying the true / false discrimination results for reconstructed periodic features, the true / false discrimination results for real periodic features can also be applied.

[0099] That is, we can obtain the acoustic features of real speech, which are denoted here as real acoustic features. Based on this, we can extract periodic features from the real acoustic features, which are denoted here as real periodic features. It is understandable that the method of extracting periodic features from real acoustic features is the same as the method of extracting periodic features from reconstructed acoustic features, and will not be elaborated here.

[0100] After obtaining the true periodic features, the true periodic features can be input into the discrimination model. The discrimination model will then determine whether the input true periodic features are true or false, thus obtaining the true or false discrimination result of the true periodic features output by the discrimination model.

[0101] Based on this, the authenticity judgment results of the real periodic features and the reconstructed periodic features can be combined to conduct adversarial training on the generative and discriminative models. During this process, the authenticity judgment results of the reconstructed periodic features reflect the difference between the reconstructed periodic features output by the generative model and the real periodic features. This allows us to determine the generation loss for the generative model, guiding it to iterate and update in a direction where the output periodic features are closer to the real periodic features. Furthermore, the authenticity judgment results of both the real and reconstructed periodic features reflect the accuracy of the discriminative model in judging the authenticity of the input periodic features. This allows us to determine the discriminative loss for the discriminative model, guiding it to iterate and update in a direction where the output is closer to the real-world authenticity judgment results.

[0102] Based on any of the above embodiments, in step 130, the authenticity judgment result based on the real periodic features and the authenticity judgment result based on the reconstructed periodic features are used to perform adversarial training on the generative model and the discriminative model, including:

[0103] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, a discrimination loss is determined, and the discrimination model is iterated based on the discrimination loss.

[0104] Based on the true / false discrimination results of the reconstructed periodic features, the generation loss is determined, and the generation model is iterated based on the generation loss.

[0105] Specifically, for the discriminant model, both the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features can characterize the accuracy of the discriminant model in judging the true / false nature of the input periodic features. Therefore, the discriminant loss for the discriminant model can be determined by combining the true / false discrimination results of the real periodic features and the reconstructed periodic features.

[0106] Here, the discriminant loss can be the difference between the true / false classification result of the real periodic feature and the true label, and the difference between the true / false classification result of the reconstructed periodic feature and the synthetic label; the sum of these two or the result of a weighted summation. After obtaining the discriminant loss, the parameters of the discriminant model can be iterated based on the discriminant loss, thereby gradually improving the discriminant model's ability to distinguish between true and false input periodic features.

[0107] For generative models, the authenticity determination results of reconstructed periodic features can characterize the realism of the reconstructed acoustic features output by the generative model in terms of timbre. Therefore, the generation loss for the generative model can be determined based on the authenticity determination results of the reconstructed periodic features.

[0108] Here, the generation loss can be the difference between the true / false discrimination result of the reconstructed periodic features and the true label. That is, the higher the probability that the reconstructed periodic features are identified as originating from real audio by the discrimination model, the smaller the generation loss; conversely, the lower the probability that the reconstructed periodic features are identified as originating from real audio by the discrimination model, the larger the generation loss. After obtaining the generation loss, the parameters of the generation model can be iterated based on the generation loss, thereby gradually improving the ability of the generation model to generate more natural and realistic reconstructed acoustic features.

[0109] Based on any of the above embodiments Figure 4 This is the second flowchart illustrating the model training method provided by this invention. For example... Figure 4 As shown, the method may include the following steps:

[0110] Step 410: Obtain the reconstructed acoustic features and the real acoustic features.

[0111] Here, both the reconstructed acoustic features and the true acoustic features can be Mel spectrograms. The true acoustic features can be obtained first, and then the corresponding text features can be input into the generative model to obtain the reconstructed acoustic features output by the generative model. The generative model here is the acoustic model to be trained.

[0112] Step 420: Convert the acoustic features into periodic features.

[0113] Periodic feature transformations are performed on both reconstructed and actual acoustic features. Specifically, for both reconstructed and actual acoustic features, the Mel spectrum needs to be inversely transformed into a linear amplitude spectrum, and then the linear amplitude spectrum is transformed into periodic features under different periodic components.

[0114] Step 430: Use the discrimination model corresponding to different periodic components to distinguish between true and false periodic features.

[0115] Different discriminant models can be pre-set for different periodic components, so that periodic features under different periodic components can be input into the corresponding discriminant models. The discriminant models can then determine the authenticity of the input periodic features and output the results.

[0116] Step 440: Calculate the adversarial loss and conduct adversarial training.

[0117] After obtaining the true / false discrimination results, adversarial losses can be calculated based on these results. Specifically, the adversarial losses include generation loss and discrimination loss. The generation loss is obtained based on the true / false discrimination results of the reconstructed periodic features, while the discrimination loss is obtained based on the true / false discrimination results of the real periodic features and the reconstructed periodic features. After obtaining the adversarial losses, the generative and discriminative models can be adversarially trained based on these losses.

[0118] Based on any of the above embodiments Figure 5 This is a flowchart illustrating the speech synthesis method provided by the present invention. Figure 5 As shown, the method includes:

[0119] Step 510: Input the text into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model;

[0120] The speech synthesis model includes the acoustic model trained using the above-mentioned model training method.

[0121] Specifically, the trained acoustic model can be obtained based on the model training methods provided in the above embodiments. Based on this, a speech synthesis model including the acoustic model can be constructed. For example, within the speech synthesis process, the input text can first be encoded into a text vector using a front-end text and prosody analysis model. Then, the acoustic features can be reconstructed from the text vector using a back-end acoustic model. Finally, the acoustic features can be restored into a speech waveform using a vocoder, thereby outputting synthesized speech.

[0122] Therefore, after obtaining the speech synthesis model, text can be input into the speech synthesis model, and the model will output synthesized speech with content consistent with the text. Since the speech synthesis model includes the acoustic model obtained based on the above embodiments, the naturalness and realism of the timbre in the synthesized speech are guaranteed.

[0123] In the method provided in this embodiment of the invention, speech synthesis is performed based on a speech synthesis model that includes an acoustic model with timbre reconstruction as the optimization objective, which can ensure the naturalness and realism of the synthesized speech timbre.

[0124] The model training apparatus provided by the present invention is described below. The model training apparatus described below and the model training method described above can be referred to in correspondence.

[0125] Figure 6 This is a schematic diagram of the model training device provided by the present invention. Figure 6 As shown, the device includes:

[0126] The feature extraction unit 610 is used to obtain the reconstructed acoustic features output by the generation model and extract the reconstruction periodicity features of the reconstructed acoustic features.

[0127] The authenticity discrimination unit 620 is used to input the reconstructed periodic features into the discrimination model and obtain the authenticity discrimination result of the reconstructed periodic features output by the discrimination model;

[0128] The adversarial training unit 630 is used to perform adversarial training on the generative model and the discriminative model based on the true / false discrimination results, and to use the trained generative model as the acoustic model.

[0129] In the apparatus provided in this embodiment of the invention, the reconstructed periodic features of the reconstructed acoustic features output by the generative model are input into the discriminative model for true / false discrimination. This allows for adversarial training between the generative model and the discriminative model. In this process, the periodic features that reflect timbre are explicitly used as optimization targets. As a result, when the generative model is applied to speech synthesis, the timbre of the synthesized speech is more natural and realistic.

[0130] Furthermore, the aforementioned training device involves two types of models: generative models and discriminative models. It has a simple structure, is easy to port, and can be applied to any acoustic model that targets Mel spectrograms or linear amplitude spectrograms to improve the timbre reconstruction capability of such acoustic models.

[0131] Based on any of the above embodiments, the feature extraction unit is specifically used for:

[0132] Extract the reconstructed periodic features of the reconstructed acoustic features under multiple periodic components;

[0133] The authenticity discrimination unit is specifically used for:

[0134] The reconstructed periodic features under each periodic component are input into the discrimination model corresponding to each periodic component, and the true and false discrimination results output by the discrimination model corresponding to each periodic component are obtained.

[0135] Based on any of the above embodiments, the plurality of periodic components are a plurality of periodic components based on prime numbers.

[0136] Based on any of the above embodiments, the feature extraction unit is specifically used for:

[0137] The reconstructed acoustic features are converted into a linear amplitude spectrum.

[0138] Reconstructed periodic features under multiple periodic components are extracted from the linear amplitude spectrum.

[0139] Based on any of the above embodiments, the adversarial training unit is specifically used for:

[0140] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, the generative model and the discriminative model are subjected to adversarial training.

[0141] The authenticity determination result of the real periodic feature is output by the discrimination model based on the real periodic feature, which is extracted from the real acoustic features.

[0142] Based on any of the above embodiments, the adversarial training unit is specifically used for:

[0143] Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, a discrimination loss is determined, and the discrimination model is iterated based on the discrimination loss.

[0144] Based on the true / false discrimination results of the reconstructed periodic features, the generation loss is determined, and the generation model is iterated based on the generation loss.

[0145] The speech synthesis apparatus provided by the present invention will be described below. The speech synthesis apparatus described below can be referred to in correspondence with the speech synthesis method described above.

[0146] Speech synthesis device, including:

[0147] A speech synthesis unit is used to input text into a speech synthesis model and obtain synthesized speech output by the speech synthesis model.

[0148] The speech synthesis model includes the acoustic model trained by the model training method provided in the above embodiments.

[0149] In the device provided in the embodiments of the present invention, speech synthesis is performed based on a speech synthesis model that includes an acoustic model with timbre reconstruction as the optimization objective, which can ensure the naturalness and realism of the synthesized speech timbre.

[0150] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a model training method, which includes:

[0151] Obtain the reconstructed acoustic features output by the generative model, and extract the reconstruction periodicity features of the reconstructed acoustic features;

[0152] The reconstructed periodic features are input into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model;

[0153] Based on the true / false discrimination results, the generative model and the discrimination model are subjected to adversarial training, and the resulting generative model is used as the acoustic model.

[0154] Alternatively, processor 710 may invoke logical instructions in memory 730 to execute a speech synthesis method, the method comprising:

[0155] Text is input into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model;

[0156] The speech synthesis model includes an acoustic model trained using a model training method.

[0157] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0158] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the model training method provided by the above methods, the method comprising:

[0159] Obtain the reconstructed acoustic features output by the generative model, and extract the reconstruction periodicity features of the reconstructed acoustic features;

[0160] The reconstructed periodic features are input into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model;

[0161] Based on the true / false discrimination results, the generative model and the discrimination model are subjected to adversarial training, and the resulting generative model is used as the acoustic model.

[0162] Alternatively, the computer can execute the speech synthesis methods provided by the above methods, which include:

[0163] Text is input into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model;

[0164] The speech synthesis model includes an acoustic model trained using a model training method.

[0165] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the model training methods provided by the methods described above, the method comprising:

[0166] Obtain the reconstructed acoustic features output by the generative model, and extract the reconstruction periodicity features of the reconstructed acoustic features;

[0167] The reconstructed periodic features are input into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model;

[0168] Based on the true / false discrimination results, the generative model and the discrimination model are subjected to adversarial training, and the resulting generative model is used as the acoustic model.

[0169] Alternatively, when the computer program is executed by a processor, it is implemented to perform the speech synthesis methods provided by the methods described above, the method comprising:

[0170] Text is input into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model;

[0171] The speech synthesis model includes an acoustic model trained using a model training method.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model training method, characterized in that, include: Obtain the reconstructed acoustic features output by the generative model, and extract the reconstructed periodic features of the reconstructed acoustic features. The input of the generative model is text, and the reconstructed acoustic features are Mel spectrograms. The reconstructed periodic features are input into the discrimination model to obtain the true or false discrimination result of the reconstructed periodic features output by the discrimination model; Based on the true / false discrimination results, the generative model and the discrimination model are subjected to adversarial training, and the resulting generative model is used as the acoustic model. The extraction of the reconstructed periodic features of the reconstructed acoustic features includes: converting the reconstructed acoustic features into a linear amplitude spectrum; extracting the reconstructed periodic features under multiple periodic components from the linear amplitude spectrum; wherein the multiple periodic components are multiple periodic components based on prime numbers.

2. The model training method according to claim 1, characterized in that, The step of inputting the reconstructed periodic features into the discrimination model to obtain the true / false discrimination result of the reconstructed periodic features output by the discrimination model includes: The reconstructed periodic features under each periodic component are input into the discrimination model corresponding to each periodic component, and the true and false discrimination results output by the discrimination model corresponding to each periodic component are obtained.

3. The model training method according to claim 1 or 2, characterized in that, The adversarial training of the generative model and the discriminative model based on the authenticity discrimination results includes: Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, the generative model and the discriminative model are subjected to adversarial training. The authenticity determination result of the real periodic feature is output by the discrimination model based on the real periodic feature, which is extracted from the real acoustic features.

4. The model training method according to claim 3, characterized in that, The true / false discrimination results based on the real periodic features and the true / false discrimination results based on the reconstructed periodic features are used to perform adversarial training on the generative model and the discriminative model, including: Based on the true / false discrimination results of the real periodic features and the true / false discrimination results of the reconstructed periodic features, a discrimination loss is determined, and the discrimination model is iterated based on the discrimination loss. Based on the true / false discrimination results of the reconstructed periodic features, the generation loss is determined, and the generation model is iterated based on the generation loss.

5. A speech synthesis method, characterized in that, include: Text is input into the speech synthesis model to obtain the synthesized speech output by the speech synthesis model; The speech synthesis model includes the acoustic model trained by the model training method as described in any one of claims 1 to 4.

6. A model training device, characterized in that, include: The feature extraction unit is used to obtain the reconstructed acoustic features output by the generative model and extract the reconstruction periodicity features of the reconstructed acoustic features. The input of the generative model is text, and the reconstructed acoustic features are Mel spectrograms. The authenticity discrimination unit is used to input the reconstructed periodic features into the discrimination model and obtain the authenticity discrimination result of the reconstructed periodic features output by the discrimination model; The adversarial training unit is used to perform adversarial training on the generative model and the discriminative model based on the true / false discrimination results, and to use the trained generative model as the acoustic model. The extraction of the reconstructed periodic features of the reconstructed acoustic features includes: converting the reconstructed acoustic features into a linear amplitude spectrum; extracting the reconstructed periodic features under multiple periodic components from the linear amplitude spectrum; wherein the multiple periodic components are multiple periodic components based on prime numbers.

7. A speech synthesis device, characterized in that, include: A speech synthesis unit is used to input text into a speech synthesis model and obtain synthesized speech output by the speech synthesis model. The speech synthesis model includes the acoustic model trained by the model training method as described in any one of claims 1 to 4.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the model training method as described in any one of claims 1 to 4, or the speech synthesis method as described in claim 5.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model training method as described in any one of claims 1 to 4, or the speech synthesis method as described in claim 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the model training method as described in any one of claims 1 to 4, or the speech synthesis method as described in claim 5.

Citation Information

Patent Citations

  • Speech synthesis method and speech synthesis system

    CN112908294A

  • Speech synthesis method based on generative adversarial network

    CN113066475A