Acoustic Model Training and Speech Synthesis Method, Device, System and Storage Medium

By introducing a generative adversarial network architecture in acoustic model training, the acoustic model and discriminator are subject to adversarial training, the problem of poor performance of acoustic models in the existing technology is solved, more realistic and accurate acoustic information generation is achieved, and the performance of speech synthesis system is improved.

CN114299918BActive Publication Date: 2025-06-20DATABAKER (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111582248.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-06-20
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The existing acoustic models directly calculate the gap between predicted acoustic information and real acoustic information during training, resulting in poor performance and easy to generate bad examples.

Method used

The generative adversarial network architecture is adopted to conduct adversarial training on the acoustic model and discriminator, and optimize the real discriminant results and predicted discriminant results to improve the generation ability of the acoustic model.

Benefits of technology

It significantly improves the authenticity and accuracy of the acoustic information generated by the acoustic model, reduces the generation of bad cases, and improves the overall performance and user experience of the speech synthesis system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299918B_ABST
    Figure CN114299918B_ABST
Patent Text Reader

Abstract

The present invention provides an acoustic model training and speech synthesis method, device, system and storage medium. The training method includes: obtaining text information and initial true acoustic information, where the text information includes training text or a text feature sequence related to the training text, and the initial true acoustic information includes initial true speech or an initial true acoustic feature sequence related to the initial true speech; inputting the text information into an acoustic model to obtain initial predicted acoustic information output by the acoustic model, where the form of the initial predicted acoustic information is consistent with that of the initial true acoustic information; inputting the initial true acoustic information and the initial predicted acoustic information into a discriminator respectively to obtain a true discrimination result and a predicted discrimination result output by the discriminator; and performing adversarial training on the acoustic model and the discriminator at least based on the true discrimination result and the predicted discrimination result. This method can improve the performance of the trained acoustic model, enabling it to generate more accurate and real acoustic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing. Specifically, it relates to an acoustic model training method, apparatus, and system, a storage medium, a speech synthesis method, apparatus, and system, and a storage medium. Background Art

[0002] Speech synthesis technology is a technology that converts text information into voice information. Speech synthesis technology can provide speech synthesis services for a wide range of users and target applications. Speech synthesis systems are widely used nowadays.

[0003] Speech synthesis requires the use of an acoustic model to achieve text-to-speech conversion. Before using the acoustic model for speech synthesis, it is usually necessary to train the acoustic model.

[0004] When existing acoustic models are trained, the model parameters are adjusted by directly calculating the difference between the predicted acoustic information generated by the acoustic model and the true acoustic information. This training method is relatively simple, and the performance of the trained acoustic model is not good enough. Summary of the Invention

[0005] In order to at least partially solve the problems existing in the prior art, there is provided an acoustic model training method, apparatus, and system, a storage medium, a speech synthesis method, apparatus, and system, and a storage medium.

[0006] According to one aspect of the present invention, there is provided an acoustic model training method, including: obtaining text information and initial true acoustic information, where the text information includes training text or a text feature sequence related to the training text, and the initial true acoustic information includes initial true speech or an initial true acoustic feature sequence related to the initial true speech; inputting the text information into an acoustic model to obtain initial predicted acoustic information output by the acoustic model, where the initial predicted acoustic information is in the same form as the initial true acoustic information; inputting the initial true acoustic information and the initial predicted acoustic information into a discriminator respectively to obtain a true discrimination result and a predicted discrimination result output by the discriminator, where the true discrimination result corresponds to the initial true acoustic information, and the predicted discrimination result corresponds to the initial predicted acoustic information; and performing adversarial training on the acoustic model and the discriminator at least based on the true discrimination result and the predicted discrimination result.

[0007] Exemplarily, adversarial training of the acoustic model and the discriminator based at least on the true discrimination result and the predicted discrimination result includes: when the acoustic model is fixed, performing discriminator training operations, and when the discriminator is fixed, performing acoustic model training operations; wherein, the discriminator training operations include: calculating a true loss based on the true discrimination result; calculating a predicted loss based on the predicted discrimination result; calculating a discriminator loss based on the true loss and the predicted loss; optimizing the parameters of the discriminator based on the discriminator loss; wherein, the acoustic model training operations include: calculating an information loss based on the initial true acoustic information and the initial predicted acoustic information; calculating an adversarial loss based on the predicted discrimination result; calculating a generator loss based on the information loss and the adversarial loss; optimizing the parameters of the acoustic model based on the generator loss.

[0008] Exemplarily, the discriminator includes n sub-discriminators, where n is a positive integer greater than 1. Inputting the initial true acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the true discrimination result and the predicted discrimination result output by the discriminator includes: performing n - 1 downsampling operations on the initial true acoustic information respectively to obtain n - 1 groups of downsampled true acoustic information, wherein the downsampling scales of any two of the n - 1 downsampling operations are different; performing n - 1 downsampling operations on the initial predicted acoustic information respectively to obtain n - 1 groups of downsampled predicted acoustic information; inputting the initial true acoustic information and the n - 1 groups of downsampled true acoustic information into the n sub-discriminators one by one to obtain n sub-true discrimination results output by the n sub-discriminators, and the true discrimination result includes the n sub-true discrimination results; inputting the initial predicted acoustic information and the n - 1 groups of downsampled predicted acoustic information into the n sub-discriminators one by one to obtain n sub-predicted discrimination results output by the n sub-discriminators, and the predicted discrimination result includes the n sub-predicted discrimination results.

[0009] Exemplarily, calculating the true loss based on the true discrimination result includes:

[0010] Calculating the true loss real-loss through the following formula:

[0011] real-loss = E s [max(0, 1 - D k (s))], k = 1, 2, 3…n;

[0012] Calculating the predicted loss based on the predicted discrimination result includes:

[0013] Calculating the predicted loss fake-loss through the following formula:

[0014] fake-loss = E x [max(0, 1 + D k(G(x))), k = 1, 2, 3…n;

[0015] Calculating the adversarial loss based on the prediction discrimination result includes:

[0016] Calculating the adversarial loss adv-loss through the following formula:

[0017] adv-loss = E x [-D k (G(x))), k = 1, 2, 3…n;

[0018] where D k represents the k-th sub-discriminator among n sub-discriminators, s represents the initial true acoustic information, x represents the text information, G represents the acoustic model, G(x) represents the initial predicted acoustic information, D k (s) represents the sub-true discrimination result corresponding to the k-th sub-discriminator, D k (G(x)) represents the sub-predicted discrimination result corresponding to the k-th sub-discriminator.

[0019] Exemplarily, the i-th downsampling operation among n - 1 downsampling operations is used to downsample the corresponding acoustic information by 2i times, i = 1, 2, 3……n - 1.

[0020] Exemplarily, calculating the discriminator loss based on the true loss and the prediction loss includes: weighted summing the true loss and the prediction loss to obtain the discriminator loss; and / or calculating the generator loss based on the information loss and the adversarial loss includes: weighted summing the information loss and the adversarial loss to obtain the generator loss.

[0021] Exemplarily, calculating the information loss based on the initial true acoustic information and the initial predicted acoustic information includes: substituting the initial true acoustic information and the initial predicted acoustic information into the mean square error function or the squared absolute error function to calculate the information loss.

[0022] Exemplarily, the discriminator includes n sub-discriminators, where n is a positive integer greater than 1. Inputting the initial real acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the real discrimination result and the predicted discrimination result output by the discriminator includes: performing n - 1 downsampling operations on the initial real acoustic information respectively to obtain n - 1 groups of downsampled real acoustic information, where any two of the n - 1 downsampling operations have different downsampling scales; performing n - 1 downsampling operations on the initial predicted acoustic information respectively to obtain n - 1 groups of downsampled predicted acoustic information; inputting the initial real acoustic information and the n - 1 groups of downsampled real acoustic information into the n sub-discriminators one by one to obtain n sub-real discrimination results output by the n sub-discriminators, and the real discrimination result includes the n sub-real discrimination results; inputting the initial predicted acoustic information and the n - 1 groups of downsampled predicted acoustic information into the n sub-discriminators one by one to obtain n sub-predicted discrimination results output by the n sub-discriminators, and the predicted discrimination result includes the n sub-predicted discrimination results.

[0023] According to another aspect of the present invention, there is also provided a speech synthesis method, including: obtaining the text to be synthesized; using the acoustic model trained by the above acoustic model training method to perform speech synthesis on the text to be synthesized to obtain the target speech.

[0024] According to another aspect of the present invention, there is also provided an acoustic model training device, including: an acquisition module for acquiring text information and initial real acoustic information, where the text information includes training text or a text feature sequence related to the training text, and the initial real acoustic information includes initial real speech or an initial real acoustic feature sequence related to the initial real speech; a first input module for inputting the text information into the acoustic model to obtain the initial predicted acoustic information output by the acoustic model, where the initial predicted acoustic information is in the same form as the initial real acoustic information; a second input module for inputting the initial real acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the real discrimination result and the predicted discrimination result output by the discriminator, where the real discrimination result corresponds to the initial real acoustic information, and the predicted discrimination result corresponds to the initial predicted acoustic information; and a training module for performing adversarial training on the acoustic model and the discriminator at least based on the real discrimination result and the predicted discrimination result.

[0025] According to another aspect of the present invention, there is also provided a speech synthesis device, including: an acquisition module for acquiring the text to be synthesized; a synthesis module for using the acoustic model trained by the above acoustic model training method to perform speech synthesis on the text to be synthesized to obtain the target speech.

[0026] According to another aspect of the present invention, there is also provided an acoustic model training system, including a processor and a memory. Among them, computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above-mentioned acoustic model training method.

[0027] According to another aspect of the present invention, there is also provided a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned acoustic model training method.

[0028] According to another aspect of the present invention, there is also provided a speech synthesis system, including a processor and a memory. Among them, computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the above-mentioned speech synthesis method.

[0029] According to another aspect of the present invention, there is also provided a storage medium, on which program instructions are stored, and when the program instructions are run, they are used to execute the above-mentioned speech synthesis method.

[0030] For the acoustic model training method, device, system and storage medium according to the embodiments of the present invention, and the speech synthesis method, device, system and storage medium, the acoustic model is regarded as a generator, and the acoustic model is trained based on the generative adversarial network architecture. The acoustic model trained in this way can generate more accurate and clear acoustic information, can effectively improve the authenticity of the generated acoustic information, and is closer to the real acoustic information to a greater extent. Therefore, this training method can significantly reduce the bad cases of the acoustic model generating acoustic information. Further, in the case of using the above acoustic model for speech synthesis, higher-quality speech can be generated, which helps to improve the overall performance and user experience of the speech synthesis system.

[0031] A series of simplified concepts are introduced in the summary of the invention, which will be further described in detail in the detailed implementation section. The summary of the invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0032] The following will detail the advantages and features of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The following drawings of the present invention are hereby incorporated as part of the present invention to understand the present invention. The embodiments and descriptions of the present invention are shown in the drawings to explain the principles of the present invention. In the drawings,

[0034] Figure 1 A schematic flowchart showing an acoustic model training method according to an embodiment of the present invention;

[0035] Figure 2Shows a schematic flowchart of acoustic model training according to an embodiment of the present invention;

[0036] Figure 3 Shows a schematic flowchart of a speech synthesis method according to an embodiment of the present invention

[0037] Figure 4 Shows a schematic block diagram of an acoustic model training device according to an embodiment of the present invention;

[0038] Figure 5 Shows a schematic block diagram of an acoustic model training system according to an embodiment of the present invention;

[0039] Figure 6 Shows a schematic block diagram of a speech synthesis device according to an embodiment of the present invention;

[0040] Figure 7 Shows a schematic block diagram of a speech synthesis system according to an embodiment of the present invention; Detailed implementation manners

[0041] In the following description, a large number of details are provided to enable a thorough understanding of the present invention. However, those skilled in the art can understand that the following description only exemplarily shows the preferred embodiments of the present invention, and the present invention can be implemented without one or more of such details. In addition, to avoid confusion with the present invention, some well-known technical features in the art are not described in detail.

[0042] To at least partially solve the above technical problems, embodiments of the present invention provide an acoustic model training method and device. In the present invention, the acoustic model is regarded as a generator and forms a generative adversarial network with a discriminator for adversarial training. Through this adversarial training, the performance of the acoustic model can be effectively improved, making the converted acoustic information more real and accurate. When the acoustic model trained by the acoustic model training method is applied to speech synthesis, the above solution helps to further improve the accuracy of subsequent speech synthesis, thereby greatly improving the user experience of the speech synthesis system.

[0043] A Generative Adversarial Network (GAN) is a method of unsupervised learning that learns by having two neural networks play against each other. A generative adversarial network consists of a generative network (i.e., a generator) and a discriminative network (i.e., a discriminator). The generative network randomly samples from the latent space as input, and its output results need to mimic the real samples in the training set as much as possible. The input of the discriminative network is either real samples or the output of the generative network, and its purpose is to distinguish the output of the generative network from the real samples as much as possible. And the generative network has to deceive the discriminative network as much as possible. The two networks confront each other and continuously adjust their parameters. The ultimate goal is to make the discriminative network unable to determine whether the output results of the generative network are real. Generative adversarial networks are mostly used in the field of images. The inventors of the present invention creatively thought of applying it to the training of acoustic models. The following describes the acoustic model training method based on generative adversarial networks.

[0044] According to one aspect of the present invention, an acoustic model training method is provided. Figure 1 The schematic flowchart of an acoustic model training method 100 according to an embodiment of the present invention is shown. As Figure 1 shown, the acoustic model training method 100 includes steps S110, S120, S130, and S140.

[0045] In step S110, text information and initial real acoustic information are obtained. The text information includes training text or a text feature sequence related to the training text, and the initial real acoustic information includes initial real speech or an initial real acoustic feature sequence related to the initial real speech.

[0046] An exemplary process of traditional speech synthesis is as follows: first, the text is converted into a text feature sequence, then the text feature sequence is converted into an acoustic feature sequence through an acoustic model, and then the acoustic feature sequence is converted into a speech waveform. In this case, the acoustic model is a feature conversion model that generates acoustic features based on text features. In addition, speech synthesis can also be achieved through an end-to-end acoustic model. In this case, it is not necessary to separately extract the text feature sequence, nor is it necessary to separately generate the speech waveform based on the acoustic feature sequence. The acoustic models required for the above different synthesis methods can all be trained by the acoustic model training method 100 provided by the present invention.

[0047] In one example, the acoustic model described herein may be a feature conversion model that generates acoustic features based on text features, that is, its input is text features (specifically, a sequence of text features), and its output is acoustic features (specifically, a sequence of acoustic features). In this case, the text information may be or include a sequence of text features related to the training text, and the initial true acoustic information may be or include an initial true sequence of acoustic features related to the initial true speech. In another example, the acoustic model may be an end-to-end speech synthesis model, that is, its input is text and its output is speech synthesized based on the text. In this case, the text information may be or include the training text, and the initial true acoustic information may be or include the initial true speech. Of course, optionally, the acoustic model may also be set to implement other conversion functions, and thus have other different input-output combinations. For example, the acoustic model may be used to convert the text features herein into speech, in which case the text information may be or include a sequence of text features related to the training text, and the initial true acoustic information may be or include the initial true speech. Another example is that the acoustic model may also be used to convert text into acoustic features, in which case the text information may be or include the training text, and the initial true acoustic information may be or include an initial true sequence of acoustic features related to the initial true speech.

[0048] The sequence of text features being related to the training text means that there is a corresponding relationship between the two, and the semantic content expressed by the two is the same. In this document, a sequence of text features can be generated based on the training text. For example, a sequence of text features related to the training text can be obtained by performing text analysis on the training text. Exemplarily, the text analysis may include operations such as text regularization, word segmentation, part-of-speech prediction, disambiguation of polyphonic characters, prosody prediction, etc. The above text analysis can be implemented using any existing or future possible text analysis methods, and the present invention does not limit this.

[0049] The sequence of text features may include text features corresponding one-to-one to multiple frames. Those skilled in the art can understand the meaning of "frame" in the field of speech processing and its division method, and can also understand the meaning of "text features" and the information it contains. This document will not elaborate.

[0050] The initial true sequence of acoustic features being related to the initial true speech means that there is a corresponding relationship between the two, and the semantic content expressed by the two is the same. In this document, the initial true speech can be generated based on the initial true sequence of acoustic features.

[0051] The various sequences of acoustic features described herein (such as the initial true sequence of acoustic features, the initial predicted sequence of acoustic features, the downsampled true sequence of acoustic features, the downsampled predicted sequence of acoustic features, etc.) can all include acoustic features corresponding one-to-one to multiple frames. Those skilled in the art can understand the meaning of "acoustic features" in the field of speech processing and the information it contains. This document will not elaborate.

[0052] Exemplarily and non - restrictively, the text features described herein may include text feature information such as phonetic symbols and prosody. Exemplarily and non - restrictively, the acoustic features described herein may include acoustic feature information such as Mel - Frequency Cepstral Coefficients (MFCC) and fundamental frequency (F0). The above description is only an example and not a limitation of the present invention, and any suitable existing or future - emerging text features and acoustic features adopted in the field of speech should fall within the protection scope of the present invention.

[0053] Optionally, the initial true acoustic information may be acoustic information corresponding to the text information, that is, the semantic content expressed by the initial true acoustic information is consistent with the semantic content expressed by the text information. In this case, the initial true acoustic information can be understood as the annotation data (ground truth) of the text information. However, it should be understood that the above - mentioned embodiments are only examples and not limitations of the present invention, and the present invention is not limited to this implementation scheme. For example, the semantic content expressed by the initial true acoustic information may also be different from the semantic content expressed by the text information.

[0054] In step S120, the text information is input into the acoustic model to obtain the initial predicted acoustic information output by the acoustic model, where the initial predicted acoustic information is in the same form as the initial true acoustic information.

[0055] In the case where the initial true acoustic information is or includes an initial true acoustic feature sequence, the initial predicted acoustic information is or includes an initial predicted acoustic feature sequence. In the case where the initial true acoustic information is or includes an initial true speech, the initial predicted acoustic information is or includes an initial predicted speech. That is to say, the form of the initial predicted acoustic information is consistent with that of the initial true acoustic information.

[0056] The acoustic model is used to convert text information into corresponding acoustic information. The acoustic model described herein can be any suitable existing or future - emerging acoustic model that can be used for speech synthesis, and the present invention does not limit this.

[0057] Exemplarily and non - restrictively, the acoustic model described herein may include one or more of the following network models: Deep Neural Networks (DNN), FastSpeech1 / 2 model, tacotron1 / 2 model, etc.

[0058] The acoustic model adopted in step S120 may be an initialized or certain - trained acoustic model, and the present invention does not limit this.

[0059] Input text information into an acoustic model, which can output corresponding initial predicted acoustic information. The acoustic model can be regarded as the generator in a generative adversarial network, and its purpose is to make the output initial predicted acoustic information as close as possible to the real acoustic information corresponding to the text feature sequence.

[0060] In step S130, input the initial real acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the real discrimination result and the predicted discrimination result output by the discriminator. The real discrimination result corresponds to the initial real acoustic information, and the predicted discrimination result corresponds to the initial predicted acoustic information.

[0061] The discriminator can be implemented using any suitable existing or future possible discriminator model, and the present invention does not limit this. Exemplarily rather than restrictively, the discriminator can be implemented using a convolutional neural network (CNN), a deep convolutional neural network (DNN), etc.

[0062] The discriminator is used to judge whether the input acoustic information is real or fake (fake means generated by the acoustic model). Input the initial real acoustic information into the discriminator to obtain the corresponding real discrimination result. Input the initial predicted acoustic information into the discriminator to obtain the corresponding predicted discrimination result.

[0063] The discriminator can be implemented using a single network structure. Of course, the discriminator can also be implemented using a multiple network structure. For example, the discriminator can include multiple sub-discriminators, which will be described below.

[0064] In step S140, perform adversarial training on the acoustic model and the discriminator based at least on the real discrimination result and the predicted discrimination result.

[0065] The adversarial training can be implemented using a conventional adversarial training method. Those skilled in the art can understand that the adversarial training can be implemented using an alternating training method, that is, the acoustic model and the discriminator can train the other while fixing one of them, and the two are alternately trained until a preset requirement is met. The preset requirement can be, for example, that the entire generative adversarial network converges to a Nash equilibrium point.

[0066] Exemplarily, a real loss can be calculated based on the real discrimination result, a predicted loss can be calculated based on the predicted discrimination result, and the two losses can be combined to obtain the total discriminator loss, and the discriminator can be trained based on the discriminator loss.

[0067] Exemplarily, an adversarial loss can be calculated based on the prediction discrimination result, and the acoustic model can be trained at least based on the adversarial loss. In one example, the acoustic model can be trained solely based on the adversarial loss, which can be implemented whether the initial true acoustic information corresponds to the text information (i.e., the expressed speech content is consistent) or not (i.e., the expressed speech content is inconsistent). In another example, an information loss can also be calculated based on the true acoustic information and the predicted acoustic information, and the total generator loss can be obtained by combining the information loss and the adversarial loss. Subsequently, the acoustic model can be trained based on the generator loss. The second example can be implemented when the initial true acoustic information corresponds to the text information.

[0068] The scheme of training the generator or discriminator based on the loss can be implemented by the backpropagation algorithm. Those skilled in the art can understand the implementation method of training the generator or discriminator based on the loss, which will not be elaborated herein.

[0069] The existing acoustic model training methods are relatively simple, and the performance of the trained acoustic models is poor, and there are very likely to be some random bad cases. However, according to the acoustic model training method of the embodiments of the present invention, the acoustic model is regarded as a generator, and the acoustic model is trained based on the generative adversarial network architecture. The acoustic model trained in this way can generate more accurate and clear acoustic information, can effectively improve the authenticity of the generated acoustic information, and is closer to the true acoustic information to a greater extent. Therefore, this training method can greatly reduce the bad cases of the acoustic model generating acoustic information. Further, in the case of using the above acoustic model for speech synthesis, higher-quality speech can be generated, which helps to improve the overall performance and user experience of the speech synthesis system.

[0070] According to the embodiments of the present invention, the adversarial training (step S140) of the acoustic model and the discriminator based at least on the true discrimination result and the prediction discrimination result may include: performing a discriminator training operation when the acoustic model is fixed, and performing an acoustic model training operation when the discriminator is fixed; wherein, the discriminator training operation includes: calculating a true loss based on the true discrimination result; calculating a prediction loss based on the prediction discrimination result; calculating a discriminator loss based on the true loss and the prediction loss; optimizing (i.e., updating) the parameters of the discriminator based on the discriminator loss; wherein, the acoustic model training operation includes: calculating an information loss based on the initial true acoustic information and the initial predicted acoustic information; calculating an adversarial loss based on the prediction discrimination result; calculating a generator loss based on the information loss and the adversarial loss; optimizing (i.e., updating) the parameters of the acoustic model based on the generator loss.

[0071] As described above, in the case where the initial true acoustic information corresponds to the text information, it is possible to further calculate an information loss based on the initial true acoustic information and the initial predicted acoustic information, and further train the acoustic model in combination with this information loss.

[0072] When a conventional generative adversarial network is trained, consistency is not usually required between the true samples and the false samples used. The false samples are usually generated by superimposing noise on certain images. Due to this noise-based generation method, when calculating the loss of the generator in the existing generative adversarial network, usually only the loss brought by the predicted results of the false samples is calculated.

[0073] Although the acoustic model of the present invention is regarded as a generator, it is different from a conventional generator. The initial predicted acoustic information generated by it is not a simple "false sample", but itself is a prediction of the corresponding acoustic information of the text information. Therefore, in the case where the initial true acoustic information corresponds to the text information, there is a consistency relationship between the initial true acoustic information and the initial predicted acoustic information, that is, the more consistent the two are, the better. Therefore, in this case, an information loss can be calculated based on the initial true acoustic information and the initial predicted acoustic information.

[0074] Compared with the traditional acoustic model training method, the acoustic model trained by the above-mentioned adversarial training method based on multiple losses can generate more accurate and real acoustic features. And compared with the training of the traditional generative adversarial network, the loss information obtained by the above-mentioned training scheme that combines multiple losses (especially adding information loss) is richer, which helps to improve the robustness of the entire generative adversarial network.

[0075] According to an embodiment of the present invention, the discriminator includes n sub-discriminators, where n is a positive integer greater than 1. Inputting the initial true acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the true discrimination result and the predicted discrimination result output by the discriminator (step S130) may include: performing n - 1 downsampling operations on the initial true acoustic information respectively to obtain n - 1 groups of downsampled true acoustic information respectively, where the downsampling scales of any two of the n - 1 downsampling operations are different; performing n - 1 downsampling operations on the initial predicted acoustic information respectively to obtain n - 1 groups of downsampled predicted acoustic information respectively; inputting the initial true acoustic information and the n - 1 groups of downsampled true acoustic information into the n sub-discriminators one by one to obtain n sub-true discrimination results output by the n sub-discriminators, and the true discrimination result includes the n sub-true discrimination results; inputting the initial predicted acoustic information and the n - 1 groups of downsampled predicted acoustic information into the n sub-discriminators one by one to obtain n sub-predicted discrimination results output by the n sub-discriminators, and the predicted discrimination result includes the n sub-predicted discrimination results.

[0076] In the case where the initial true acoustic information is or includes an initial true acoustic feature sequence, the n-1 sets of downsampled true acoustic information can be or include n-1 downsampled true acoustic feature sequences, and the n-1 sets of downsampled predicted acoustic information can be or include n-1 downsampled predicted acoustic feature sequences.

[0077] The discriminator may include multiple sub-discriminators, each sub-discriminator being respectively used to determine the authenticity of acoustic information with different sampling scales. Optionally, these sub-discriminators may have the same network structure but operate on different sampling scales of the acoustic information. Any two different sub-discriminators may respectively operate on different sampling scales. The sampling scale described herein may be understood as the sampling rate. The downsampling scale described herein may be understood as the downsampling factor or the sampling rate after downsampling.

[0078] For example, use D k to represent the k-th sub-discriminator among the n sub-discriminators, where k = 1, 2, 3... n. Then, D1 can operate on the sampling scale of the original acoustic information (such as the initial true acoustic information and the initial predicted acoustic information), while D2, D3... D n operate on the sampling scales after the original acoustic information is respectively downsampled. Exemplarily and non-limitingly, the initial true acoustic information and the initial predicted acoustic information may have the same sampling scale as their respective source audio. And, the initial true acoustic information and the initial predicted acoustic information have the same sampling scale as each other.

[0079] Figure 2 FIG. shows a schematic flow diagram of acoustic model training according to an embodiment of the present invention. It should be noted that although in Figure 2 it is shown that the text information is a text feature sequence and the initial true acoustic information is an initial true acoustic feature sequence, as described above, this is only an example and not a limitation of the present invention.

[0080] See Figure 2, showing multiple sub-discriminators, namely sub-discriminator 1, sub-discriminator 2... sub-discriminator n. Although not shown, it can be understood that before inputting into sub-discriminators 2 to n, the initial predicted acoustic information and the initial true acoustic information can be downsampled to obtain their respective corresponding downsampled acoustic information. Of course, the above downsampling scheme is only an example and not a limitation to the present invention. For example, the downsampling operation required for each sub-discriminator can also be implemented inside the sub-discriminator, that is, each of sub-discriminators 2 to n can include its own required downsampling layer for implementing the corresponding downsampling operation. Except for the downsampling layer, the remaining network layers included in each of sub-discriminators 1 to n (sub-discriminator 1 is all network layers) can have the same network structure. Of course, optionally, in the scheme of implementing the downsampling operation outside the sub-discriminator, any two of sub-discriminators 1 to n can have different network structures from each other. And in the scheme of implementing the downsampling operation inside the sub-discriminator, except for the downsampling layer, the remaining network layers included in any two of sub-discriminators 1 to n can also have different network structures from each other.

[0081] The high-frequency acoustic information is relatively sparse and is difficult to be effectively learned during training. By using multiple sub-discriminators to provide feedback on the acoustic information at different sampling scales, it can better assist in the training of the acoustic model and better learn the high-frequency features of the acoustic information. Therefore, the acoustic model trained in this way of training can generate a clearer acoustic spectrogram, which helps to greatly reduce the background noise of the finally synthesized sound.

[0082] According to an embodiment of the present invention, the i-th downsampling operation among the n - 1 downsampling operations is used to downsample the corresponding acoustic information by 2^i times, where i = 1, 2, 3... n - 1.

[0083] It can be understood that when performing n - 1 downsampling operations on the initial true acoustic information respectively, the i-th downsampling operation among the n - 1 downsampling operations is used to downsample the initial true acoustic information by 2^i times. When performing n - 1 downsampling operations on the initial predicted acoustic information respectively, the i-th downsampling operation among the n - 1 downsampling operations is used to downsample the initial predicted acoustic information by 2^i times.

[0084] The setting of the downsampling scale in this embodiment is only an example and not a limitation to the present invention. The downsampling scale of each downsampling operation can be set to any appropriate downsampling scale according to needs.

[0085] According to an embodiment of the present invention, the discriminator includes n sub-discriminators. The n sub-discriminators include an original sub-discriminator and n-1 downsampling sub-discriminators, where n is a positive integer greater than 1. Inputting the initial true acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the true discrimination result and the predicted discrimination result output by the discriminator (step S130) may include: inputting the initial true acoustic information into the n sub-discriminators respectively to obtain n sub-true discrimination results output by the n sub-discriminators, and the true discrimination result includes the n sub-true discrimination results; and inputting the initial predicted acoustic information into the n sub-discriminators respectively to obtain n sub-predicted discrimination results output by the n sub-discriminators, and the predicted discrimination result includes the n sub-predicted discrimination results; wherein, inputting the initial true acoustic information into the n sub-discriminators respectively to obtain the n sub-true discrimination results output by the n sub-discriminators includes: when the current sub-discriminator is a downsampling sub-discriminator, inputting the initial true acoustic information into the downsampling layer of the current sub-discriminator for downsampling operation to obtain the downsampled true acoustic information corresponding to the current sub-discriminator, wherein the downsampling scales of any two of the n-1 downsampling sub-discriminators are different; inputting the downsampled true acoustic information into the remaining network layer of the current sub-discriminator to obtain the sub-true discrimination result corresponding to the current sub-discriminator; wherein, inputting the initial predicted acoustic information into the n sub-discriminators respectively to obtain the n sub-predicted discrimination results output by the n sub-discriminators includes: when the current sub-discriminator is a downsampling sub-discriminator, inputting the initial predicted acoustic information into the downsampling layer of the current sub-discriminator for downsampling operation to obtain the downsampled predicted acoustic information corresponding to the current sub-discriminator; inputting the downsampled predicted acoustic information into the remaining network layer of the current sub-discriminator to obtain the sub-predicted discrimination result corresponding to the current sub-discriminator.

[0086] According to an embodiment of the present invention, the downsampling layer of the i-th downsampling sub-discriminator among the n-1 downsampling sub-discriminators is used to downsample the corresponding acoustic information by 2i times, where i = 1, 2, 3... n-1.

[0087] The embodiment of implementing downsampling inside the sub-discriminator has been described above. This embodiment can be understood by referring to the above description and will not be elaborated here.

[0088] According to an embodiment of the present invention, calculating the true loss based on the true discrimination result includes:

[0089] Calculating the true loss real-loss through the following formula:

[0090] real-loss = E s [max(0, 1 - D k (s))], k = 1, 2, 3... n Formula (1)

[0091] Calculating the prediction loss based on the prediction discrimination result includes:

[0092] Calculate the prediction loss fake-loss through the following formula:

[0093] fake-loss = E x [max(0, 1 + D k (G(x))), k = 1, 2, 3…n Formula (2)

[0094] Calculating the adversarial loss based on the prediction discrimination result includes:

[0095] Calculate the adversarial loss adv-loss through the following formula:

[0096] adv-loss = E x [-D k (G(x))], k = 1, 2, 3…n Formula (3)

[0097] Among them, D k represents the k-th sub-discriminator among n sub-discriminators, s represents the initial true acoustic information, x represents the text information, G represents the acoustic model, G(x) represents the initial predicted acoustic information, D k (s) represents the sub-true discrimination result corresponding to the k-th sub-discriminator, D k (G(x)) represents the sub-prediction discrimination result corresponding to the k-th sub-discriminator.

[0098] In the case of training with multiple sub-discriminators, the discrimination results of multiple sub-discriminators can be integrated through the above formulas (1)-(3), and the true loss, prediction loss, and adversarial loss can be calculated based on the integrated results.

[0099] In each formula in this article, E represents taking the average. Formula (1) represents that for each sub-discriminator, taking the maximum value between 0 and (1 - D k (s)), and taking the average of all the maximum value results (a total of n maximum value results) calculated respectively in the case of k = 1, 2, 3…n.

[0100] Formula (2) represents that for each sub-discriminator, taking the maximum value between 0 and (1 + D k (G(x))), and taking the average of all the maximum value results (a total of n maximum value results) calculated respectively in the case of k = 1, 2, 3…n.

[0101] Formula (3) represents taking the average of -D k (G(x)).

[0102] According to an embodiment of the present invention, calculating the information loss based on the initial true acoustic information and the initial predicted acoustic information includes: substituting the initial true acoustic information and the initial predicted acoustic information into a mean squared error (MSE) function or a mean absolute error (MAE) function to calculate the information loss.

[0103] In one example, the information loss mel-loss can be calculated by the following formula:

[0104] mel-loss = E (s,x) [(s - G(x)) 2 Formula (4)

[0105] In another example, the information loss mel-loss can be calculated by the following formula:

[0106] mel-loss = E (s,x) [|s - G(x)|] Formula (5)

[0107] According to an embodiment of the present invention, calculating the discriminator loss based on the true loss and the predicted loss includes: weighted summing the true loss and the predicted loss to obtain the discriminator loss; and / or, calculating the generator loss based on the information loss and the adversarial loss includes: weighted summing the information loss and the adversarial loss to obtain the generator loss.

[0108] Exemplarily, the generator loss g-loss can be calculated by the following formula:

[0109] g-loss = mel-loss + adv-loss Formula (6)

[0110] Exemplarily, the discriminator loss d-loss can be calculated by the following formula:

[0111] d-loss = real-loss + fake-loss Formula (7)

[0112] In Formulas (6) and (7), the weight of each loss is 1, but this is not a limitation of the present invention. For example, other appropriate weights can be set for the true loss, the predicted loss, the information loss, and the adversarial loss as needed.

[0113] Figure 2 Examples of calculating the generator loss and the discriminator loss are shown and can be referred to Figure 2 to understand the above embodiments.

[0114] According to another aspect of the present invention, a speech synthesis method is further provided. Figure 3 A schematic flowchart of a speech synthesis method 300 according to an embodiment of the present invention is shown. As Figure 3As shown in the figure, the speech synthesis method 300 includes steps S310 and S320.

[0115] In step S310, the text to be synthesized is obtained.

[0116] In step S320, the text to be synthesized is subjected to speech synthesis by using the acoustic model obtained by training through the above acoustic model training method 100 to obtain the target speech.

[0117] Exemplarily, the acoustic model may be a feature conversion model that generates acoustic features based on text features. At this time, step S320 may include: performing text analysis on the text to be synthesized to obtain the text feature sequence to be synthesized; inputting the text feature sequence to be synthesized into the acoustic model obtained by training through the above acoustic model training method 100 to obtain the target acoustic feature sequence output by the acoustic model; and inputting the target acoustic feature sequence into a vocoder to obtain the target speech.

[0118] Referring to the above, the speech synthesis process can be divided into a front-end process and a back-end process. The front-end process may include performing text analysis on the text to be synthesized. For example, performing regularization, word segmentation, part-of-speech prediction, disambiguation of polyphonic characters, prosody prediction, etc. of the text. The front-end process may adopt a unified front-end model for the same language.

[0119] The back-end process may include using the acoustic model and the text analysis result to perform speech synthesis to obtain the speech data corresponding to the text to be synthesized. For example, inputting information such as word segmentation, phonetic notation, and prosody into the acoustic model of the target object, the acoustic features corresponding to the speech, such as spectral envelope, fundamental frequency, duration, etc. information can be obtained. The acoustic features are the features reflecting the timbre of each object, and the acoustic models of different objects are usually different, and the obtained acoustic features are usually also different. Subsequently, the acoustic features (target acoustic feature sequence) can be input into a vocoder to obtain the final waveform file (i.e., speech data).

[0120] As described above, the acoustic model obtained by training according to the acoustic model training method of the embodiment of the present invention can generate more accurate and clear acoustic information, can effectively improve the authenticity of the generated acoustic information, and is closer to the real acoustic information to a greater extent. In the case of performing speech synthesis by using the above acoustic model, higher-quality speech can be generated, which helps to improve the overall performance and user experience of the speech synthesis system.

[0121] According to another aspect of the present invention, an acoustic model training device is provided. Figure 4 The schematic block diagram of an acoustic model training device 400 according to an embodiment of the present invention is shown.

[0122] As Figure 4As shown, the acoustic model training device 400 according to an embodiment of the present invention includes an acquisition module 410, a first input module 420, a second input module 430, and a training module 440. Each of the modules can respectively execute the various steps / functions of the acoustic model training method 100 described above in conjunction with Figure 1-2 the description of the acoustic model training method 100. Only the main functions of the components of the acoustic model training device 400 will be described below, while omitting the details already described above.

[0123] The acquisition module 410 is configured to acquire text information and initial true acoustic information. The text information includes training text or a text feature sequence related to the training text, and the initial true acoustic information includes initial true speech or an initial true acoustic feature sequence related to the initial true speech.

[0124] The first input module 420 is configured to input the text information into the acoustic model to obtain initial predicted acoustic information output by the acoustic model, where the initial predicted acoustic information is in the same form as the initial true acoustic information.

[0125] The second input module 430 is configured to input the initial true acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain a true discrimination result and a predicted discrimination result output by the discriminator. The true discrimination result corresponds to the initial true acoustic information, and the predicted discrimination result corresponds to the initial predicted acoustic information.

[0126] The training module 440 is configured to perform adversarial training on the acoustic model and the discriminator at least based on the true discrimination result and the predicted discrimination result.

[0127] According to another aspect of the present invention, an acoustic model training system is provided. Figure 5 The schematic block diagram of an acoustic model training system 500 according to an embodiment of the present invention is shown. The acoustic model training system 500 includes a processor 510 and a memory 520.

[0128] The memory 520 stores computer program instructions for implementing the corresponding steps in the acoustic model training method 100 according to an embodiment of the present invention.

[0129] The processor 510 is configured to run the computer program instructions stored in the memory 520 to execute the corresponding steps of the acoustic model training method 100 according to an embodiment of the present invention.

[0130] According to another aspect of the present invention, there is provided a storage medium on which program instructions are stored. When the program instructions are run by a computer or a processor, they are used to execute the corresponding steps of the acoustic model training method 100 according to an embodiment of the present invention, and are used to implement the corresponding modules in the acoustic model training device 400 according to an embodiment of the present invention. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0131] According to another aspect of the present invention, there is provided a speech synthesis device. Figure 6 The schematic block diagram of a speech synthesis device 600 according to an embodiment of the present invention is shown.

[0132] As Figure 6 shown, the speech synthesis device 600 according to an embodiment of the present invention includes an acquisition module 610 and a synthesis module 620. Each of the modules can respectively execute the corresponding steps / functions of the speech synthesis method 400 described above in combination with Figure 4 The following only describes the main functions of the components of the speech synthesis device 600, and omits the details that have been described above.

[0133] The acquisition module 610 is used to acquire the text to be synthesized.

[0134] The analysis module 620 is used to perform speech synthesis on the text to be synthesized by using the acoustic model trained through the above acoustic model training method 100 to obtain the target speech.

[0135] According to another aspect of the present invention, there is provided a speech synthesis system. Figure 7 The schematic block diagram of a speech synthesis system 700 according to an embodiment of the present invention is shown. The speech synthesis system 700 includes a processor 710 and a memory 720.

[0136] The memory 720 stores computer program instructions for implementing the corresponding steps in the speech synthesis method 400 according to an embodiment of the present invention.

[0137] The processor 710 is used to run the computer program instructions stored in the memory 720 to execute the corresponding steps of the speech synthesis method 400 according to an embodiment of the present invention.

[0138] According to another aspect of the present invention, there is provided a storage medium on which program instructions are stored. When the program instructions are run by a computer or a processor, they are used to execute the corresponding steps of the speech synthesis method 400 according to the embodiments of the present invention, and are used to implement the corresponding modules in the speech synthesis device 600 according to the embodiments of the present invention. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0139] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0140] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0141] Similarly, it should be understood that in order to streamline the present invention and help understand one or more of the various aspects of the invention, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting the intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in being able to solve the corresponding technical problems with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present invention.

[0142] Those skilled in the art will appreciate that, except where features are mutually exclusive, any combination can be employed to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings), as well as all the processes or units of any method or device so disclosed. Each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose, unless expressly stated otherwise.

[0143] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that in practice, a microprocessor or a digital signal processor (DSP) can be used to implement some or all of the functions of some of the modules in the acoustic model training system or the speech synthesis system according to the embodiments of the present invention. The present invention can also be implemented as a device program (e.g., a computer program and a computer program product) for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0144] It should be noted that the above embodiments illustrate rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0145] As described above, this is only the specific implementation manner of the present invention or the description of the specific implementation manner, and the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all such changes or substitutions should be covered by the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An acoustic model training method, comprising: Obtain text information and initial true acoustic information, where the text information includes a training text or a text feature sequence related to the training text, and the initial true acoustic information includes initial true speech or an initial true acoustic feature sequence related to the initial true speech; Input the text information into an acoustic model to obtain initial predicted acoustic information output by the acoustic model, where the initial predicted acoustic information is in the same form as the initial true acoustic information; Input the initial true acoustic information and the initial predicted acoustic information into a discriminator respectively to obtain a true discrimination result and a predicted discrimination result output by the discriminator, where the true discrimination result corresponds to the initial true acoustic information, and the predicted discrimination result corresponds to the initial predicted acoustic information; and Perform adversarial training on the acoustic model and the discriminator based at least on the true discrimination result and the predicted discrimination result.

2. The method according to claim 1, wherein The performing adversarial training on the acoustic model and the discriminator based at least on the true discrimination result and the predicted discrimination result includes: When the acoustic model is fixed, perform discriminator training operations, and when the discriminator is fixed, perform acoustic model training operations; Wherein, the discriminator training operations include: Calculate a true loss based on the true discrimination result; Calculate a predicted loss based on the predicted discrimination result; Calculate a discriminator loss based on the true loss and the predicted loss; Optimize the parameters of the discriminator based on the discriminator loss; Wherein, the acoustic model training operations include: Calculate an information loss based on the initial true acoustic information and the initial predicted acoustic information; Calculate an adversarial loss based on the predicted discrimination result; Calculate a generator loss based on the information loss and the adversarial loss; Optimize the parameters of the acoustic model based on the generator loss.

3. The method according to claim 2, wherein The discriminator includes n sub-discriminators, where n is a positive integer greater than 1. The inputting the initial true acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the true discrimination result and the predicted discrimination result output by the discriminator includes: Perform n - 1 downsampling operations on the initial true acoustic information respectively to obtain n - 1 groups of downsampled true acoustic information, where the downsampling scales of any two of the n - 1 downsampling operations are different; Perform the n - 1 downsampling operations on the initial predicted acoustic information respectively to obtain n - 1 groups of downsampled predicted acoustic information; Input the initial true acoustic information and the n - 1 groups of downsampled true acoustic information into the n sub-discriminators one by one to obtain n sub-true discrimination results output by the n sub-discriminators, where the true discrimination result includes the n sub-true discrimination results; Input the initial predicted acoustic information and the n - 1 groups of downsampled predicted acoustic information into the n sub-discriminators one by one to obtain n sub-predicted discrimination results output by the n sub-discriminators, where the predicted discrimination result includes the n sub-predicted discrimination results.

4. The method according to claim 3, wherein Calculating the real loss based on the real discrimination result includes: Calculating the real loss real-loss through the following formula: real-loss = E s [max(0, 1 - D k (s))], k = 1, 2, 3…n; Calculating the predicted loss based on the predicted discrimination result includes: Calculating the predicted loss fake-loss through the following formula: fake-loss = E x [max(0, 1 + D k (G(x))), k = 1, 2, 3…n; Calculating the adversarial loss based on the predicted discrimination result includes: Calculating the adversarial loss adv-loss through the following formula: adv-loss = E x [-D k (G(x))], k = 1, 2, 3…n; Among them, D k represents the k-th sub-discriminator among the n sub-discriminators, s represents the initial true acoustic information, x represents the text information, G represents the acoustic model, G(x) represents the initial predicted acoustic information, D k (s) represents the sub-true discrimination result corresponding to the k-th sub-discriminator, D k (G(x)) represents the sub-predicted discrimination result corresponding to the k-th sub-discriminator.

5. The method according to claim 3 or 4, wherein The i-th downsampling operation among the n - 1 downsampling operations is used to downsample the corresponding acoustic information by 2^i times, where i = 1, 2, 3... n - 1.

6. The method according to any one of claims 2 to 4, wherein Calculating the discriminator loss based on the real loss and the predicted loss includes: Weighted summing the real loss and the predicted loss to obtain the discriminator loss; and / or Calculating the generator loss based on the information loss and the adversarial loss includes: Weighted summing the information loss and the adversarial loss to obtain the generator loss.

7. The method according to any one of claims 2 to 4, wherein, Calculating the information loss based on the initial real acoustic information and the initial predicted acoustic information includes: Substituting the initial real acoustic information and the initial predicted acoustic information into the mean square error function or the squared absolute error function to calculate the information loss.

8. The method according to claim 1, wherein, The discriminator includes n sub-discriminators, where n is a positive integer greater than 1. Inputting the initial real acoustic information and the initial predicted acoustic information into the discriminator respectively to obtain the real discrimination result and the predicted discrimination result output by the discriminator includes: Performing n - 1 downsampling operations on the initial real acoustic information respectively to obtain n - 1 groups of downsampled real acoustic information, where the downsampling scales of any two of the n - 1 downsampling operations are different; Performing the n - 1 downsampling operations on the initial predicted acoustic information respectively to obtain n - 1 groups of downsampled predicted acoustic information; Inputting the initial real acoustic information and the n - 1 groups of downsampled real acoustic information into the n sub-discriminators one by one to obtain n sub-real discrimination results output by the n sub-discriminators, and the real discrimination result includes the n sub-real discrimination results; Inputting the initial predicted acoustic information and the n - 1 groups of downsampled predicted acoustic information into the n sub-discriminators one by one to obtain n sub-predicted discrimination results output by the n sub-discriminators, and the predicted discrimination result includes the n sub-predicted discrimination results.

9. A speech synthesis method, comprising: Obtain the text to be synthesized; Perform speech synthesis on the text to be synthesized by using the acoustic model trained by the acoustic model training method according to any one of claims 1 to 8 to obtain the target speech.

10. An acoustic model training device, comprising: An acquisition module for acquiring text information and initial real acoustic information, where the text information includes training text or a text feature sequence related to the training text, and the initial real acoustic information includes initial real speech or an initial real acoustic feature sequence related to the initial real speech; A first input module, configured to input the text information into an acoustic model to obtain initial predicted acoustic information output by the acoustic model, wherein the form of the initial predicted acoustic information is the same as that of the initial true acoustic information; A second input module, configured to input the initial true acoustic information and the initial predicted acoustic information into a discriminator respectively to obtain a true discrimination result and a predicted discrimination result output by the discriminator, the true discrimination result corresponding to the initial true acoustic information, and the predicted discrimination result corresponding to the initial predicted acoustic information; and A training module, configured to perform adversarial training on the acoustic model and the discriminator at least based on the true discrimination result and the predicted discrimination result.

11. A speech synthesis device, comprising: An acquisition module, configured to acquire a text to be synthesized; A synthesis module, configured to perform voice synthesis on the text to be synthesized by using an acoustic model obtained by training through the acoustic model training method according to any one of claims 1 to 8 to obtain a target voice.

12. An acoustic model training system, comprising a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the acoustic model training method according to any one of claims 1 to 8.

13. A speech synthesis system, comprising a processor and a memory, wherein, Computer program instructions are stored in the memory, and when the computer program instructions are run by the processor, they are used to execute the voice synthesis method according to claim 9.

14. A storage medium, on which program instructions are stored, and the program instructions are used to execute the acoustic model training method according to any one of claims 1 to 8 when running.

15. A storage medium, on which program instructions are stored, and the program instructions are used to execute the speech synthesis method according to claim 9 when running.

Citation Information

Patent Citations

  • Voice synthesis model training method and system, voice synthesis method and system, equipment and medium

    CN111627418A