Voice synthesis model training method, voice synthesis method, and related device
Through adversarial generative training, the speech synthesis model is trained using an adversarial discriminative model, which solves the problem of insufficient correlation between the time domain and frequency domain when generating Mel spectrum in the speech synthesis model, and achieves the effects of clear articulation, natural pronunciation, and strong sense of rhythm.
Patent Information
- Application Number
- CN202211128018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2042-09-16
AI Technical Summary
Existing speech synthesis models have poor correlation between the time domain and frequency domain when generating Mel spectrograms, resulting in the generated speech being unclear and overly smooth.
Adversarial generative training is adopted to train the speech synthesis model using N discriminators in the adversarial discriminative model. By dividing the Mel spectrum into multiple spectrum segments, the speech synthesis model is subjected to adversarial generative training to improve the time domain and frequency domain correlation of the Mel spectrum.
The clarity of speech generated by the speech synthesis model has been improved, with more natural pronunciation, better rhythm and rhyme, and closer to real-life voices.
Smart Images

Figure CN116129853B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech processing, and in particular to a speech synthesis model training method, a speech synthesis method and related equipment. BACKGROUND
[0002] As an important part of artificial intelligence technology, intelligent speech technology has been applied in many fields, such as e-book reading, digital artificial customer service, etc. Intelligent speech synthesis is used in these fields.
[0003] A TTS (Text-to-Speech) model is the core of intelligent speech synthesis technology, which can be divided into two categories: autoregressive TTS models and non-autoregressive TTS models. The non-autoregressive TTS model generates complete mel-spectrogram at one time, while the autoregressive TTS model generates each spectrogram image frame and then combines them into a complete mel-spectrogram. Compared with the two models, the processing method of the non-autoregressive TTS model ignores the correlation between the time domain and the frequency domain, which causes the generated speech to have unclear pronunciation and be too smooth. The autoregressive TTS model generates frame by frame, and the generation of the next frame depends on the generation of the previous frame. The correlation between the time domain and the frequency domain is better than that of the non-autoregressive TTS model. However, the quality of the synthesized speech is not very ideal no matter which method is used.
[0004] That is, the existing speech synthesis model has poor correlation between the time domain and the frequency domain when generating mel-spectrogram, which causes the generated speech to have unclear pronunciation and be too smooth. SUMMARY
[0005] The present application provides a speech synthesis model training method, a speech synthesis method and related equipment to solve the technical problem that the existing speech synthesis model has poor correlation between the time domain and the frequency domain when generating mel-spectrogram, which causes the generated speech to have unclear pronunciation and be too smooth.
[0006] In a first aspect, the present application provides a speech synthesis model training method, comprising:
[0007] Obtaining training data, the training data comprising a target speech and a phoneme sequence corresponding to the target speech;
[0008] Preprocessing the target speech to determine a target mel-spectrogram; and inputting the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel-spectrogram;
[0009] The target mel-frequency spectrum and the predicted mel-frequency spectrum are respectively segmented and paired according to the sound rules of the target voice, to obtain N pairs of spectrum segments, each pair of spectrum segments including a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment being obtained by segmenting the predicted mel-frequency spectrum, and the second spectrum segment being obtained by segmenting the target mel-frequency spectrum; N is an integer greater than 1.
[0010] The N discriminators in the adversarial discriminant model are used to respectively perform adversarial generation training on the speech synthesis model based on the N pairs of spectrum segments, and the trained speech synthesis model is used to synthesize the to-be-synthesized text into synthesized speech.
[0011] In a second aspect, the present application provides a speech synthesis method, comprising:
[0012] Obtaining a phoneme sequence corresponding to the to-be-synthesized text;
[0013] Performing speech synthesis processing on the phoneme sequence by a speech synthesis model to obtain synthesized speech corresponding to the to-be-synthesized text; the speech synthesis model is trained by using any one of the possible speech synthesis model training methods provided in the first aspect
[0014] In a third aspect, the present application provides a speech synthesis model training device, comprising:
[0015] An obtaining module is configured to obtain training data, wherein the training data comprises a target voice and a phoneme sequence corresponding to the target voice;
[0016] A processing module is configured to:
[0017] Preprocess the target voice to determine a target mel-frequency spectrum, and input the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel-frequency spectrum;
[0018] Segment and pair the target mel-frequency spectrum and the predicted mel-frequency spectrum according to the sound rules of the target voice to obtain N pairs of spectrum segments, each pair of spectrum segments including a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment being obtained by segmenting the predicted mel-frequency spectrum, and the second spectrum segment being obtained by segmenting the target mel-frequency spectrum; N is an integer greater than 1.
[0019] The N discriminators in the adversarial discriminant model are used to respectively perform adversarial generation training on the speech synthesis model based on the N pairs of spectrum segments, and the trained speech synthesis model is used to synthesize the to-be-synthesized text into synthesized speech.
[0020] In a fourth aspect, the present application provides a speech synthesis device, comprising:
[0021] an acquisition module configured to acquire a phoneme sequence corresponding to the text to be synthesized;
[0022] a synthesis module configured to perform speech synthesis processing on the phoneme sequence by using a speech synthesis model to obtain synthesized speech corresponding to the text to be synthesized, wherein the speech synthesis model is trained by using any one of the training methods for a speech synthesis model provided in the first aspect.
[0023] In a fifth aspect, the present application provides an electronic device, comprising:
[0024] a memory configured to store program instructions;
[0025] a processor configured to invoke and execute the program instructions in the memory to perform the method provided in the first aspect or the second aspect.
[0026] In a sixth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is used to perform the method provided in the first aspect or the second aspect.
[0027] In a seventh aspect, the present application further provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the method provided in the first aspect or the second aspect.
[0028] In the present application, when training a speech synthesis model, first, training data is acquired, the training data comprising target speech and a phoneme sequence corresponding to the target speech; the target speech is preprocessed to determine a target mel-frequency spectrum; and the phoneme sequence is input into the speech synthesis model for synthesis processing to obtain a predicted mel-frequency spectrum; the target mel-frequency spectrum and the predicted mel-frequency spectrum are respectively split and paired according to the sound rules of the target speech to obtain N pairs of spectrum segments; and the N discriminators in the adversarial discriminant model are used to respectively perform adversarial generation training on the speech synthesis model based on the N pairs of spectrum segments, and the trained speech synthesis model is used to synthesize text to be synthesized into synthesized speech. As can be seen, the speech synthesis model is trained by using the adversarial generation training method, so that the time domain and frequency domain correlation of the mel-frequency spectrum generated by the speech synthesis model is strengthened, and by dividing the mel-frequency spectrum into multiple spectrum segments, the expression of harmonic energy in the mel-frequency spectrum is clearer, the outline of the high-frequency spectrum graph is clearer, and the corresponding energy points are clearer, so that the pronunciation of the synthesized speech generated by the speech synthesis model is more accurate, and the technical problems of unclear pronunciation and excessive smoothness are solved. The technical effects of clear pronunciation, more natural pronunciation, better rhythm and prosody, and closer to real human voice are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0029] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, further serve to explain the principles of the application.
[0030] Figure 1 A flowchart of a training method of a speech synthesis model provided for an embodiment of the application;
[0031] Figure 2 A training scenario diagram of a speech synthesis model provided for an embodiment of the application;
[0032] Figure 3 A flowchart of a speech synthesis method provided for an embodiment of the application;
[0033] Figure 4 A flowchart of another training method of a speech synthesis model provided for an embodiment of the application;
[0034] Figure 5 A training scenario diagram of another speech synthesis model provided for an embodiment of the application;
[0035] Figure 6 A structural diagram of a training device of a speech synthesis model provided for an embodiment of the application;
[0036] Figure 7 A structural diagram of a speech synthesis device provided for an embodiment of the application;
[0037] Figure 8 A structural diagram of an electronic device provided for the application.
[0038] Through the above-mentioned drawings, the specific embodiments of the application have been shown, and will be described in more detail hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the application by any means, but to illustrate the concept of the application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0039] In order to make the objects, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without any creative work, including but not limited to combinations of multiple embodiments, all belong to the scope of protection of the application.
[0040] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, if any, are used for distinguishing between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the use of these terms herein is merely for distinguishing between the similar objects and the same can be referenced to other terms. A single component can be designated by different terminologies and the same terminology can represent different components. Terminologies or words used in the description and the claims of the present application are not intended to limit the scope of the application, and the meaning of a term can be "converted" to the meaning of the other term when it is clear that the "converted" meaning of the term is intended in the relevant place. For example, when a certain component is referred to as "comprising" or "including" a list of functions or units, it means that the component can further include the other functions or units not listed or stated.
[0041] The professional terms related to the present application are explained as follows:
[0042] embedding: In the field of deep learning, it refers to the features extracted from the original data, that is, the low-dimensional vector after the original data is mapped through the neural network.
[0043] FastSpeech2: A TTS (Text-to-Speech) model jointly proposed by Microsoft Asia Research Institute and Zhejiang University. On the basis of the FastSpeech1 model, the Teacher-Student knowledge distillation framework is abandoned, the training complexity is reduced, and the real speech data is directly used as the training target, thereby avoiding information loss, and more accurate duration information and other variable information in speech, such as pitch and energy, are introduced to improve the quality of synthesized speech.
[0044] Mel spectrogram: Mel spectrogram, which is obtained by applying Mel filter bank to the multi-frame spectrogram composed of power spectrum.
[0045] MFCC (Mel-frequency cepstral coefficients): Mel-frequency cepstral coefficients. A feature widely used in speaker segmentation, voiceprint recognition, speech recognition, and speech synthesis. Mel frequency is based on the characteristics of human ear hearing, which has a nonlinear correspondence relationship with Hertz Hz frequency. Mel-frequency cepstral coefficients are calculated using this relationship between them. Mainly used for speech data feature extraction.
[0046] TTS (Text-to-Speech) model: converting text into corresponding speech.
[0047] over-smoothed: during the training process, as the number of network layers increases and the number of iterations increases, the hidden layer representation of each node tends to converge to the same value (i.e., the same position in space).
[0048] GAN (Generative Adversarial Nets): In the GAN model, there are a generation model (Genertive Model or Genertor) and a discriminative model (Discriminative Model or Dicriminator). It is generally applied in the field of image generation.
[0049] Non-autoregressive TTS models have attracted increasing attention from both industry and academia, but non-autoregressive TTS models generate mel-spectrograms all at once, ignoring the correlation between time and frequency domains, resulting in unclear pronunciation and over-smoothed speech generated by non-autoregressive TTS models. Autoregressive TTS models, such as tacotron1 / 2 models, generate mel-spectrograms one frame at a time, with the next frame depending on the previous frame, so the correlation between time and frequency domains is stronger, and the naturalness of the generated speech is better than that of non-autoregressive TTS models, but it is not ideal enough.
[0050] The reason is that the inventors of the present application found that in the related art, the TTS model generates mel-spectrograms, and the correlation between the time domain and the frequency domain is not strong enough, resulting in a blurred contour of the generated mel-spectrogram, and the harmonic energy clarity is not enough, and the characteristics of frequency fluctuations at different times in long-range long waves cannot be well handled. Thus, the synthesized speech by the TTS model is unclear and over-smoothed, lacking the naturalness of a real person.
[0051] In summary, the existing speech synthesis model has the technical problem of poor correlation between the time domain and the frequency domain when generating mel-spectrograms, resulting in unclear pronunciation and over-smoothed speech.
[0052] To solve the above problems, the inventive concept of the present application is:
[0053] This application discovered that the Mel spectrum can be regarded as a two-dimensional image. Therefore, deep learning models generally used to process images, such as the GAN model, can be used to optimize the Mel spectrum generated by the speech synthesis model. The generative model in the GAN model is replaced by the TTS speech synthesis model, and the discriminative model in the GAN model is retained. This structure is used to train the TTS speech synthesis model. By improving the clarity of the Mel spectrum corresponding to the synthesized speech instead of directly changing the synthesized speech, this method can improve the clarity of the synthesized speech and make the rhythm of the speech more natural, thereby solving the problems of unclear and overly smooth pronunciation.
[0054] It should be noted that the application scenarios of the training method of the speech synthesis model provided in this application include: e-book reading, digital human customer service, map navigation, speech synthesis model training system, server with speech synthesis function and equipment with TTS speech conversion, etc.
[0055] The following describes in detail the training method of the state speech synthesis model provided by this application.
[0056] Figure 1 A flowchart of a method for training a speech synthesis model provided in an embodiment of the present application. Figure 1 As shown, the specific steps of the training method of the speech synthesis model include:
[0057] S101. Obtain training data.
[0058] In this step, the training data includes: the target speech and the phoneme sequence corresponding to the target speech.
[0059] The target speech includes the preset content recorded by one or more speakers. The phoneme sequence consists of multiple phonemes. A phoneme is the smallest phonetic unit that can distinguish a character or word, also known as a phoneme. In the field of Chinese speech synthesis, a phoneme is composed of pinyin plus prosody (i.e., the length of pauses between words).
[0060] For example, the phoneme sequence is: han2 guo2 7 zui4 da4 de5 7 dao6 yu6 7 ji3zhou1 dao3, which corresponds to Jeju Island, the largest island in South Korea. It should be noted that the 2 after the pinyin "han2" represents the tone: 1 represents the first tone, 2 represents the second tone, 3 represents the third tone, 4 represents the fourth tone, 5 represents the light tone, 6 represents the altered tone, and 7, 8, and 9 represent different rhythmic pause lengths: 7 represents a short pause, 9 represents a long pause, and 8 represents a center pause.
[0061] The manner of obtaining the phoneme sequence corresponding to the target voice includes: converting the target voice into a sequence of pinyin through a pinyin or pronunciation manner corresponding to the target voice and a preset rhythm rule, and then adding a rhythm to each pinyin in the sequence of pinyin according to the preset rhythm rule, so as to obtain the phoneme sequence.
[0062] Specifically, the target voice and the phoneme corresponding thereto can be pre-stored in a training database. Obtaining the training data can mean obtaining the target voice and the phoneme sequence corresponding thereto from the training database.
[0063] S102, pre-processing the target voice to determine a target mel-frequency spectrum; and inputting the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel-frequency spectrum.
[0064] In this step, the mel-frequency spectrum is a spectrogram obtained by filtering a spectrogram composed of a plurality of power spectrums (also referred to as a plurality of power frames) through a mel filter bank. Therefore, pre-processing the target voice to determine the target mel-frequency spectrum can mean: converting the target voice into a plurality of power spectrums by using a mel-frequency spectrum extraction module, then inputting the power spectrums into a mel filter in the mel-frequency spectrum extraction module for filtering, and then combining the filtering results into a complete spectrogram to obtain the target mel-frequency spectrum.
[0065] The speech synthesis model includes a Fastspeech2 model, and the function of the speech synthesis model is to convert text into corresponding speech. The speech synthesis model can include an encoder encoder, a variance adjuster Variance Adoptor and a mel-frequency spectrum demodulator Mel-spectrogram Decoder connected in sequence. Inputting the phoneme sequence into the speech synthesis model for synthesis processing to obtain the predicted mel-frequency spectrum means: inputting the phoneme sequence into the Fastspeech2 model, then the Fastspeech2 model performs embedding processing on the phoneme sequence to generate a vector corresponding to each phoneme, then inputting the vectors into the encoder encoder for encoding processing to obtain an encoded vector, then inputting the encoded vector into the variance adjuster Variance Adoptor for processing, and inputting the processing result into the mel-frequency spectrum demodulator Mel-spectrogram Decoder to extract the predicted mel-frequency spectrum.
[0066] S103, according to the sound rule of the target voice, respectively cutting and grouping the target mel-frequency spectrum and the predicted mel-frequency spectrum to obtain N spectrum segment pairs.
[0067] In this step, each spectrum segment pair includes a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment is obtained by dividing the predicted mel spectrum, and the second spectrum segment is obtained by dividing the target mel spectrum; N is an integer greater than 1.
[0068] Specifically, the sound rule of the target voice indicates that different rhythms or prosodies in the target voice are reflected by the frequency of the target voice in the frequency domain. Therefore, the target mel spectrum and the predicted mel spectrum can be divided into multiple spectrum segment pairs according to different frequency ranges.
[0069] The segmentation processing according to the sound rule of the target voice ensures that the voice spectrum at each rhythm or prosody pause is not ignored, more spectral contour details of the target voice are obtained, the problem of over-smoothness and unnaturalness of the synthesized voice caused by ignoring the above details is avoided, and the module can learn the voice features under different rhythms and prosodies, so that the finally synthesized voice has more rhythm and is more in line with the prosody.
[0070] S104, using N discriminators in the adversarial discriminative model, respectively based on the N spectrum segment pairs, performing adversarial generation training on the speech synthesis model.
[0071] In this step, the adversarial discriminative model includes N discriminators, and each discriminator includes a discriminative model (Discriminator) in a GAN model.
[0072] Specifically, the N discriminators are used to determine whether each spectrum segment pair meets the preset training requirement; wherein, any spectrum segment pair meeting the preset training requirement means that the difference between the first spectrum segment and the second spectrum segment in any spectrum segment pair meets the preset training requirement; if the number of spectrum segment pairs meeting the preset training requirement in the N spectrum segment pairs is less than a threshold, the speech synthesis model is subjected to back propagation training.
[0073] In an embodiment, each discriminator includes a feature extraction network; the plurality of spectrum segment pairs include a target spectrum segment pair, the target spectrum segment pair corresponds to a target discriminator in the N discriminators; the N discriminators are used to determine whether each spectrum segment pair meets the preset training requirement, including:
[0074] The feature extraction network in the target discriminator is used to extract features from the first spectrum segment and the second spectrum segment in the target spectrum segment pair, to determine the predicted voice feature and the target voice feature.
[0075] Using a preset loss function, the similarity between the predicted speech features and the target speech features is calculated; a determination is made as to whether the similarity is greater than or equal to a similarity threshold; if the similarity is greater than or equal to the similarity threshold, the target spectrum segment pair is determined to meet preset training requirements. The preset loss function includes a cosine loss function.
[0076] The embodiment of the present application trains the speech synthesis model by adopting adversarial generative training, so that the time domain and frequency domain correlation of the mel spectrum generated by the speech synthesis model is strengthened, and by dividing the mel spectrum into multiple spectrum segments, the expression of the harmonic energy in the mel spectrum is clearer, the outline of the high-frequency spectrum graph is clearer, and its corresponding energy points are also clearer, thereby making the speech pronunciation generated by the speech synthesis model more accurate, solving the technical problems of fuzzy pronunciation and overly smooth pronunciation. The technical effect of achieving clear pronunciation, more natural pronunciation, better rhythm and rhythm, and closer to real human voice is achieved.
[0077] Based on the training method of the speech synthesis model, the present embodiment provides a training scenario for the speech synthesis model. Figure 2 , is a schematic diagram of a training scenario of a speech synthesis model provided in an embodiment of the present application. Figure 2 As shown in , two types of training data are needed for training. One type is phoneme 101, which is used for speech synthesis. The other type is target speech 102, which is the speech directly recorded by the target speaker. The corresponding factor sequence is the pinyin plus rhythm corresponding to the target speech. Figure 2 As shown, a target speech 102 and multiple phoneme phonemes 101 in a phoneme sequence are first loaded from a training database 100. A mel-spectrogram extraction module 104 then preprocesses the target speech 102 to extract a target mel-spectrogram 103 corresponding to the target speech 102. Simultaneously, the training phoneme phonemes 101 are input into a speech synthesis model 200 for TTS speech synthesis processing, converting them into predicted mel-spectrograms 201. Then, the predicted Mel spectrum 201 and the target Mel spectrum 103 are segmented according to the sound rules of the target speech 102, the predicted Mel spectrum 201 is segmented into n first spectrum segments 2011, and the target Mel spectrum 103 is segmented into n second spectrum segments 1031, and one first spectrum segment 2011 corresponds to one second spectrum segment 1031 one to one, and are combined into spectrum segment pairs 401. Then, each spectrum segment pair 401 is input into the corresponding discriminator 301 in the adversarial discriminant model 300, and each discriminator 301 determines whether the difference between the first spectrum segment 2011 and the second spectrum segment 1031 meets the preset training requirements; if not, the speech synthesis model 200 is back-propagation trained.
[0078] The above steps are repeated for adversarial training and model iteration until the preset number of discriminators 301 in the adversarial discriminant model 300 cannot recognize the difference between the corresponding first frequency spectrum segment 2011 and the second frequency spectrum segment 1031, or the difference between the two is small enough, or the difference is less than the preset difference threshold, which proves that the speech synthesis model 200 has been trained.
[0079] It should be noted that the present application introduces the principle of adversarial training of GAN model in the field of image synthesis into the field of speech synthesis, breaks through the calculation barrier between the two fields, and compensates for the lack of speaker features extracted by the speech synthesis model through adversarial training. The way of thinking or technical inertia in the related art of obtaining more speaker feature vectors, or increasing the dimension number of feature vectors, or increasing the number of neural network layers or complexity to change the feature extraction method. Simplify the way to add speaker features for synthesized speech, and improve the training efficiency of the model.
[0080] After the training of the above embodiments, a trained speech synthesis model is obtained. The following describes a method of using the trained speech synthesis model for speech synthesis.
[0081] Figure 3 A flowchart of a speech synthesis method provided by an embodiment of the present application is shown in FIG. 8. As shown in FIG. 8, the speech synthesis method uses the trained speech synthesis model in the embodiment shown in FIG. 7, and the specific steps include: Figure 3 Figure 1
[0082] S301, obtaining a phoneme sequence corresponding to the text to be synthesized.
[0083] In this step, the text to be synthesized is the text content input by the user through the input interface on the terminal or selected from the selection box of the input interface. After the terminal obtains the text to be synthesized, it finds the phoneme sequence corresponding to the text to be synthesized from the background database. Alternatively, after the terminal obtains the text to be synthesized, the same method as in step S201 is used to obtain the phoneme sequence corresponding to the text to be synthesized.
[0084] S302, performing speech synthesis processing on the phoneme sequence by the speech synthesis model to obtain a synthesized speech corresponding to the text to be synthesized.
[0085] In this step, the speech synthesis model is trained by the method shown in FIG. 7. Figure 1 The training method of the speech synthesis model of the illustrated embodiment is trained. The speech synthesis model comprises a FastSpeech2 model trained by the above training method. The FastSpeech2 model vectorizes the phonemes in the phoneme sequence, then encodes the phoneme vector through an encoder, and adds feature vectors such as the speech, rhythm, intonation, prosody, and timbre of the target speaker through an adjustor, to combine to obtain a speech vector corresponding to the synthesized speech, and then decode through a decoder to output the synthesized speech.
[0086] The embodiment of the present application combines the phonemes in the phoneme sequence through the FastSpeech2 model, and adds the rhythm and intonation of the speaker, so that the resulting synthesized speech sounds more natural, avoids the problem of overly smooth and mechanical synthesized speech, improves the user's experience, and makes the synthesized speech more intelligent.
[0087] Figure 4 Another flowchart of the training method of the speech synthesis model provided by the embodiment of the present application is shown in Figure 4 The specific steps of the training method of the speech synthesis model include:
[0088] S401, obtaining training data.
[0089] In this step, the training data includes a target speech and a phoneme sequence corresponding to the target speech.
[0090] S402, pre-processing the target speech to determine a target mel-frequency spectrum, and inputting the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel-frequency spectrum.
[0091] For the nomenclature and implementation principles of S401 and S402, reference can be made to S101-S102, which will not be repeated here.
[0092] S403, respectively cutting and grouping the target mel-frequency spectrum and the predicted mel-frequency spectrum according to the sound rules of the target speech to obtain N spectrum segment pairs.
[0093] In this step, the sound rules of the target speech reflect different rhythms or prosodies in the target speech through the high and low frequencies in the frequency domain of the target speech. A spectrum segment pair includes a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment is obtained by cutting the predicted mel-frequency spectrum, and the second spectrum segment is obtained by cutting the target mel-frequency spectrum; N is an integer greater than 1.
[0094] When N is 3, the sound rules include a first frequency threshold and a second frequency threshold, and the first frequency threshold is less than the second frequency threshold.
[0095] In the embodiment, the steps of specific cutting and grouping of the processing include:
[0096] The first low-frequency spectrum segment, the first medium-frequency spectrum segment, and the first high-frequency spectrum segment of the predicted mel spectrum are respectively taken as the first spectrum segment.
[0097] The second low-frequency spectrum segment, the second medium-frequency spectrum segment, and the second high-frequency spectrum segment of the target mel spectrum are respectively taken as the second spectrum segment.
[0098] The first low-frequency spectrum segment and the second low-frequency spectrum segment form a spectrum segment pair, the first medium-frequency spectrum segment and the second medium-frequency spectrum segment form a spectrum segment pair, and the first high-frequency spectrum segment and the second high-frequency spectrum segment form a spectrum segment pair.
[0099] It is worth noting that the lengths of the spectrum segments obtained by cutting the predicted mel spectrum or the target mel spectrum are not necessarily equal, which can increase the diversity of the spectrum segments. In addition, the first frequency threshold and the second frequency threshold can also be set to randomly change over time, which can further increase the diversity of the spectrum segments, better handle the influence of fluctuations of the mel spectrum at different times, better extract the details in the spectrum, and avoid the problem of ignoring details to make the synthesized speech too smooth and unnatural.
[0100] Next, it is necessary to use N discriminators to determine whether each spectrum segment pair meets the preset training requirements. If the number of spectrum segment pairs that meet the preset training requirements in the N spectrum segment pairs is less than the number threshold, the speech synthesis model is trained by back propagation.
[0101] It is worth noting that each discriminator includes a feature extraction network; the multiple spectrum segment pairs include a target spectrum segment pair, and the target spectrum segment pair corresponds to a target discriminator in the N discriminators.
[0102] S404, the feature extraction network in the target discriminator is used to extract features from the first spectrum segment and the second spectrum segment in the target spectrum segment pair, to determine the predicted speech features and the target speech features.
[0103] In the embodiment, the feature extraction network of the target discriminator includes multiple feature extraction nodes, and each feature extraction node includes a two-dimensional feature extractor. This step specifically includes:
[0104] S4041, input the first spectrum segment and the second spectrum segment in the target spectrum segment pair into the two-dimensional feature extractor of each feature extraction node of the target discriminator respectively for feature extraction, to obtain a predicted feature vector corresponding to each feature extraction node and a target feature vector corresponding to each feature extraction node.
[0105] Specifically, the two-dimensional feature extractor includes a Conv2D. The Conv2D is a two-dimensional feature extraction function in a Convolution Layers convolution layer. Since the mel spectrum can be understood as a two-dimensional picture containing multiple frames, each spectrum segment can be understood as one or more two-dimensional pictures, and the Conv2D can be used to extract a feature vector in the spectrum segment.
[0106] S4042, determine a predicted speech feature according to the predicted feature vector corresponding to each feature extraction node, and determine a target speech feature according to the target feature vector corresponding to each feature extraction node.
[0107] In a possible design, each feature extraction node further includes a fitting processing layer and a normalization processor. The fitting processing layer includes a Dropout. The Dropout can significantly reduce the overfitting phenomenon by ignoring half of the feature detectors (letting half of the hidden layer nodes have a value of 0) in each training batch. This way can reduce the interaction between the feature detectors (hidden layer nodes), and the detector interaction refers to that some detectors depend on other detectors to function. That is, during the forward propagation, the activation value of a certain neuron stops working with a certain probability p, so that the model has stronger generalization ability because it does not rely too much on some local features.
[0108] The normalization processor includes a BatchNorm, which is often used in a deep neural network, is an algorithm for accelerating neural network training, accelerating convergence speed and stability, and can be said to be an indispensable part of the current deep neural network. The essence of the neural network learning process is to learn the data distribution. If we do not do normalization processing, the distribution of each batch of training data is different. From a large perspective, the neural network needs to find a balance point in the multiple distributions. From a small perspective, since the input data distribution of each network layer is constantly changing, it will also cause each network layer to find a balance point. Obviously, the neural network is difficult to converge. Of course, if we only normalize the input data (such as dividing the input image by 255 to make it between 0 and 1), we can only ensure that the input layer data distribution is the same, and we cannot ensure that the input data distribution of each network layer is the same, so we need to add normalization processing to the middle layer of the neural network. The essence of the neural network learning process is to learn the data distribution. If the distribution of the training data and the test data is different, the generalization ability of the network will be severely reduced. Suppose the input image contains four dimensions: N, C, H, and W. The calculation of BatchNorm is to separately normalize the N, H, and W dimensions of each channel.
[0109] At this time, the determination of the predicted speech feature according to the predicted feature vector corresponding to each feature extraction node and the determination of the target speech feature according to the target feature vector corresponding to each feature extraction node in S4042 include:
[0110] The predicted feature vector corresponding to each feature extraction node and the target feature vector corresponding to each feature extraction node are activated by using a preset activation function (such as LeakyReLU) to determine a plurality of predicted activation vectors and a plurality of target activation vectors; a predicted activation vector is obtained by activating the predicted feature vector corresponding to a feature extraction node, and a target activation vector is obtained by activating the target feature vector corresponding to a feature extraction node; the plurality of predicted activation vectors and the plurality of target activation vectors are input into the fitting processing layer (such as Dropout) of the corresponding feature extraction node for fitting processing, and the number of neurons in the fitting processing layer is reduced by using the allocation probability function to determine a plurality of predicted fitting vectors and a plurality of target fitting vectors; the plurality of predicted fitting vectors and the plurality of target fitting vectors are normalized by using a normalization processor (such as BatchNorm) to obtain the predicted speech feature and the target speech feature.
[0111] S405, using a preset loss function, calculating the similarity of the predicted speech feature and the target speech feature, and determining whether the similarity is greater than or equal to a similarity threshold.
[0112] In this step, if the similarity is greater than or equal to the similarity threshold, the target spectral segment pair is determined to meet the preset training requirements. It is worth noting that this step requires evaluation of all target spectral segments. If the number of spectral segment pairs meeting the preset training requirements is less than the preset threshold, S406 is executed. Otherwise, the speech synthesis model has been trained and training ends.
[0113] S406: Perform back-propagation training and model iteration on the speech synthesis model.
[0114] This application trains a speech synthesis model by adopting adversarial generative training, so that the time domain and frequency domain correlation of the mel spectrum generated by the speech synthesis model is strengthened, and by dividing the mel spectrum into multiple spectrum segments, the expression of the harmonic energy in the mel spectrum is clearer, the outline of the high-frequency spectrum graph is clearer, and its corresponding energy points are also clearer, so that the speech pronunciation generated by the speech synthesis model is more accurate, and the technical problems of fuzzy pronunciation and over-smoothing are solved. The technical effect of achieving clear pronunciation, more natural pronunciation, better rhythm and rhythm, and closer to real human voice is achieved.
[0115] To facilitate understanding, the following examples are given to illustrate each of the above steps. Figure 5 Schematic diagram of another training scenario of a speech synthesis model provided in an embodiment of the present application. Figure 5 As shown, the adversarial discriminator model 330 includes: multiple discriminators 331, each of which includes: a two-dimensional feature extractor 3311, an activation function 3312, a fitting processing layer 3313, a normalization processor 3314, and a preset loss function 3315. The structure composed of the two-dimensional feature extractor 3311, the activation function 3312, the fitting processing layer 3313, and the normalization processor 3314 is called a feature extraction node.
[0116] like Figure 5 As shown, the target speech 312 is preprocessed by using the Mel spectrum extraction module 314 to extract the target Mel spectrum corresponding to the target speech 312, namely the target Mel spectrum 313. In this embodiment, the speech synthesis model includes: FastSpeech2 model. Figure 5 As shown, the training phoneme 311 is input into the FastSpeech2 model 320 for TTS speech synthesis processing, and the speech synthesized by the FastSpeech2 model 320 is converted into a predicted mel spectrum 321 by the mel spectrum demodulator 322.
[0117] Then, the adversarial discriminator model 330 will randomly divide the predicted mel spectrogram 321 and the target mel spectrogram 313 into three different length segments, respectively. In this way, the diversity of the samples is increased. The data in the same batch size, that is, the same size of the sample group, is randomly divided into mel spectrogram segments at different times. This indirectly increases the diversity of the speaker features of the mel spectrogram speech synthesis model during training, and can better handle the fluctuations in tone, intonation, rhythm, volume, etc. of the speaker features at different times.
[0118] The adversarial discriminator model 330 includes three discriminators 331, each of which corresponds to processing a first frequency spectrum segment 3211 and a second frequency spectrum segment 3131 corresponding to the first frequency spectrum segment 3211. Specifically, the first frequency spectrum segment 3211 is input into the two-dimensional feature extractor 3311 to obtain a predicted segment feature vector. Similarly, the second frequency spectrum segment 3131 is input into the two-dimensional feature extractor 3311 to obtain a target segment feature vector. The two-dimensional feature extractor 3311 includes a Conv2D function in the GAN model.
[0119] Then, the predicted segment feature vector and the target segment feature vector are respectively activated by using a preset activation function 3312 to obtain a predicted segment activation vector and a target segment activation vector. In this embodiment, the activation function 3312 includes LeakyReLU.
[0120] Next, the predicted segment activation vector and the target segment activation vector are input into the fitting processing layer 3313 for fitting processing to determine a predicted segment fitting vector and a target segment fitting vector. In this embodiment, the fitting processing layer 3313 includes Dropout. The implementation principle of Dropout is to reduce the number of neurons by using an assignment probability function to prevent overfitting problems caused by a small number of samples.
[0121] Then, the predicted segment fitting vector and the target segment fitting vector are input into the normalization processor 3314 for normalization processing to determine a predicted segment speech feature and a target segment speech feature. In this embodiment, the normalization processor 3314 includes BatchNorm.
[0122] Then, a preset loss function 3315 is used to calculate the similarity between the predicted segment voice feature and the target segment voice feature. If the similarity is greater than or equal to a preset similarity threshold, it is determined that the preset training requirement is met, otherwise, the model parameters in the FastSpeech2 model 320 are adjusted according to the difference between the predicted segment voice feature and the target segment voice feature, that is, back propagation and model iteration are performed, and then the above process is repeated until the similarity is greater than or equal to the preset similarity threshold, and the training of the FastSpeech2 model 320 is completed. In this embodiment, the preset loss function includes: Least SquaresGAN loss.
[0123] The embodiment of the application provides a speech synthesis model training method, which splits the entire mel spectrum into multiple mel spectrum segments, increases the diversity of the mel spectrum, and thus makes the extracted speaker features more accurate and rich. The technical problem of low quality of speech synthesis caused by insufficient speaker feature information in the training of the existing speech synthesis model is solved. The technical effect of adding the speaker timbre feature to the training of the language synthesis model and improving the quality of the synthesized speech output by the speech synthesis model is achieved.
[0124] Figure 6 A structural schematic diagram of a speech synthesis model training device provided by the embodiment of the application is provided. The speech synthesis model training device 600 can be realized by software, hardware or a combination of both.
[0125] As shown in Figure 6 The speech synthesis model training device 600 includes:
[0126] The acquisition module 601 is configured to acquire training data, and the training data includes target speech and a phoneme sequence corresponding to the target speech.
[0127] The processing module 602 is configured to:
[0128] preprocess the target speech to determine a target mel spectrum, and input the phoneme sequence into a speech synthesis model for synthesis processing to obtain a predicted mel spectrum;
[0129] The target mel spectrum and the predicted mel spectrum are respectively cut and paired according to the sound rules of the target speech to obtain N spectrum segment pairs, each spectrum segment pair including a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment being obtained by cutting the predicted mel spectrum, and the second spectrum segment being obtained by cutting the target mel spectrum; N is an integer greater than 1.
[0130] The N discriminators in the adversarial discriminant model are used to perform adversarial generation training on the speech synthesis model based on N spectrum segment pairs respectively, and the trained speech synthesis model is used to synthesize the to-be-synthesized text into synthesized speech.
[0131] In a possible design, the processing module 602 is configured to:
[0132] Each of the N discriminators is used to determine whether each spectrum segment pair meets the preset training requirement; wherein, the first spectrum segment and the second spectrum segment in any spectrum segment pair meet the preset training requirement means that the difference between the first spectrum segment and the second spectrum segment in any spectrum segment pair meets the preset training requirement;
[0133] If the number of spectrum segment pairs meeting the preset training requirement in the N spectrum segment pairs is less than the number threshold, the speech synthesis model is subjected to back propagation training.
[0134] In a possible design, each discriminator includes a feature extraction network; the plurality of spectrum segment pairs include a target spectrum segment pair, and the target spectrum segment pair corresponds to a target discriminator in the N discriminators; and the processing module 602 is configured to:
[0135] The feature extraction network in the target discriminator is used to perform feature extraction on the first spectrum segment and the second spectrum segment in the target spectrum segment pair respectively, to determine the predicted speech feature and the target speech feature;
[0136] A preset loss function is used to calculate the similarity between the predicted speech feature and the target speech feature;
[0137] It is determined whether the similarity is greater than or equal to a similarity threshold;
[0138] If the similarity is greater than or equal to the similarity threshold, it is determined that the target spectrum segment pair meets the preset training requirement.
[0139] In a possible design, the feature extraction network of the target discriminator includes a plurality of feature extraction nodes, and each feature extraction node includes at least one two-dimensional feature extractor; correspondingly, the processing module 602 is configured to:
[0140] The first spectrum segment and the second spectrum segment in the target spectrum segment pair are input into the two-dimensional feature extractor of each feature extraction node of the target discriminator for feature extraction, to obtain a predicted feature vector corresponding to each feature extraction node and a target feature vector corresponding to each feature extraction node;
[0141] The predicted speech feature is determined according to the predicted feature vector corresponding to each feature extraction node, and the target speech feature is determined according to the target feature vector corresponding to each feature extraction node.
[0142] In a possible design, each feature extraction node further includes at least one fitting processing layer and at least one normalization processor, the feature vector includes a predicted feature vector and a target feature vector, and the processing module 602 is configured to:
[0143] perform activation processing on the predicted feature vector corresponding to each feature extraction node and the target feature vector corresponding to each feature extraction node by using a preset activation function, to determine a plurality of predicted activation vectors and a plurality of target activation vectors; one predicted activation vector is obtained by performing activation processing on the predicted feature vector corresponding to one feature extraction node, and one target activation vector is obtained by performing activation processing on the target feature vector corresponding to one feature extraction node;
[0144] perform fitting processing on the plurality of predicted activation vectors and the plurality of target activation vectors by using the fitting processing layer of the corresponding feature extraction node, and reduce the number of neurons in the fitting processing layer by using an allocation probability function, to determine a plurality of predicted fitting vectors and a plurality of target fitting vectors;
[0145] perform normalization processing on the plurality of predicted fitting vectors and the plurality of target fitting vectors by using the normalization processor, to obtain predicted speech features and target speech features.
[0146] In a possible design, the processing module 602 is configured to:
[0147] The sound rule of the target speech reflects different rhythms or cadences in the target speech through the frequency of the target speech in the frequency domain, and the sound rule includes a first frequency threshold and a second frequency threshold, the first frequency threshold is smaller than the second frequency threshold; when the value of N is 3, the processing module 602 is configured to:
[0148] take a first low-frequency spectral segment smaller than the first frequency threshold, a first medium-frequency spectral segment greater than the first frequency threshold and smaller than the second frequency threshold, and a first high-frequency spectral segment greater than the second frequency threshold in the predicted mel spectrum as a first spectral segment, respectively;
[0149] take a second low-frequency spectral segment smaller than the first frequency threshold, a second medium-frequency spectral segment greater than the first frequency threshold and smaller than the second frequency threshold, and a second high-frequency spectral segment greater than the second frequency threshold in the target mel spectrum as a second spectral segment, respectively;
[0150] take the first low-frequency spectral segment and the second low-frequency spectral segment as a spectral segment pair, take the first medium-frequency spectral segment and the second medium-frequency spectral segment as a spectral segment pair, and take the first high-frequency spectral segment and the second high-frequency spectral segment as a spectral segment pair.
[0151] It is worth noting that, Figure 6The apparatus provided by the embodiments shown can perform the training method provided in the embodiments of any of the voice synthesis model training methods described above, and the specific implementation principles, technical features, professional term explanations, and technical effects are similar, and will not be repeated here.
[0152] Figure 7 A structural schematic diagram of a voice synthesis device provided by the embodiments of the present application is shown. The voice synthesis model training device 700 can be implemented by software, hardware, or a combination of both.
[0153] As shown in the figure, the voice synthesis model training device 700 includes: Figure 7
[0154] The acquisition module 701 is configured to acquire a phoneme sequence corresponding to the text to be synthesized.
[0155] The synthesis module 702 is configured to perform voice synthesis processing on the phoneme sequence by using a voice synthesis model to obtain a synthesized voice corresponding to the text to be synthesized. The voice synthesis model is trained by using the voice synthesis model training method described above.
[0156] It is worth noting that, Figure 7 The apparatus provided by the embodiments shown can perform the voice synthesis method provided in any of the method embodiments described above, and the specific implementation principles, technical features, professional term explanations, and technical effects are similar, and will not be repeated here.
[0157] Figure 8 A structural schematic diagram of an electronic device provided by the embodiments of the present application is shown. As shown in the figure, Figure 8 The electronic device 800 can include at least one processor 801 and a memory 802. Figure 8 It is shown that the electronic device takes one processor as an example.
[0158] The memory 802 is configured to store a program. Specifically, the program can include program code, and the program code includes computer operation instructions for implementing the voice synthesis model training method or the voice synthesis method provided by the method embodiments described above.
[0159] The memory 802 can include a high-speed RAM memory, and can also include a non-volatile memory such as at least one disk memory.
[0160] The processor 801 is configured to execute the computer execution instructions stored in the memory 802 to implement the voice synthesis model training method or the voice synthesis method described in the above method embodiments.
[0161] The processor 801 can be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to perform the operations of the embodiments of the present application.
[0162] The memory 802 can be independent or integrated with the processor 801. When the memory 802 is independent of the processor 801, the electronic device 800 can further include:
[0163] The bus 803 is used to connect the processor 801 and the memory 802. The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like, but does not mean that there is only one bus or one type of bus.
[0164] Optionally, in a specific implementation, if the memory 802 and the processor 801 are integrated on a chip, the memory 802 and the processor 801 can communicate through an internal interface.
[0165] The embodiments of the present application also provide a computer readable storage medium, which can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes, and specifically, the computer readable storage medium stores program instructions. The program instructions are used for the training method of the speech synthesis model or the speech synthesis method in the above-mentioned method embodiments.
[0166] The embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the training method of the speech synthesis model or the speech synthesis method in the above-mentioned method embodiments is implemented.
[0167] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0168] It is to be understood that the application is not limited to the precise construction herein disclosed and shown in the drawings, and that various changes in shape, size and arrangements of parts can be made without departing from the scope of the application. The scope of the application is limited only by the claims that follow.
Claims
1. A method of training a speech synthesis model, the method comprising: The method comprises: obtaining training data, wherein the training data comprises target speech and a phoneme sequence corresponding to the target speech; preprocessing the target speech to obtain a target mel spectrum, and inputting the phoneme sequence into a speech synthesis model to obtain a predicted mel spectrum through synthesis processing; performing segmentation and group pairing processing on the target mel spectrum and the predicted mel spectrum according to a sound rule of the target speech to obtain N spectrum segment pairs, wherein each spectrum segment pair comprises a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment is obtained by performing segmentation processing on the predicted mel spectrum, and the second spectrum segment is obtained by performing segmentation processing on the target mel spectrum; N is an integer greater than 1; performing adversarial generation training on the speech synthesis model based on the N spectrum segment pairs by using N discriminators in an adversarial discriminative model, and the trained speech synthesis model is used to synthesize text to be synthesized into synthesized speech.
2. The training method of claim 1, wherein, The method of performing adversarial generation training on the speech synthesis model based on the N spectrum segment pairs by using N discriminators in an adversarial discriminative model comprises: using the N discriminators to determine whether each spectrum segment pair meets a preset training requirement; wherein the first spectrum segment and the second spectrum segment in any spectrum segment pair meet the preset training requirement, which means that a difference between the first spectrum segment and the second spectrum segment in the any spectrum segment pair meets the preset training requirement; if a number of spectrum segment pairs meeting the preset training requirement in the N spectrum segment pairs is less than a number threshold, performing back propagation training on the speech synthesis model.
3. The training method of claim 2, wherein, Each discriminator comprises a feature extraction network; the N spectrum segment pairs comprise a target spectrum segment pair, and the target spectrum segment pair corresponds to a target discriminator in the N discriminators; the method of using the N discriminators to determine whether each spectrum segment pair meets the preset training requirement comprises: using the feature extraction network in the target discriminator to perform feature extraction on the first spectrum segment and the second spectrum segment in the target spectrum segment pair to determine predicted speech features and target speech features; using a preset loss function to calculate a similarity between the predicted speech features and the target speech features; determining whether the similarity is greater than or equal to a similarity threshold; if the similarity is greater than or equal to the similarity threshold, determining that the target spectrum segment pair meets the preset training requirement.
4. The training method of claim 3, wherein, The feature extraction network in the target discriminator comprises a plurality of feature extraction nodes, and each feature extraction node comprises a two-dimensional feature extractor. The method of using the feature extraction network in the target discriminator to perform feature extraction on the first spectrum segment and the second spectrum segment in the target spectrum segment pair to determine predicted speech features and target speech features comprises: The first spectrum segment and the second spectrum segment in the target spectrum segment pair are respectively input into a two-dimensional feature extractor of each feature extraction node of the target discriminator for feature extraction, to obtain a predicted feature vector corresponding to each feature extraction node and a target feature vector corresponding to each feature extraction node; The predicted voice feature is determined according to the predicted feature vector corresponding to each feature extraction node, and the target voice feature is determined according to the target feature vector corresponding to each feature extraction node.
5. The training method of claim 4, wherein, The feature extraction node further comprises a fitting processing layer and a normalization processor, and the predicted voice feature is determined according to the predicted feature vector corresponding to each feature extraction node, and the target voice feature is determined according to the target feature vector corresponding to each feature extraction node. The predicted feature vector corresponding to each feature extraction node and the target feature vector corresponding to each feature extraction node are activated by using a preset activation function, to determine a plurality of predicted activation vectors and a plurality of target activation vectors; one predicted activation vector is obtained by activating the predicted feature vector corresponding to one feature extraction node, and one target activation vector is obtained by activating the target feature vector corresponding to one feature extraction node; The plurality of predicted activation vectors and the plurality of target activation vectors are input into the fitting processing layer of the corresponding feature extraction node for fitting processing, and the number of neurons in the fitting processing layer is reduced by using an allocation probability function, to determine a plurality of predicted fitting vectors and a plurality of target fitting vectors; The plurality of predicted fitting vectors and the plurality of target fitting vectors are normalized by using the normalization processor, to obtain the predicted voice feature and the target voice feature.
6. The training method of claim 1, wherein, The sound rule of the target voice indicates that the frequency high and low in the frequency domain of the target voice reflects different rhythms or rhythms in the target voice, and the sound rule comprises a first frequency threshold and a second frequency threshold, and the first frequency threshold is less than the second frequency threshold; The value of N is 3, the target mel spectrum and the predicted mel spectrum are respectively segmented and paired according to the sound features of the target voice, to obtain N spectrum segment pairs, which comprise: The first low-frequency spectrum segment less than the first frequency threshold, the first medium-frequency spectrum segment greater than the first frequency threshold and less than the second frequency threshold, and the first high-frequency spectrum segment greater than the second frequency threshold in the predicted mel spectrum are respectively taken as the first spectrum segment; And the second low-frequency spectrum segment less than the first frequency threshold, the second medium-frequency spectrum segment greater than the first frequency threshold and less than the second frequency threshold, and the second high-frequency spectrum segment greater than the second frequency threshold in the target mel spectrum are respectively taken as the second spectrum segment; The first low-frequency spectrum segment and the second low-frequency spectrum segment form a spectrum segment pair, the first medium-frequency spectrum segment and the second medium-frequency spectrum segment form a spectrum segment pair, and the first high-frequency spectrum segment and the second high-frequency spectrum segment form a spectrum segment pair.
7. A speech synthesis method characterized by, Comprise: obtain a phoneme sequence corresponding to the text to be synthesized; perform speech synthesis processing on the phoneme sequence by using a speech synthesis model to obtain synthesized speech corresponding to the text to be synthesized; the speech synthesis model is obtained by using the training method of the speech synthesis model according to any one of claims 1-6.
8. A voice synthesis model training apparatus characterized by comprising: comprising: an obtaining module, configured to obtain training data, the training data comprising: target speech and a phoneme sequence corresponding to the target speech; a processing module, configured to: perform preprocessing on the target speech to obtain a target mel spectrum; and input the phoneme sequence into a speech synthesis model to obtain a predicted mel spectrum; perform segmentation and pair processing on the target mel spectrum and the predicted mel spectrum according to the sound rules of the target speech to obtain N pairs of spectrum segments, each pair of spectrum segments comprising a first spectrum segment and a second spectrum segment corresponding to each other, the first spectrum segment being obtained by segmenting the predicted mel spectrum, and the second spectrum segment being obtained by segmenting the target mel spectrum; N is an integer greater than 1; perform adversarial generation training on the speech synthesis model based on the N pairs of spectrum segments by using N discriminators in an adversarial discriminant model, and the trained speech synthesis model is used to synthesize text to be synthesized into synthesized speech.
9. A speech synthesis apparatus characterized by comprising: comprising: an obtaining module, configured to obtain a phoneme sequence corresponding to the text to be synthesized; a synthesizing module, configured to perform speech synthesis processing on the phoneme sequence by using a speech synthesis model to obtain synthesized speech corresponding to the text to be synthesized; the speech synthesis model is obtained by using the training method of the speech synthesis model according to any one of claims 1-6.
10. An electronic device, comprising: comprising: a processor; and a memory, configured to store a computer program of the processor; wherein the processor is configured to execute the training method of the speech synthesis model according to any one of claims 1-6 or the speech synthesis method according to claim 7 by executing the computer program.
Citation Information
Patent Citations
Training method and device of vocoder, method for synthesizing audio signal and vocoder
CN113436603A
Speech synthesis model training method, speech synthesis method, speech synthesis device and medium
CN114038447A
Speech enhancement method of deep generative adversarial network
CN114446314A