Speech synthesis model product
By using the pronunciation posteriori diagram to carry accent information in the speech synthesis model and predicting the pronunciation characteristics and posteriori diagrams from phoneme vectors, the problem of difficulty in synthesizing accents in the prior art is solved, and efficient non-Mandarin speech synthesis is achieved.
Patent Information
- Application Number
- CN202211024404.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-12
AI Technical Summary
Existing speech synthesis models are difficult to synthesize speech with accents, especially when less non-Mandarin audio is used.
Phonetic Posterior Grams (PPGs) are used to carry accent information, and predict speech characteristics and speech posterior graphs from the phoneme vectors of the text to be synthesized through neural network models to generate target speech matching the accent.
It realizes the automatic synthesis of voices with accents when using less non-Mandarin audio, which improves the richness of voice synthesis and solves the problem of insufficient accent information in the prior art.
Smart Images

Figure CN115294963B_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese patent application with an application date of April 12, 2022, an application number of 202210377265.8, and an invention title of "Speech Synthesis Method, Neural Network Model Training Method, and Speech Synthesis Model". Technical Field
[0002] Embodiments of the present application relate to the field of neural network technologies, and in particular, to a speech synthesis method, a neural network model training method, and a speech synthesis model product. Background Art
[0003] Currently, end-to-end models based on neural networks are constantly improving, and the modeling ability of speech synthesis models is continuously enhanced, making the synthesis time of speech shorter, the speed faster, the effect more robust, and the synthesized speech more natural. However, existing speech synthesis models require a large database and a large amount of computing resources. On the other hand, in daily life, due to geographical influence, dialects with strong accents are widely used, but existing speech synthesis models are difficult to synthesize speech audio with accents. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a speech synthesis solution to at least partially solve the above problems.
[0005] According to a first aspect of embodiments of the present application, a speech synthesis method is provided, including: obtaining a phoneme vector of a text to be synthesized; predicting, from the phoneme vector, speech features and a speech posteriorgram corresponding to each phoneme, where the speech posteriorgram carries accent information; generating a speech spectrum according to the speech features and the speech posteriorgram; and outputting, based on the speech spectrum, a target speech corresponding to the text to be synthesized, where the accent of the target speech matches the accent information.
[0006] According to a second aspect of embodiments of the present application, a speech synthesis model product is provided, including an encoder, a decoder, and a vocoder. The encoder is configured to predict speech features and a speech posteriorgram from a phoneme vector of a text to be synthesized, where the speech posteriorgram carries accent information. The decoder is configured to determine a speech spectrum based on the speech features and the speech posteriorgram. The vocoder is configured to generate a target speech corresponding to the text to be synthesized according to the speech spectrum, where the accent of the target speech matches the accent information in the speech posteriorgram.
[0007] According to a third aspect of the embodiments of the present application, a method for training a neural network model is provided. The method is used to train the above speech synthesis model, and the method includes: training the speech synthesis model using audio samples corresponding to a first accent to obtain an initially trained speech synthesis model; training the initially trained speech synthesis model using audio samples corresponding to a second accent to obtain a secondarily trained speech synthesis model, wherein the duration of the audio samples corresponding to the first accent is greater than that of the audio samples corresponding to the second accent.
[0008] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the method described in the first aspect.
[0009] According to a fifth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.
[0010] According to a sixth aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the above method.
[0011] In this way, real target speech with a non-Mandarin accent can be generated, thereby enhancing the richness of the synthesizable speech. In this embodiment, phonetic posterior grams (PPGs) are innovatively applied to the synthesis of speech with a strong accent (i.e., non-Mandarin), thereby realizing the automatic synthesis of accented speech with less non-Mandarin audio, and solving the problem in the prior art that due to the shortage of accented audio, accented speech cannot be synthesized. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0013] Figure 1A Schematic diagram of a speech synthesis model according to Embodiment 1 of the present application;
[0014] Figure 1BSchematic diagram of an encoder and a decoder in a speech synthesis model according to Embodiment 1 of the present application;
[0015] Figure 1C Schematic diagram of the variance adapter of the encoder in the speech synthesis model according to Embodiment 1 of the present application;
[0016] Figure 2 Flowchart of the steps of a speech synthesis method according to Embodiment 1 of the present application;
[0017] Figure 3 Sub - flowchart of the sub - steps of step S204 of a speech synthesis method according to Embodiment 1 of the present application;
[0018] Figure 4 Sub - flowchart of the sub - steps of step S206 of a speech synthesis method according to Embodiment 1 of the present application;
[0019] Figure 5 Flowchart of the steps of a neural network model training method according to Embodiment 2 of the present application;
[0020] Figure 6 Sub - flowchart of the sub - steps of step 502 of a neural network model training method according to Embodiment 2 of the present application;
[0021] Figure 7 Another sub - flowchart of the sub - steps of step 502 of a neural network model training method according to Embodiment 2 of the present application;
[0022] Figure 8 Block diagram of the structure of a speech synthesis device according to Embodiment 3 of the present application;
[0023] Figure 9 Block diagram of the structure of a neural network model training device according to Embodiment 4 of the present application;
[0024] Figure 10 Schematic diagram of the structure of an electronic device according to Embodiment 5 of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application shall fall within the protection scope of the embodiments of the present application.
[0026] The following further illustrates the specific implementation of the embodiments of the present application in conjunction with the accompanying drawings of the embodiments of the present application.
[0027] Example 1
[0028] In this embodiment, a new neural network model (which can also be called a speech synthesis model) capable of synthesizing target speech with an accent (i.e., non-Mandarin) is provided. For ease of understanding, before explaining the implementation process of the speech synthesis method, the speech synthesis model will be described first.
[0029] Referring to Figure 1A , a schematic diagram of a speech synthesis model is shown. The model includes an encoder, a decoder, and a vocoder.
[0030] Among them, the encoder (i.e., Figure 1A the encoder shown in
[0031] ) is used to predict speech features and a phonetic posteriorgram from the phoneme vectors of the text to be synthesized, and the phonetic posteriorgram carries accent information.
[0032] The speech features may include the fundamental frequency (F0) and energy information (energy) of each phoneme in the target speech to be synthesized, but are not limited thereto. The encoder in this embodiment can not only predict the fundamental frequency and energy information, but also predict the phonetic posteriorgram (Phonetic PosteriorGrams, PPGs). The phonetic posteriorgram can extract language-independent phoneme posterior probabilities to form a phonetic posteriorgram. The phonetic posteriorgram can retain sound-related information (such as accent information) while excluding the influence of the speaker, so that the phonetic posteriorgram can serve as a bridge between the speaker and the speech. Through the accent corresponding to each phoneme indicated in the phonetic posteriorgram and the duration of each phoneme, the accent of the subsequent synthesized target speech can be well controlled, solving the problem that it is difficult to synthesize non-Mandarin-accented speech due to different phonemes and rhythms of different-accented speech. Figure 1A The decoder (i.e.,
[0033] the decoder shown in Figure 1A ) is used to determine the speech spectrum based on the speech features and the phonetic posteriorgram. The speech spectrum can be a Mel spectrum.
[0034] The encoder, decoder, and vocoder will be described below by way of example. As Figure 1B shown, the encoder includes a variance adapter and multiple encoding modules. The encoding modules are used to extract context information from the phoneme vectors of the text to be synthesized, and the variance adapter is used to predict the speech features and the phonetic posteriorgram based on the output data of the encoding modules.
[0035] The number of encoding modules can be determined according to requirements, and there is no limit to this. For example, the number of encoding modules can be 6, but this number is only an example and is not limited thereto.
[0036] Among them, the encoding module can include an encoding multi-head self-attention layer (multihead attention), an encoding normalization layer (add&norm), an encoding one-dimensional convolutional layer (conv1D), etc.
[0037] After concatenating the phoneme vector with the position information (position encoding) of each phoneme, it is used as the encoding input data and input into the encoding multi-head self-attention layer, and the first feature information is extracted by the encoding multi-head self-attention layer.
[0038] In the encoding normalization layer connected to the encoding multi-head self-attention layer, the first feature information and the encoding input data are normalized to obtain the first normalization result.
[0039] The first normalization result is input into the encoding one-dimensional convolutional layer to further extract the second feature information. In the encoding normalization layer connected to the encoding one-dimensional convolutional layer, the second feature information and the first normalization result are normalized to obtain the output data. The context information of the text to be synthesized is carried in the output data.
[0040] The output data of multiple encoding modules is input into a variance adaptor, as Figure 1C shown, an example variance adaptor includes a fundamental frequency prediction unit (pitch predictor), an energy prediction unit (energy predictor), and a phonetic posteriorgram prediction unit (PPG predictor). The phonetic posteriorgram prediction unit can use an LSTM network or other appropriate neural networks, and there is no limit to this.
[0041] The output data of the encoding module is input into the variance adaptor. On the one hand, the output data is processed by the fundamental frequency prediction unit to output the fundamental frequency corresponding to each phoneme. In order to further solve the problem that the pronunciation and prosody of phonemes in non-Mandarin Chinese are different from those of phonemes in Mandarin Chinese, the fundamental frequency prediction unit in this embodiment outputs the normalized logarithmic scale fundamental frequency (i.e., Log-F0).
[0042] On the other hand, the output data is processed by the energy prediction unit to output the energy corresponding to each phoneme. On the other hand, the output data is processed by the phonetic posteriorgram unit to output the phonetic posteriorgram, which contains the duration and accent information of each phoneme.
[0043] The variance adapter splices the output fundamental frequency, energy, and speech posterior map into the output data of the encoding module to form the encoded data output by the encoder. Generally speaking, the variance adapter uses the hidden sequence (i.e., the output data of the encoding module) as the input and uses the MSE (mean square error) loss function to predict the fundamental frequency, energy, and speech posterior map of each speech frame corresponding to each phoneme.
[0044] Similarly to the encoder, the decoder includes a plurality of decoding modules, and the decoding modules are used to generate a speech spectrum based on the input speech features, speech posterior map, and a preset speaker vector. For example, the number of decoding modules is the same as the number of encoding modules, which is also 6. Of course, this is just an example and does not limit it to other numbers.
[0045] The decoding module includes a decoding multi-head self-attention layer, a decoding normalization layer, a decoding one-dimensional convolutional layer, etc. The splicing position information of the encoded data output by the variance adapter is used as the decoding input data and input into the decoding module. The decoding multi-head self-attention layer of the decoding module processes the decoding input data to obtain the third feature information.
[0046] The third feature information and the decoding input data are input into the decoding normalization layer connected to the decoding multi-head self-attention layer, and the third normalization result is output. The third normalization result is input into the decoding one-dimensional convolutional layer to obtain the fourth feature information output by it. The fourth feature information and the third normalization result are input into the decoding normalization layer connected to the decoding one-dimensional convolutional layer to obtain the output fourth normalization result.
[0047] In this way, through the processing of multiple decoding modules, the fourth normalization result output by the decoding module is input into the linear layer, and the output speech spectrum is obtained.
[0048] The target speech is generated according to the speech spectrum by a vocoder. In this way, it is possible to synthesize the target speech with an accent other than Mandarin, thereby enriching the function of speech synthesis. The vocoder can be LPCnet, and of course it can also be other neural networks that can convert the speech spectrum into the target speech, and this is not limited.
[0049] The speech synthesis model can be an end-to-end neural network model called PPG_FS. The encoder and decoder of this model are non-autoregressive structures. Through training, the attention alignment mechanism can be extracted from the encoder-decoder based teacher model to improve the accuracy. The LPCNet vocoder is used to convert acoustic features (such as speech spectrum) into speech frames in the target speech to realize the synthesis of the target speech.
[0050] Next, in combination with the speech synthesis model, the speech synthesis method of this embodiment will be described. Of course, it should be noted that this method is not limited to using the speech synthesis model exemplified in this embodiment, but can also be applied to other models.
[0051] As Figure 2 shown, the method includes the following steps:
[0052] Step S202: Obtain the phoneme vectors of the text to be synthesized.
[0053] The text sequence to be synthesized can be a sentence or a paragraph. When converted into speech, it can be segmented into one or more phonemes, and phoneme embeddings are generated based on the segmented phonemes.
[0054] Step S204: Predict the speech features and speech posteriorgrams corresponding to each phoneme from the phoneme vectors.
[0055] In one example, a trained speech synthesis model is used. As described above, the speech synthesis model includes an encoder, a decoder, and a vocoder.
[0056] Among them, as Figure 3 shown, step S204 can be implemented through the following sub-steps:
[0057] Sub-step S2041: Construct input data based on the phoneme vectors.
[0058] For example, the phoneme vectors are concatenated with the position information of the text to be synthesized under each phoneme to form input data.
[0059] Sub-step S2042: Input the input data into the encoder of the trained speech synthesis model, and obtain the fundamental frequency and the energy information corresponding to each phoneme output by the encoder as the speech features.
[0060] The input data is input into the encoder, and the encoder processes the input data to output the fundamental frequency (F0) and energy information (energy) corresponding to each phoneme.
[0061] The fundamental frequency affects the pitch of the synthesized target speech to a certain extent, and thus also has a certain impact on the tone of the target speech.
[0062] The energy information affects the volume of the target speech and can also reflect the logical stress of the target speech, etc.
[0063] Sub-step S2043: Obtain the speech posteriorgram carrying accent information output by the encoder.
[0064] Since the pronunciation corresponding to each phoneme and the duration are indicated in the phonetic posteriorgram, and these two represent the accent information. The phonetic posteriorgram is the one predicted by the model according to the phoneme and is for the required accent.
[0065] Step S206: Generate a voice spectrum according to the voice feature and the phonetic posteriorgram.
[0066] In a feasible manner, as Figure 4 shown, step S206 can be implemented through the following sub-steps:
[0067] Sub-step S2061: Obtain a speaker vector which carries the timbre information of the speaker.
[0068] In order to make the generated target voice more real and make the timbre of the target voice closer to the real human voice, in the decoder, not only the voice feature (including the fundamental frequency and energy corresponding to each phoneme) and the phonetic posteriorgram are input, but also the speaker vector can be input. The speaker information carries the timbre of the speaker. The speaker vector can be obtained during the decoder training phase.
[0069] Sub-step S2062: Input the voice feature, the phonetic posteriorgram and the speaker vector into the decoder of the voice synthesis model, and obtain the Mel spectrum output by the decoder as the voice spectrum.
[0070] The Mel spectrum can accurately indicate the acoustic features of the target voice, thus ensuring the authenticity of the generated target voice.
[0071] Step S208: Output a target voice corresponding to the text to be synthesized based on the voice spectrum, and the accent of the target voice matches the accent information.
[0072] In a feasible manner, step S208 can be implemented as: Input the voice spectrum into a vocoder, and obtain multiple voice frames output by the vocoder as the target voice corresponding to the text to be synthesized.
[0073] In this way, a real target voice with a non-Mandarin accent can be generated, thus enhancing the richness of the synthesizable voice. In this embodiment, the phonetic posteriorgram (PPGs) is innovatively applied to the synthesis of accented (i.e., non-Mandarin) voices, thus realizing the automatic synthesis of accented voices with less non-Mandarin audio, and solving the problem in the prior art that due to the lack of accented audio, accented voices cannot be synthesized.
[0074] Embodiment 2
[0075] Refer toFigure 5 , showing the schematic flow chart of a neural network model training method according to the second embodiment of the present application.
[0076] This method is used to train the aforementioned speech synthesis model, and it includes the following steps:
[0077] Step S502: Train the speech synthesis model using the audio samples corresponding to the first accent to obtain an initially trained speech synthesis model.
[0078] Among them, the first accent can be Mandarin, or other accent types with a relatively large number of samples (i.e., relatively long audio duration).
[0079] In this embodiment, in order to solve the problem that the audio samples of non-Mandarin accents are insufficient and it is difficult to train a usable, accented speech synthesis model using non-Mandarin accent audio samples, the audio samples of Mandarin with a large number of samples are used to train the speech synthesis model to obtain an initially trained speech synthesis model.
[0080] In one example, as Figure 6 shown, step S502 includes the following sub-steps:
[0081] Sub-step S5021: Extract speech features and speech posteriorgrams from the audio samples corresponding to the first accent.
[0082] The speech features include, but are not limited to, the fundamental frequency and energy of each phoneme in the audio samples corresponding to the first accent. The acquisition method can be obtained by any appropriate and known method, and no limitation is imposed on this.
[0083] The speech posteriorgram can be extracted using Pytorch-Kaldi, but it is not limited to this, and it can also be extracted by other methods capable of extracting speech posteriorgrams.
[0084] Sub-step S5022: Obtain the speaker vector.
[0085] During the training process, at the first training, the speaker vector can be randomly initialized. When it is not the first training, the speaker vector can be the adjusted speaker vector.
[0086] Sub-step S5023: Input the speech features, speech posteriorgrams, and the speaker vector into the decoder of the speech synthesis model, and obtain the speech spectrum output by the decoder.
[0087] Sub-step S5024: Use the vocoder of the speech synthesis model to generate the target speech based on the speech spectrum.
[0088] Sub-step S5025: Adjust the speaker vector according to the target speech and the audio sample, and use the adjusted speaker vector as the new speaker vector. Return to continue executing by inputting the speech feature, speech posteriorgram, and the speaker vector into the decoder of the speech synthesis model until the first termination condition is met, so as to obtain the trained decoder and speaker vector.
[0089] During the training process, since there may be differences between the timbre represented by the speaker vector and the real speaker, and the accuracy of the decoder in extracting the features of the audio sample may be insufficient, there may be a deviation between the synthesized target speech and the audio sample. Therefore, the loss value can be calculated based on the target speech and the audio sample, and then the parameters of the speaker vector and the decoder can be adjusted according to the loss value.
[0090] Use the adjusted speaker vector as the new speaker vector, and return to continue executing sub-step S5023 until the first termination condition is met. The termination condition can be reaching the set number of training times, or the decoder meeting the convergence condition, which is not limited. Through the above sub-steps, the trained decoder can be obtained.
[0091] As Figure 7 shown, based on the trained decoder, the encoder can be trained through the following sub-steps.
[0092] Sub-step S5026: Obtain the phoneme vector sample of the text sample corresponding to the audio sample of the first accent.
[0093] Any appropriate and known method can be used to obtain the phoneme vector sample, which is not limited.
[0094] Sub-step S5027: Input the phoneme vector sample into the encoder of the speech synthesis model, and obtain the speech feature and speech posteriorgram output by the encoder.
[0095] As mentioned above, the encoder can predict the fundamental frequency and energy of each phoneme as the speech feature, and predict the speech posteriorgram.
[0096] Sub-step S5028: Input the speech feature, the speech posteriorgram, and the trained speaker vector into the trained decoder to obtain the speech spectrum output by the trained decoder.
[0097] Sub-step S5029: Use the vocoder of the speech synthesis model to generate the target speech based on the speech spectrum.
[0098] Sub-step S50210: Adjust the encoder according to the target speech and the audio sample, and return to continue executing the step of inputting the phoneme vector sample into the encoder of the speech synthesis model until the termination condition is met, so as to obtain a trained encoder.
[0099] Calculate the loss value based on the target speech and the audio sample, adjust the encoder according to the loss value, and then return to sub-step S5027 to continue executing until the second termination condition is met. The second termination condition can be that the number of training times is satisfied, or the encoder converges, and there is no limitation on this.
[0100] Step S504: Use the audio sample corresponding to the second accent to train the initially trained speech synthesis model to obtain a secondarily trained speech synthesis model, where the duration of the audio sample corresponding to the first accent is longer than that of the audio sample corresponding to the second accent.
[0101] In order to enable the speech synthesis model to synthesize a real target speech with a non-Mandarin accent, after training the speech synthesis model with the audio sample of Mandarin, the audio sample of the second accent can also be used to adjust it, so as to obtain a speech synthesis model corresponding to the second accent.
[0102] The second accent can be a non-Mandarin accent, and the duration of the required audio sample can be shorter than that of the audio sample of the first accent, which ensures that a better speech synthesis model can be trained with fewer audio samples. The process of using the audio sample of the second accent to perform secondary training on the initial speech synthesis model is similar to the process of the foregoing sub-steps S5021 to S50210, so it will not be elaborated here.
[0103] In this embodiment, the speech synthesis model can be improved based on the FastSpeech model, so that the improved speech synthesis model can predict the phonetic posterior grams (PPGs) that cross the speaker and language boundaries, and then use the phonetic posterior grams and the normalized logarithmic scale fundamental frequency (Log-F0), energy, etc. to solve the problems of speech and prosody mismatch between Mandarin and other languages, so as to implement an end-to-end speech synthesis model that can synthesize heavily accented speech.
[0104] Due to the introduction of the language posteriorgram, the speech synthesis model well adapts to the problem that there is insufficient audio sample volume of other accents, making it difficult to directly train a speech synthesis model capable of synthesizing accented speech on the FastSpeech model, overcomes the difficulty of sparse accented audio samples (the audio sample duration used for training the FastSpeech model needs nearly 24 hours, while the improved speech synthesis model (which can also be called PPG_FS) only requires two hours of accented audio sample data), and also solves the problems of difficult accented language annotation, the difference in pitch between the pronunciation of accented phonemes and their own pronunciation, and the uncontrollable synthesis effect, and can easily adapt to the training of an accented language with limited data resources and realize speaker voice conversion.
[0105] We have tried on the low-data-resource Hunan and Northeast accented data. Experiments show that this model can synthesize highly intelligible and natural-sounding accented speech.
[0106] Since during the speech synthesis model training stage, Mandarin audio samples (also called corpora) are used, but when fine-tuning with accented data, the trained Mandarin speaker vectors remain unchanged, enabling the speech synthesis model to synthesize the effect of speaking an accented language with the voice of a Mandarin speaker when finally used, further enriching the target speech.
[0107] The speech posteriorgram is incorporated into this model, and the speech posteriorgram can exclude the speaker identity while retaining the voice information. Therefore, it can serve as a bridge between the speaker and language information. Utilizing the language-independent characteristics of the speech posteriorgram, cross-Mandarin and accented language phoneme posterior representations are carried out, fully solving the problem that different models are trained with different audio samples of different accents due to different matching relationships between phonemes and rhythms, thus achieving a better training effect. In addition, the fundamental frequency is used to further compensate for the mismatch between rhythm and phonemes, and a method of training with a large amount of Mandarin corpora and fine-tuning with a small amount of accented corpora is implemented to achieve accented speech synthesis with speaker voice conversion, that is, speaking the target speech with other accents with the timbre of a Mandarin speaker. Using the aforementioned model and method, text can be converted into realistic speech, and then artificial intelligence services can be carried out.
[0108] Embodiment III
[0109] Refer to Figure 8 , which shows the structural block diagram of the speech synthesis device according to Embodiment III of the present application.
[0110] In this embodiment, the device includes:
[0111] An acquisition module 802, configured to acquire the phoneme vector of the text to be synthesized;
[0112] A prediction module 804, configured to predict speech features and a speech posterior map corresponding to each phoneme from the phoneme vectors, where the speech posterior map carries accent information;
[0113] A generation module 806, configured to generate a speech spectrum according to the speech features and the speech posterior map;
[0114] A synthesis module 808, configured to output a target speech corresponding to the text to be synthesized based on the speech spectrum, where the accent of the target speech matches the accent information.
[0115] Optionally, the speech features include a fundamental frequency corresponding to each phoneme and energy information corresponding to each phoneme. The prediction module 804 is configured to construct input data based on the phoneme vectors; input the input data into an encoder of a trained speech synthesis model, and obtain the fundamental frequency corresponding to each phoneme and the energy information output by the encoder as the speech features; and obtain the speech posterior map carrying the accent information output by the encoder.
[0116] Optionally, the generation module 806 is configured to obtain a speaker vector, where the speaker vector carries timbre information of a speaker; input the speech features, the speech posterior map, and the speaker vector into a decoder of the speech synthesis model, and obtain a Mel spectrum output by the decoder as the speech spectrum.
[0117] Optionally, the synthesis module 808 is configured to input the speech spectrum into a vocoder, and obtain a plurality of speech frames output by the vocoder as the target speech corresponding to the text to be synthesized.
[0118] The device in this embodiment is configured to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein. In addition, the function implementation of each module in the device in this embodiment may refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated herein either.
[0119] Embodiment 4
[0120] Refer to Figure 9 , which shows a structural block diagram of a neural network model training device according to Embodiment 4 of the present application.
[0121] The device includes:
[0122] A first training module 902 is configured to train the speech synthesis model using audio samples corresponding to a first accent to obtain an initially trained speech synthesis model;
[0123] The second training module 904 is used to train the initially trained speech synthesis model using the audio samples corresponding to the second accent to obtain a secondarily trained speech synthesis model, and the duration of the audio samples corresponding to the first accent is greater than that of the audio samples corresponding to the second accent.
[0124] Optionally, the first training module 902 is used to extract speech features and speech posteriorgrams from the audio samples corresponding to the first accent; obtain a speaker vector; input the speech features, speech posteriorgrams, and the speaker vector into the decoder of the speech synthesis model, and obtain the speech spectrum output by the decoder; use the vocoder of the speech synthesis model to generate a target speech based on the speech spectrum; adjust the speaker vector according to the target speech and the audio samples, and use the adjusted speaker vector as the new speaker vector, and return to input the speech features, speech posteriorgrams, and the speaker vector into the decoder of the speech synthesis model to continue execution until the first termination condition is met to obtain a trained decoder and speaker vector.
[0125] Optionally, the first training module 902 is further used to obtain a phoneme vector sample of the text sample corresponding to the audio sample of the first accent; input the phoneme vector sample into the encoder of the speech synthesis model, and obtain the speech features and speech posteriorgrams output by the encoder; input the speech features, the speech posteriorgrams, and the trained speaker vector into the trained decoder to obtain the speech spectrum output by the trained decoder; use the vocoder of the speech synthesis model to generate a target speech based on the speech spectrum; adjust the encoder according to the target speech and the audio samples, and return to the step of inputting the phoneme vector sample into the encoder of the speech synthesis model to continue execution until the second termination condition is met to obtain a trained encoder.
[0126] The device of this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here. In addition, the function implementation of each module in the device of this embodiment can refer to the description of the corresponding part in the foregoing method embodiments, which will not be elaborated here either.
[0127] Embodiment Five
[0128] Referring to Figure 10 , a schematic structural diagram of an electronic device according to Embodiment Five of the present application is shown, and the specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0129] As Figure 10As shown, the electronic device may include: a processor 1002, a communications interface 1004, a memory 1006, and a communication bus 1008.
[0130] Wherein:
[0131] The processor 1002, the communications interface 1004, and the memory 1006 communicate with each other via the communication bus 1008.
[0132] The communications interface 1004 is used to communicate with other electronic devices or servers.
[0133] The processor 1002 is used to execute the program 1010, and specifically may execute the relevant steps in the foregoing method embodiments.
[0134] Specifically, the program 1010 may include program code, and the program code includes computer operation instructions.
[0135] The processor 1002 may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0136] The memory 1006 is used to store the program 1010. The memory 1006 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0137] The program 1010 is specifically used to cause the processor 1002 to perform the corresponding operations of the foregoing method.
[0138] For the specific implementation of each step in the program 1010, reference may be made to the corresponding steps and descriptions in the corresponding units in the foregoing method embodiments, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein.
[0139] The embodiments of the present application also provide a computer program product, including computer instructions, and the computer instructions instruct a computing device to perform the corresponding operations of any one of the foregoing method embodiments.
[0140] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0141] The methods according to the embodiments of the present application described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the methods described herein can be stored on such software processes on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0142] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0143] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application, and the patent protection scope of the embodiments of the present application shall be defined by the claims.
Claims
1. A speech synthesis model product, comprising an encoder, a decoder, and a vocoder; The encoder is an encoder trained based on the trained decoder. The encoder is used to predict speech features and a speech posterior map from the phoneme vectors of the text to be synthesized, and the speech posterior map carries accent information; The decoder is a trained decoder. The decoder is used to determine a speech spectrum based on the speech features and the speech posterior map. The speech posterior map is used to indicate the accent corresponding to each phoneme and the duration of each phoneme. The training of the decoder precedes the training of the encoder; The vocoder is a trained vocoder. The vocoder is used to generate the target speech corresponding to the text to be synthesized according to the speech spectrum, and the accent of the target speech matches the accent information in the speech posterior map.
2. The speech synthesis model product according to claim 1, wherein, The decoder is first trained in the following manner: extracting speech features and a speech posterior map from an audio sample corresponding to a first accent; obtaining a speaker vector; inputting the speech features, the speech posterior map, and the speaker vector into the decoder to obtain the speech spectrum output by the decoder; using the vocoder to generate a target speech based on the speech spectrum; adjusting the speaker vector according to the target speech and the audio sample, and using the adjusted speaker vector as a new speaker vector, and returning to input the speech features, the speech posterior map, and the speaker vector into the decoder to continue execution until a first termination condition is satisfied to obtain a trained decoder and a speaker vector; The encoder is trained based on the trained decoder in the following manner: obtaining a phoneme vector sample of a text sample corresponding to the audio sample of the first accent; inputting the phoneme vector sample into the encoder and obtaining the speech features and the speech posterior map output by the encoder; inputting the speech features, the speech posterior map, and the trained speaker vector into the trained decoder to obtain the speech spectrum output by the trained decoder; using the vocoder to generate a target speech based on the speech spectrum; adjusting the encoder according to the target speech and the audio sample, and returning to the step of inputting the phoneme vector sample into the encoder to continue execution until a second termination condition is satisfied to obtain a trained encoder.
3. The speech synthesis model product according to claim 1, wherein, The encoder includes a plurality of encoding modules and a variance adapter. The encoding modules are used to extract context information from the phoneme vectors of the text to be synthesized, and the variance adapter is used to predict the speech features and the speech posterior map based on the output data of the encoding modules.
4. The speech synthesis model product according to claim 3, wherein, The variance adapter includes: A fundamental frequency prediction unit, configured to output the fundamental frequency corresponding to each phoneme based on the output data of the encoding module; An energy prediction unit, configured to output the energy corresponding to each phoneme based on the output data of the encoding module; A voice posteriorgram prediction unit, configured to output a voice posteriorgram based on the output data of the encoding module.
5. The voice synthesis model product according to claim 4, wherein, the fundamental frequency is a normalized logarithmic scale fundamental frequency.
6. The voice synthesis model product according to claim 1, wherein, the decoder includes a plurality of decoding modules, and the decoding modules are configured to generate a voice spectrum based on input voice features, a voice posteriorgram, and a preset speaker vector.
7. The voice synthesis model product according to claim 6, wherein, the decoding module includes: a multi-head self-attention layer, configured to process the decoding input data by using the encoding data splicing position information output by the encoder as the decoding input data to obtain third feature information; a decoding normalization layer, configured to output a third normalization result according to the third feature information and the decoding input data; a decoding one-dimensional convolutional layer, configured to obtain fourth feature information according to the third normalization result; another decoding normalization layer, configured to obtain a fourth normalization result according to the fourth feature information and the third normalization result; a linear layer, configured to output a voice spectrum according to the fourth normalization result.
8. The voice synthesis model product according to any one of claims 1-7, wherein, both the encoder and the decoder are non-autoregressive structures.
Citation Information
Patent Citations
Voice generation method and device, electronic equipment and readable storage medium
CN113628608A
Reference-fee foreign accent conversion system and method
WO2022046781A1