Methods, Terminals, and Storage Media for Training a Spectral Synthesis Model and Synthesizing Audio
By training the spectrum synthesis model, using text and speech feature information and intention vectors to generate synthetic audio with higher naturalness, solving the problem of rigid pronunciation of synthetic audio in the prior art.
Patent Information
- Application Number
- CN202111093218.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-09-17
AI Technical Summary
In the prior art, the synthetic audio generated based on text feature information is more rigid and mechanical and lacks naturalness.
By entering the training samples (including text samples, speech samples, and standard intent vectors) into the initial spectrum synthesis model, text feature information and speech spectrum data are extracted, and model parameters are adjusted based on the predicted intent vector and standard intent vectors to generate more natural synthetic audio.
It improves the naturalness and quality of synthetic audio, making pronunciation more natural and speech quality higher.
Smart Images

Figure CN113920982B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and particularly to a method, a terminal, and a storage medium for training a spectral synthesis model and synthesizing audio. Background Art
[0002] With the development of science and technology, audiobooks and audio information have become increasingly common, which makes the demand for automatically synthesizing audio based on text more and more urgent.
[0003] In related technologies, the scheme for generating synthesized audio according to text is as follows: extracting feature information of the target text to obtain the target text feature information corresponding to the target text. Inputting the target text feature information into a pre-trained spectral synthesis model to obtain the target spectral data corresponding to the target text. Inputting the target spectral data corresponding to the target text into a vocoder to obtain the synthesized audio corresponding to the target text. Among them, the text feature information includes phoneme feature information, word segmentation feature information, and prosody feature information.
[0004] Since the above synthesized audio is only generated based on the target text feature information, the pronunciation is relatively rigid and mechanical. Summary of the Invention
[0005] Embodiments of this application provide a method, a terminal, and a storage medium for training a spectral synthesis model and synthesizing audio. Since this application fully considers the influence of the speaker's speaking intention on pronunciation, the pronunciation of the synthesized audio is more natural, improving the quality of the synthesized audio. The technical solution is as follows:
[0006] In a first aspect, embodiments of this application provide a method for training a spectral synthesis model, and the method includes:
[0007] Inputting a training sample into an initial spectral synthesis model, where the training sample includes a text sample, a corresponding speech sample, and a standard intention vector;
[0008] Extracting the sample text feature information corresponding to the text sample, the standard spectral data corresponding to the speech sample, and the predicted intention vector corresponding to the speech sample;
[0009] Determining the predicted spectral data corresponding to the text sample according to the sample text feature information and the predicted intention vector;
[0010] Determining a first loss value according to the predicted spectral data and the standard spectral data;
[0011] Determining a second loss value according to the predicted intention vector and the standard intention vector;
[0012] Adjusting the parameters of the initial spectral synthesis model according to the first loss value and the second loss value;
[0013] If the preset training end condition is satisfied, the initial spectrum synthesis model after parameter adjustment is determined as the trained spectrum synthesis model.
[0014] If the preset training end condition is not satisfied, continue to adjust the parameters of the initial spectrum synthesis model after parameter adjustment according to other training samples.
[0015] Optionally, the spectrum synthesis model includes a text encoder, a voice encoder, and a first self-attention learning module.
[0016] The spectrum synthesis model includes a text encoder, a voice encoder, and a first self-attention learning module.
[0017] The extraction of the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample, and the predicted intention vector corresponding to the voice sample includes:
[0018] Input the text sample into the text encoder to obtain the sample text feature information corresponding to the text sample.
[0019] Extract the spectrum data of the voice sample as the standard spectrum data.
[0020] Input the standard spectrum data corresponding to the voice sample into the voice encoder to obtain a voice vector.
[0021] Input the voice vector into the first self-attention learning module to obtain the predicted intention vector of the voice sample, where the predicted intention vector is composed of the possibility values corresponding to each intention classification arranged in a preset order.
[0022] Optionally, the spectrum synthesis model includes the intention feature information corresponding to each intention classification and a spectrum synthesis module.
[0023] The determination of the predicted spectrum data corresponding to the text sample according to the sample text feature information and the predicted intention vector includes:
[0024] Multiply the possibility value of each intention classification in the predicted intention vector by the intention feature information corresponding to each intention classification stored in advance, and combine the multiplication results to obtain the sample intention feature information corresponding to the voice sample.
[0025] Input the sample text feature information and the sample intention feature information into the spectrum synthesis module to obtain the predicted spectrum data corresponding to the text sample.
[0026] Optionally, the adjustment of the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value includes:
[0027] Determine a comprehensive loss value according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value;
[0028] Adjust the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0029] Optionally, the determining the second loss value according to the predicted intent vector and the standard intent vector includes:
[0030] Determine the cross entropy between the predicted intent vector and the standard intent vector, and use the cross entropy as the second loss value.
[0031] In a second aspect, an embodiment of the present application provides a method for synthesizing audio, the method including:
[0032] Input the target text into a trained natural language processing model to obtain a target intent classification;
[0033] Determine a target intent vector corresponding to the target intent classification;
[0034] Input the target text and the target intent vector into a trained spectrum synthesis model, and determine target spectrum data corresponding to the target text according to the target text and the target intent vector;
[0035] Input the target spectrum data corresponding to the target text into a vocoder to obtain target speech corresponding to the target text.
[0036] Optionally, when the spectrum synthesis model includes a text encoder, intent feature information corresponding to each intent classification, and a spectrum synthesis module, the determining the target spectrum data corresponding to the target text according to the target text and the target intent vector includes:
[0037] Input the target text into the text encoder to obtain target text feature information corresponding to the target text;
[0038] Multiply the probability values of each intent classification in the target intent vector by the intent feature information corresponding to each intent classification stored in advance respectively, and combine the multiplication results to obtain target intent feature information corresponding to the target text;
[0039] Input the target text feature information and the target intent feature information into the spectrum synthesis module to obtain target spectrum data corresponding to the target text.
[0040] Optionally, the determining the target intent vector corresponding to the target intent classification includes:
[0041] Set the possibility value corresponding to the target intent classification to 1, set the possibility values corresponding to other intent classifications to 0, and sort the possibility values corresponding to each intent classification in a preset order to obtain the target intent vector corresponding to the target intent classification.
[0042] Thirdly, an embodiment of the present application provides a device for training a spectrum synthesis model, and the device includes:
[0043] An input module, configured to input a training sample into an initial spectrum synthesis model, where the training sample includes a text sample, a corresponding speech sample, and a standard intent vector;
[0044] An extraction module, configured to extract the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the speech sample, and the predicted intent vector corresponding to the speech sample;
[0045] A first determination module, configured to determine the predicted spectrum data corresponding to the text sample according to the sample text feature information and the predicted intent vector;
[0046] A second determination module, configured to determine a first loss value according to the predicted spectrum data and the standard spectrum data;
[0047] A third determination module, configured to determine a second loss value according to the predicted intent vector and the standard intent vector;
[0048] A parameter adjustment module, configured to adjust the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value;
[0049] A first judgment module, configured to determine the initial spectrum synthesis model after parameter adjustment as a trained spectrum synthesis model if a preset training end condition is satisfied;
[0050] A second judgment module, configured to continue to adjust the parameters of the initial spectrum synthesis model after parameter adjustment according to other training samples if the preset training end condition is not satisfied.
[0051] Optionally, the spectrum synthesis model includes a text encoder, a speech encoder, and a first self-attention learning module;
[0052] The extraction module is configured to:
[0053] Input the text sample into the text encoder to obtain the sample text feature information corresponding to the text sample;
[0054] Extract the spectrum data of the speech sample as the standard spectrum data;
[0055] Input the standard spectrum data corresponding to the speech sample into a speech encoder to obtain a speech vector;
[0056] Input the speech vector into a first self-attention learning module to obtain a predicted intention vector of the speech sample, where the predicted intention vector is formed by arranging the likelihood values corresponding to each intention classification in a preset order.
[0057] Optionally, the spectrum synthesis model includes intention feature information corresponding to each intention classification and a spectrum synthesis module;
[0058] The first determination module is configured to:
[0059] Multiply the likelihood value of each intention classification in the predicted intention vector by the intention feature information corresponding to each intention classification stored in advance, and combine the multiplication results to obtain sample intention feature information corresponding to the speech sample;
[0060] Input the sample text feature information and the sample intention feature information into the spectrum synthesis module to obtain predicted spectrum data corresponding to the text sample.
[0061] Optionally, the parameter adjustment module is configured to:
[0062] Determine a comprehensive loss value according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value;
[0063] Adjust the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0064] Optionally, the third determination module is configured to:
[0065] Determine the cross entropy between the predicted intention vector and the standard intention vector, and use the cross entropy as the second loss value.
[0066] In a fourth aspect, an embodiment of the present application provides a device for synthesizing audio, where the device includes:
[0067] A first input module configured to input the target text into a trained natural language processing model to obtain a target intention classification;
[0068] A determination module configured to determine a target intention vector corresponding to the target intention classification;
[0069] A second input module configured to input the target text and the target intention vector into a trained spectrum synthesis model, and determine target spectrum data corresponding to the target text according to the target text and the target intention vector;
[0070] A third input module, configured to input target spectrum data corresponding to the target text into a vocoder to obtain target speech corresponding to the target text.
[0071] Optionally, when the spectrum synthesis model includes a text encoder, intention feature information corresponding to each intention classification, and a spectrum synthesis module, the second input module is configured to:
[0072] Input the target text into the text encoder to obtain target text feature information corresponding to the target text;
[0073] Multiply the probability values of each intention classification in the target intention vector by the intention feature information corresponding to each intention classification stored in advance, and combine the multiplication results to obtain target intention feature information corresponding to the target text;
[0074] Input the target text feature information and the target intention feature information into the spectrum synthesis module to obtain target spectrum data corresponding to the target text.
[0075] Optionally, the determination module is configured to:
[0076] Set the probability value corresponding to the target intention classification to 1, set the probability values corresponding to other intention classifications to 0, and sort the probability values corresponding to each intention classification in a preset order to obtain a target intention vector corresponding to the target intention classification.
[0077] On the one hand, a terminal is provided, which includes a processor and a memory. At least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the above method.
[0078] On the one hand, a computer-readable storage medium is provided. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the above method.
[0079] On the one hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer program code. The computer program code is stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device executes the above method.
[0080] In the embodiments of the present application, a text sample, a corresponding voice sample, and a standard intent vector are input into an initial spectrum synthesis model, so that the initial spectrum synthesis model extracts the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample, and the predicted intent vector corresponding to the voice sample. Then, according to the sample text feature information and the predicted intent vector, the predicted spectrum data corresponding to the text sample is determined. Finally, according to the predicted spectrum data and the standard spectrum data, as well as the predicted intent vector and the standard intent vector, the initial spectrum synthesis model is jointly adjusted. The spectrum synthesis model obtained according to the above training method can generate spectrum data based on intent feature information and text feature information. Furthermore, the voice generated by the vocoder based on this spectrum data is more natural and has higher voice quality compared to the voice generated only based on target text feature information in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0082] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0083] Figure 2 is a flowchart of a method for training a spectrum synthesis model provided by an embodiment of the present application;
[0084] Figure 3 is a schematic diagram of a method for training a spectrum synthesis model provided by an embodiment of the present application;
[0085] Figure 4 is a flowchart of a method for synthesizing audio provided by an embodiment of the present application;
[0086] Figure 5 is a schematic diagram of a method for synthesizing audio provided by an embodiment of the present application;
[0087] Figure 6 is a schematic diagram of the structure of a device for training a spectrum synthesis model provided by an embodiment of the present application;
[0088] Figure 7 is a schematic diagram of the structure of a device for synthesizing audio provided by an embodiment of the present application;
[0089] Figure 8 is a schematic diagram of the structure of a terminal provided by an embodiment of the present application;
[0090] Figure 9 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific embodiments
[0091] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0092] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application. As Figure 1 shown, this method can be implemented by the terminal 101 or the server 102.
[0093] The terminal 101 may include components such as a processor and a memory. The processor, which can be a CPU (Central Processing Unit), etc., can be used to input training samples into the initial spectrum synthesis model, extract the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample, and the predicted intention vector corresponding to the voice sample, determine the predicted spectrum data corresponding to the text sample according to the sample text feature information and the predicted intention vector, determine the first loss value according to the predicted spectrum data and the standard spectrum data, determine the second loss value according to the predicted intention vector and the standard intention vector, adjust the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value, if the preset training end condition is satisfied, determine the adjusted initial spectrum synthesis model as the trained spectrum synthesis model, if the preset training end condition is not satisfied, continue to adjust the parameters of the adjusted initial spectrum synthesis model according to other training samples, and other processing. The memory, which can be a RAM (Random Access Memory), Flash, etc., can be used to store training samples, etc. The terminal 101 may further include a transceiver, an image detection component, a screen, an audio output component, and an audio input component, etc. Among them, the audio output component may be a speaker, a headset, etc. The audio input component may be a microphone, etc.
[0094] Server 102 may include components such as a processor and a memory. The processor, which can be a CPU (Central Processing Unit), etc., can be used to input training samples into the initial spectrum synthesis model, extract sample text feature information corresponding to text samples, standard spectrum data corresponding to speech samples and predicted intent vectors corresponding to speech samples, determine predicted spectrum data corresponding to text samples according to the sample text feature information and the predicted intent vectors, determine a first loss value according to the predicted spectrum data and the standard spectrum data, determine a second loss value according to the predicted intent vectors and the standard intent vectors, adjust the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value, if the preset training end condition is satisfied, determine the initial spectrum synthesis model after parameter adjustment as the trained spectrum synthesis model, if the preset training end condition is not satisfied, continue to adjust the parameters of the initial spectrum synthesis model after parameter adjustment according to other training samples, etc. The memory, which can be a RAM (Random Access Memory), Flash, etc., can be used to store training samples, etc.
[0095] Figure 2 It is a flowchart of a method for training a spectrum synthesis model provided by an embodiment of the present application. The embodiment of the present application is executed by an electronic device, which can be a server or a terminal. Refer to Figure 2 This embodiment includes:
[0096] Step 201: Input training samples into the initial spectrum synthesis model.
[0097] Among them, the training set includes multiple training samples, and each training sample includes a text sample, a corresponding speech sample and a standard intent vector. The standard intent vector is a vector corresponding to the standard intent classification. The standard intent classification is the classification made by technicians based on experience for speech samples. The intent classifications in the embodiment of the present application include: notification intent, question intent, suggestion intent, and intent of acceptance, rejection and command.
[0098] The steps for obtaining the standard intent vector according to the standard intent classification in the embodiment of the present application are: set the possibility value of the standard intent classification to 1, set the possibility values of other intent classifications except the standard intent classification to 0, and then obtain the possibility value of each intent classification. Arrange the possibility values of each intent classification in a preset order to obtain the standard intent vector.
[0099] For example, if the standard intent classification corresponding to a voice sample is a notification intent, the probability value of this notification intent is set to 1, and the probability values of other intent classifications except the notification intent are set to 0. If the preset order is notification intent, question intent, suggestion intent, and intent of acceptance, rejection, and command, after sorting the probability values of each intent classification in the preset order, the obtained standard intent vector is [1, 0, 0, 0].
[0100] Step 202: Extract the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample, and the predicted intent vector corresponding to the voice sample.
[0101] Among them, the sample text feature information is composed of sample phoneme feature information, sample word segmentation feature information, and sample prosody feature information. The sample phoneme feature information, sample word segmentation feature information, sample prosody feature information, and sample audio feature information all exist in the form of vectors or matrices. The spectrum data can be Mel spectrum, and the method for obtaining the Mel spectrum can be the method for obtaining the Mel spectrum in the prior art, which will not be elaborated here.
[0102] Optionally, as Figure 3 shown, the spectrum synthesis model in the embodiment of this application includes a text encoder, a voice encoder, and a first self-attention learning module. The steps of obtaining the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample, and the predicted intent vector corresponding to the voice sample include: inputting the text sample into the text encoder to obtain the sample text feature information corresponding to the text sample. Extracting the spectrum data of the voice sample as the standard spectrum data. Inputting the standard spectrum data corresponding to the voice sample into the voice encoder to obtain a voice vector. Inputting the voice vector into the first self-attention learning module to obtain the predicted intent vector corresponding to the voice sample, where the predicted intent vector is composed of the probability values corresponding to each intent classification arranged in the preset order.
[0103] The text encoder in the embodiments of the present application includes a phoneme conversion module, a word segmentation module, and a sub-encoder. Among them, the sub-encoder is a neural network module, and the phoneme conversion module and the word segmentation module can be either neural network modules or non-neural network modules. Taking the phoneme conversion module and the word segmentation module as non-neural network modules as an example in the embodiments of the present application, the specific steps to obtain the sample text feature information corresponding to the text sample by inputting the text sample into the text encoder are as follows: Input the text sample into the phoneme conversion module to obtain the sample phoneme sequence corresponding to the text sample. Among them, the corresponding relationship between characters and phonemes is stored in the phoneme conversion module. Input the text sample into the word segmentation module, and divide the text sample according to the sample word segmentation results stored in the word segmentation module to obtain at least one sample word segmentation result. Analyze each sample word segmentation result to determine the pause duration of each sample word segmentation result. Sort the pause durations in the order of the sample word segmentation results in the text sample to obtain the sample prosody sequence. After obtaining the sample phoneme sequence, the sample word segmentation results, and the sample prosody sequence, use the sub-encoder to encode the sample phoneme sequence, the sample word segmentation results, and the sample prosody sequence respectively to obtain the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information. Concatenate the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information in a preset order to obtain the sample text feature information.
[0104] If both the phoneme conversion module and the word segmentation module are neural network modules, it is necessary to train the phoneme conversion module and the word segmentation module separately, and apply the trained phoneme conversion module and the trained word segmentation module to the above process. The methods for training the phoneme conversion module and the word segmentation module are as described in the prior art, and will not be elaborated in the present application.
[0105] It should be noted that if the number of bits corresponding to the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information are the same, the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information can also be added bit by bit to obtain the sample text feature information.
[0106] In addition to the above methods for obtaining the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information, the sample phoneme feature information, the sample word segmentation feature information, and the sample prosody feature information can also be obtained according to other methods in the prior art, and the specific steps will not be elaborated.
[0107] In the process of obtaining the predicted intent vector of the voice sample, both the voice encoder and the first self-attention learning module used are neural network models. The function of the voice encoder is to encode the standard spectral data corresponding to the voice sample to obtain a voice vector. The function of the first self-attention learning module is to perform self-attention learning on the voice vector to obtain the possibility value of the voice vector under each intent classification, and then arrange the possibility values corresponding to each intent classification in a preset order to obtain the predicted intent vector.
[0108] It should be noted that the value range of the possibility value of each intent classification in the predicted intent vector is [0, 1]. For example, [0.2, 0.8, 0, 0]. And the possibility value of each intent classification in the standard intent vector is 1 or 0.
[0109] Step 203: Determine the predicted spectral data corresponding to the text sample according to the sample text feature information and the predicted intent vector.
[0110] Optionally, the spectral synthesis model further includes intent feature information corresponding to each intent classification and a spectral synthesis module. The specific steps for determining the predicted spectral data corresponding to the text sample according to the sample text feature information and the predicted intent vector are as follows: Multiply the possibility value of each intent classification in the predicted intent vector by the intent feature information corresponding to each intent classification stored in advance, and merge the multiplication results to obtain the sample intent feature information corresponding to the voice sample. Input the sample text feature information and the sample intent feature information into the spectral synthesis module to obtain the predicted spectral data corresponding to the text sample.
[0111] Among them, the spectral synthesis module is a neural network module, which is used to predict the predicted spectral data corresponding to the text sample according to the sample text feature information and the sample intent feature information.
[0112] In implementation, multiply the possibility value of each intent classification by the intent feature information corresponding to each intent classification stored in advance, and splice the multiplication results to obtain the sample intent feature information corresponding to the voice sample. Or, add the multiplication results bit by bit to obtain the sample intent feature information corresponding to the voice sample. Input the sample text feature information and the sample intent feature information into the spectral synthesis module to obtain the predicted spectral data corresponding to the text sample.
[0113] It should be noted that in the first training process, that is, the intent feature information corresponding to each intent classification stored in the initial spectral synthesis model is preset by technicians, and in the subsequent training process, that is, the intent feature information corresponding to each intent classification stored in the spectral synthesis model is the intent feature information adjusted during the previous training process.
[0114] Further, the spectrum synthesis module includes a merging module, a second self-attention learning module, and a decoder. The merging module is an algorithm module for merging two pieces of feature information into one piece of feature information, that is, merging two vectors into one vector. The second self-attention learning module is a machine learning module for performing self-attention learning on the feature information to obtain the feature information after self-attention learning. The decoder is a machine learning module for decoding the feature information to obtain predicted spectrum data.
[0115] In implementation, the sample text feature information and the sample intent feature information are input into the merging module in the spectrum synthesis module to obtain the merged first sample feature information. The merged first sample feature information is input into the second self-attention learning module to obtain the second sample feature information. The second sample feature information is input into the decoder to obtain the predicted spectrum data corresponding to the sample text.
[0116] Step 204: Determine a first loss value according to the predicted spectrum data and the standard spectrum data.
[0117] In implementation, the predicted spectrum data and the standard spectrum data are input into the first loss function to obtain the first loss value.
[0118] Step 205: Determine a second loss value according to the predicted intent vector and the standard intent vector.
[0119] Optionally, through the second loss function, determine the cross-entropy of the predicted intent vector and the standard intent vector, and use the cross-entropy as the second loss value. Among them, the second loss function is a cross-entropy function.
[0120] Step 206: Adjust the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value.
[0121] Optionally, according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value, determine the comprehensive loss value. Adjust the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0122] Among them, the first weight and the second weight are values preset by technicians.
[0123] In implementation, multiply the first loss value by the first weight to obtain the weighted first loss value. Multiply the second loss value by the second weight to obtain the weighted second loss value. Add the weighted first loss value and the weighted second loss value to obtain the comprehensive loss value. Adjust the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0124] Or, directly add the first loss value and the second loss value to obtain the comprehensive loss value. Adjust the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0125] It should be noted that when adjusting the parameters of the initial spectrum synthesis model, the modules to be adjusted and the parameters to be adjusted in the initial spectrum model are adjusted. Among them, the modules to be adjusted include the sub-encoder in the text encoder, the voice encoder, the first self-attention learning module, and the second self-attention learning module and decoder in the spectrum synthesis module. The parameter to be adjusted is the intention feature information corresponding to each intention classification.
[0126] Step 207: If the preset training end condition is satisfied, the initial spectrum synthesis model after parameter adjustment is determined as the trained spectrum synthesis model.
[0127] Among them, the preset training end condition can be that the comprehensive loss value is less than the preset value, or that the parameters in the initial spectrum synthesis model after parameter adjustment converge.
[0128] Step 208: If the preset training end condition is not satisfied, continue to adjust the parameters of the initial spectrum synthesis model after parameter adjustment according to other training samples.
[0129] In the embodiment of the present application, the text sample, the corresponding voice sample and the standard intention vector are input into the initial spectrum synthesis model, so that the initial spectrum synthesis model extracts the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the voice sample and the predicted intention vector corresponding to the voice sample, and then, according to the sample text feature information and the predicted intention vector, determines the predicted spectrum data corresponding to the text sample. Finally, according to the predicted spectrum data and the standard spectrum data, and the predicted intention vector and the standard intention vector, the initial spectrum synthesis model is jointly adjusted. The spectrum synthesis model obtained according to the above training method can generate spectrum data according to the intention feature information and the text feature information, and then the vocoder generates voice according to the spectrum data. Compared with the voice generated only according to the target text feature information in the prior art, the pronunciation of this voice is more natural and the voice quality is higher.
[0130] Figure 4 It is a flowchart of synthesizing audio provided by the embodiment of the present application. Refer to Figure 4 , this embodiment includes:
[0131] Step 401: Input the target text into the trained natural language processing model to obtain the target intention classification.
[0132] Among them, the target text is the text corresponding to the audio to be synthesized. This text can be a paragraph of text in an audiobook or a paragraph of text in audio information. In the embodiment of the present application, the natural language processing model is a neural network model, which is used to input the target text into the trained natural language processing model to obtain the target intention classification corresponding to the target text.
[0133] The training method of the natural language processing model in this application is independent of the training method in Embodiment 1. That is, the natural language processing model in actual use is trained separately. Among them, the specific process of training the natural language processing model is as follows: Obtain text samples and the corresponding sample intent classifications of the text samples, and use the sample intent classifications as the standard intent classifications. Input the text samples into the natural language processing model to obtain predicted intent classifications. Input the standard intent classifications and the predicted intent classifications into the third loss function to obtain a third loss value. If the third loss value is greater than a preset value, then adjust the parameters of the natural language processing model according to the third loss value to obtain an adjusted natural language processing model. Use other sample texts and the corresponding sample intent classifications to continue adjusting the parameters of the natural language processing model after the previous parameter adjustment until the third loss information is less than the preset value to obtain a trained natural language processing model.
[0134] Of course, the target intent classification corresponding to the target text can also be obtained by the method of obtaining the intent classification of the text in the prior art.
[0135] Step 402: Determine the target intent vector corresponding to the target intent classification.
[0136] Optionally, set the possibility value corresponding to the target intent classification to 1, set the possibility values corresponding to other intent classifications to 0, and sort the possibility values corresponding to each intent classification in a preset order to obtain the target intent vector corresponding to the target intent classification.
[0137] It should be noted that the preset order in the embodiments of this application is the same as the order of the preset set involved in Embodiment 1.
[0138] Step 403: Input the target text and the target intent vector into the trained spectrum synthesis model, and determine the target spectrum data corresponding to the target text according to the target text and the target intent vector.
[0139] In implementation, as Figure 4 shown, input the target text and the target intent vector into the trained spectrum synthesis model, so that the trained spectrum synthesis model can determine the target spectrum data corresponding to the target text according to the target text and the target intent vector.
[0140] Optionally, when the spectrum synthesis model includes a text encoder, intention feature information corresponding to each intention classification, and a spectrum synthesis module, determining the target spectrum data corresponding to the target text according to the target text and the target intention vector includes: inputting the target text into the text encoder to obtain the target text feature information corresponding to the target text. Multiplying the possibility value of each intention classification in the target intention vector by the intention feature information corresponding to each intention classification stored in advance, and combining the multiplication results to obtain the target intention feature information corresponding to the target text. Inputting the target text feature information and the target intention feature information into the spectrum synthesis module to obtain the target spectrum data corresponding to the target text.
[0141] For the method described in Embodiment 1, the text encoder in the embodiment of the present application includes a phoneme conversion module, a word segmentation module, and a sub-encoder. The specific steps of inputting the target text into the text encoder to obtain the target text feature information corresponding to the target text are as follows: inputting the target text into the phoneme conversion module to obtain the target phoneme sequence corresponding to the target text, where the correspondence between characters and phonemes is stored in the phoneme conversion module. Inputting the target text into the word segmentation module, and dividing the target text according to the target word segmentation result stored in the word segmentation module to obtain at least one target word segmentation result. Analyzing each target word segmentation result to determine the pause duration of each target word segmentation result. Sorting the pause durations in the order of the target word segmentation results in the target text to obtain the target prosody sequence. After obtaining the target phoneme sequence, the target word segmentation result, and the target prosody sequence, using the sub-encoder to encode the target phoneme sequence, the target word segmentation result, and the target prosody sequence respectively to obtain the target phoneme feature information, the target word segmentation feature information, and the target prosody feature information. Concatenating the target phoneme feature information, the target word segmentation feature information, and the target prosody feature information in a preset order to obtain the target text feature information.
[0142] The intention feature information corresponding to each intention classification stored in advance can be the intention feature information corresponding to each intention classification obtained after the spectrum synthesis model is trained, or can be to determine the target intention feature information corresponding to the target intention classification according to the trained spectrum synthesis model. The specific steps are as follows: According to the trained spectrum synthesis model, re-determine the first intention feature information corresponding to each training sample in the sample training set. According to the standard intention classification and the first intention feature information corresponding to each training sample, determine the multiple first intention feature information corresponding to each intention classification. For each intention classification, perform an averaging process on the multiple first intention feature information corresponding to the intention classification to obtain the second intention feature information corresponding to the intention classification, and further obtain the intention feature information corresponding to each intention classification stored in advance.
[0143] The spectrum synthesis module described in Embodiment 1, which includes a merging module, a second self-attention learning module, and a decoder. Input the target text feature information and the target intent feature information into the spectrum synthesis module to obtain the target spectrum data corresponding to the target text, including: input the target text feature information and the target intent feature information into the merging module in the spectrum synthesis module to obtain the merged first target feature information. Input the merged first target feature information into the second self-attention learning module to obtain the second target feature information. Input the second target feature information into the decoder to obtain the predicted spectrum data corresponding to the target text.
[0144] Step 404: Input the target spectrum data corresponding to the target text into the vocoder to obtain the target voice corresponding to the target text.
[0145] Among them, the vocoder is used to generate synthetic audio from the spectrum data. The vocoder can be a machine learning model or a non-machine learning model. For example, the vocoder can be a waveRNN vocoder, a waveNet vocoder, a melGan vocoder, a HifiGan vocoder, an LPCNet vocoder, etc.
[0146] In implementation, as Figure 5 shown, input the target spectrum data corresponding to the target text into the vocoder to obtain the target voice corresponding to the target text.
[0147] If the vocoder is a machine learning model, then the vocoder needs to be trained separately. The method for training the vocoder can be: obtain the speech samples and the spectrum data of the speech samples, and use the speech samples as the reference audio. Input the spectrum data into the vocoder to obtain the predicted speech. Input the speech samples and the predicted speech into the fourth loss function to obtain the fourth loss value. When the fourth loss value does not meet the preset conditions, adjust the training of the vocoder. Use other speech samples to adjust the training of the vocoder until the fourth loss value meets the preset conditions to obtain the trained vocoder.
[0148] Figure 6 is a schematic structural diagram of a device for training a spectrum synthesis model provided by an embodiment of the present application. Refer to Figure 6 , the device includes:
[0149] An input module 610, configured to input training samples into the initial spectrum synthesis model, where the training samples include text samples, corresponding speech samples, and standard intent vectors;
[0150] An extraction module 620, configured to extract the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the speech sample, and the predicted intent vector corresponding to the speech sample;
[0151] The first determination module 630 is configured to determine the predicted spectral data corresponding to the text sample according to the sample text feature information and the predicted intent vector;
[0152] The second determination module 640 is configured to determine a first loss value according to the predicted spectral data and the standard spectral data;
[0153] The third determination module 650 is configured to determine a second loss value according to the predicted intent vector and the standard intent vector;
[0154] The parameter tuning module 660 is configured to perform parameter tuning on the initial spectral synthesis model according to the first loss value and the second loss value;
[0155] The first judgment module 670 is configured to, if a preset training end condition is satisfied, determine the parameter-tuned initial spectral synthesis model as the trained spectral synthesis model;
[0156] The second judgment module 680 is configured to, if the preset training end condition is not satisfied, continue to perform parameter tuning on the parameter-tuned initial spectral synthesis model according to other training samples.
[0157] Optionally, the spectral synthesis model includes a text encoder, a voice encoder, and a first self-attention learning module;
[0158] The extraction module 620 is configured to:
[0159] Input the text sample into the text encoder to obtain the sample text feature information corresponding to the text sample;
[0160] Extract the spectral data of the voice sample as the standard spectral data;
[0161] Input the standard spectral data corresponding to the voice sample into the voice encoder to obtain a voice vector;
[0162] Input the voice vector into the first self-attention learning module to obtain the predicted intent vector of the voice sample, where the predicted intent vector is formed by arranging the possibility values corresponding to each intent classification in a preset order.
[0163] Optionally, the spectral synthesis model includes the intent feature information corresponding to each intent classification and a spectral synthesis module;
[0164] The first determination module 630 is configured to:
[0165] Multiply the possibility values of each intention classification in the predicted intention vector by the intention feature information corresponding to each intention classification stored in advance, and combine the multiplication results to obtain the sample intention feature information corresponding to the speech sample;
[0166] Input the sample text feature information and the sample intention feature information into the spectrum synthesis module to obtain the predicted spectrum data corresponding to the text sample.
[0167] Optionally, the parameter tuning module 660 is configured to:
[0168] Determine a comprehensive loss value according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value;
[0169] Tune the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
[0170] Optionally, the third determination module 650 is configured to:
[0171] Determine the cross entropy between the predicted intention vector and the standard intention vector, and use the cross entropy as the second loss value.
[0172] It should be noted that when the device for training a spectrum synthesis model provided in the above embodiment trains the spectrum synthesis model, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for training a spectrum synthesis model provided in the above embodiment and the method embodiment for training a spectrum synthesis model belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0173] Figure 7 It is a schematic structural diagram of a device for synthesizing audio provided by an embodiment of the present application. Refer to Figure 7 and the device includes:
[0174] The first input module 710 is configured to input the target text into the trained natural language processing model to obtain a target intention classification;
[0175] The determination module 720 is configured to determine a target intention vector corresponding to the target intention classification;
[0176] The second input module 730 is configured to input the target text and the target intention vector into the trained spectrum synthesis model, and determine the target spectrum data corresponding to the target text according to the target text and the target intention vector;
[0177] A third input module 740, configured to input target spectrum data corresponding to the target text into a vocoder to obtain target speech corresponding to the target text.
[0178] Optionally, when the spectrum synthesis model includes a text encoder, intention feature information corresponding to each intention classification, and a spectrum synthesis module, the second input module 730 is configured to:
[0179] Input the target text into the text encoder to obtain target text feature information corresponding to the target text;
[0180] Multiply the probability values of each intention classification in the target intention vector by the intention feature information corresponding to each intention classification stored in advance, and combine the multiplication results to obtain target intention feature information corresponding to the target text;
[0181] Input the target text feature information and the target intention feature information into the spectrum synthesis module to obtain target spectrum data corresponding to the target text.
[0182] Optionally, the determination module 720 is configured to:
[0183] Set the probability value corresponding to the target intention classification to 1, set the probability values corresponding to other intention classifications to 0, and sort the probability values corresponding to each intention classification in a preset order to obtain a target intention vector corresponding to the target intention classification.
[0184] It should be noted that when the device for synthesizing audio provided in the above embodiment synthesizes audio, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for synthesizing audio provided in the above embodiment and the method embodiment for synthesizing audio belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0185] Figure 8The structural block diagram of a terminal 800 provided by an exemplary embodiment of the present application is shown. The terminal 800 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 800 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0186] Generally, the terminal 800 includes: a processor 801 and a memory 802.
[0187] The processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 801 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0188] The memory 802 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one program code, and the at least one program code is used to be executed by the processor 801 to implement the methods for training a spectral synthesis model and synthesizing audio provided by the method embodiments in the present application.
[0189] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 808, a positioning assembly 808, and a power supply 809.
[0190] The peripheral device interface 803 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0191] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 804 may communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0192] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input to the processor 801 as control signals for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 805, which is provided on the front panel of the terminal 800; in other embodiments, there may be at least two display screens 805, which are respectively provided on different surfaces of the terminal 800 or are in a foldable design; in other embodiments, the display screen 805 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 800. Even more, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0193] The camera module 806 is used to capture images or videos. Optionally, the camera module 806 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0194] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 807 may further include a headphone jack.
[0195] The positioning component 808 is used to locate the current geographical location of the terminal 800 to achieve navigation or LBS (Location Based Service). The positioning component 808 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0196] The power supply 809 is used to supply power to each component in the terminal 800. The power supply 809 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 809 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0197] In some embodiments, the terminal 800 further includes one or more sensors 810. The one or more sensors 810 include but are not limited to: an acceleration sensor 811, a gyroscope sensor 812, a pressure sensor 813, a fingerprint sensor 814, an optical sensor 815, and a proximity sensor 816.
[0198] The acceleration sensor 811 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 800. For example, the acceleration sensor 811 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 811. The acceleration sensor 811 can also be used for collecting game or user's motion data.
[0199] The gyroscope sensor 812 can detect the body direction and rotation angle of the terminal 800. The gyroscope sensor 812 can cooperate with the acceleration sensor 811 to collect the 3D actions of the user on the terminal 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0200] The pressure sensor 813 can be disposed on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side frame of the terminal 800, it can detect the holding signal of the user on the terminal 800, and the processor 801 can identify the left and right hands or perform a quick operation according to the holding signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 805. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0201] The fingerprint sensor 814 is used to collect the fingerprint of the user. The processor 801 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 814, or the fingerprint sensor 814 can identify the user's identity according to the collected fingerprint. When the identity of the user is identified as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 814 can be disposed on the front, back, or side of the terminal 800. When there are physical buttons or a manufacturer's logo on the terminal 800, the fingerprint sensor 814 can be integrated with the physical buttons or the manufacturer's logo.
[0202] The optical sensor 815 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 according to the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera module 806 according to the ambient light intensity collected by the optical sensor 815.
[0203] The proximity sensor 816, also known as a distance sensor, is usually disposed on the front panel of the terminal 800. The proximity sensor 816 is used to collect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from the lit state to the off state; when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from the off state to the lit state.
[0204] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the terminal 800, and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0205] The computer device provided by the embodiments of the present application can be provided as a server. Figure 8 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 801 and one or more memories 802. Among them, at least one program code is stored in the memory 802, and the at least one program code is loaded and executed by the processor 801 to implement the synthetic audio method provided by each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input interface for input, and the server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0206] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code, and the above program code can be executed by a processor in a terminal or a server to complete the synthetic audio method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact-disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0207] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or by hardware related to program code. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0208] The foregoing are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for training a spectrum synthesis model, characterized in that, the method includes: Inputting training samples into an initial spectrum synthesis model, where the training samples include text samples, corresponding speech samples, and standard intent vectors; Extracting the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the speech sample, and the predicted intent vector corresponding to the speech sample; Determining the predicted spectrum data corresponding to the text sample according to the sample text feature information and the predicted intent vector; Determining a first loss value according to the predicted spectrum data and the standard spectrum data; Determining a second loss value according to the predicted intent vector and the standard intent vector; Adjusting the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value; If the preset training end condition is satisfied, determining the initial spectrum synthesis model after parameter adjustment as the trained spectrum synthesis model; If the preset training end condition is not satisfied, continuing to adjust the parameters of the initial spectrum synthesis model after parameter adjustment according to other training samples.
2. The method according to claim 1, characterized in that, the spectrum synthesis model includes a text encoder, a speech encoder, and a first self-attention learning module; The extracting the sample text feature information corresponding to the text sample, the standard spectrum data corresponding to the speech sample, and the predicted intent vector corresponding to the speech sample includes: Inputting the text sample into the text encoder to obtain the sample text feature information corresponding to the text sample; Extracting the spectrum data of the speech sample as the standard spectrum data; Inputting the standard spectrum data corresponding to the speech sample into the speech encoder to obtain a speech vector; Inputting the speech vector into the first self-attention learning module to obtain the predicted intent vector of the speech sample, where the predicted intent vector is composed of the possibility values corresponding to each intent classification arranged in a preset order.
3. The method according to claim 1, characterized in that, the spectrum synthesis model includes intent feature information corresponding to each intent classification and a spectrum synthesis module; The determining the predicted spectrum data corresponding to the text sample according to the sample text feature information and the predicted intent vector includes: Multiplying the possibility value of each intent classification in the predicted intent vector by the intent feature information corresponding to each intent classification stored in advance, and combining the multiplication results to obtain the sample intent feature information corresponding to the speech sample; Inputting the sample text feature information and the sample intent feature information into the spectrum synthesis module to obtain the predicted spectrum data corresponding to the text sample.
4. The method according to claim 1, characterized in that, the adjusting the parameters of the initial spectrum synthesis model according to the first loss value and the second loss value includes: Determining a comprehensive loss value according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value; Adjusting the parameters of the initial spectrum synthesis model according to the comprehensive loss value.
5. The method according to claim 1, It is characterized in that determining a second loss value according to the predicted intention vector and the standard intention vector includes: determining the cross entropy between the predicted intention vector and the standard intention vector, and using the cross entropy as the second loss value.
6. A method for synthesizing audio, It is characterized in that the method includes: inputting a target text into a trained natural language processing model to obtain a target intention classification; determining a target intention vector corresponding to the target intention classification; inputting the target text and the target intention vector into the trained spectrum synthesis model according to any one of claims 1-5, and determining target spectrum data corresponding to the target text according to the target text and the target intention vector; inputting the target spectrum data corresponding to the target text into a vocoder to obtain target speech corresponding to the target text.
7. The method according to claim 6, It is characterized in that when the spectrum synthesis model includes a text encoder, intention feature information corresponding to each intention classification, and a spectrum synthesis module, determining the target spectrum data corresponding to the target text according to the target text and the target intention vector includes: inputting the target text into the text encoder to obtain target text feature information corresponding to the target text; multiplying the probability value of each intention classification in the target intention vector by the intention feature information corresponding to each intention classification stored in advance respectively, and combining the multiplication results to obtain target intention feature information corresponding to the target text; inputting the target text feature information and the target intention feature information into the spectrum synthesis module to obtain target spectrum data corresponding to the target text.
8. The method according to claim 6, It is characterized in that determining the target intention vector corresponding to the target intention classification includes: setting the probability value corresponding to the target intention classification to 1, setting the probability values corresponding to other intention classifications to 0, and sorting the probability values corresponding to each intention classification in a preset order to obtain the target intention vector corresponding to the target intention classification.
9. A terminal, It is characterized in that the terminal includes a processor and a memory, and at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to implement the operations performed by the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, It is characterized in that at least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multimedia audio synthesis method and device, electronic equipment and storage medium
CN112614477A
Speech synthesis model training method, speech synthesis method and device thereof
CN113393828A