A training method, device, medium and electronic device for a speech synthesis model
By correcting errors and noise in audio data, the method trains speech synthesis models effectively using lower-quality audio, addressing the high data quality requirements of existing methods and reducing collection costs.
Patent Information
- Application Number
- CN202410248274.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-03-04
AI Technical Summary
The training data of existing speech synthesis models require the quality of pronunciation readers, resulting in high training costs and low efficiency, and the inability to effectively utilize low-quality audio.
By identifying characters and noise segments in the original audio, adjusting the original text to match the audio content, and marking the noise segments, using sample text and noise segments to train the speech synthesis model to reduce the requirements for audio quality.
The cost of the speech synthesis model training data is reduced, the recording efficiency is improved, and the high-quality speech synthesis model can be trained with more reads, less reads, misreads and low-quality audio.
Smart Images

Figure CN118230712B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, medium and electronic device for training a speech synthesis model. Background Art
[0002] With the continuous development of technology, speech synthesis technology has received extensive attention. Speech synthesis technology is a technology that converts text into speech. Among them, the speech synthesis model has a good effect of converting text into speech, so the speech synthesis model is widely used in various speech interaction scenarios, such as: reading aloud for listening, restaurant queue calling, etc.
[0003] Generally, the audio in the training data of the speech synthesis model is obtained by specifying text to a speaker and asking the speaker to read aloud according to the text. However, in order to obtain high-quality audio, the speech synthesis model has very strict requirements for training data, that is, it does not allow the speaker to stutter, cough, etc. during the recording of audio to ensure the fluency, accuracy, etc. of the audio. This adds difficulties to the training of the speech synthesis model.
[0004] Based on this, this specification provides a method for training a speech synthesis model. Summary of the Invention
[0005] This specification provides a method, apparatus, medium and electronic device for training a speech synthesis model to at least partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a method for training a speech synthesis model, the method comprising:
[0008] Obtain the original text, and obtain the original audio corresponding to the original text;
[0009] Identify the characters in the original audio;
[0010] Replace the characters in the original text with the characters identified from the original audio to obtain a sample text; and, mark the noise segments in the original audio;
[0011] Train the speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model.
[0012] Optionally, replacing the characters in the original text with the characters identified from the original audio specifically includes:
[0013] If the characters of the recognized original audio are obtained by reading more characters in the original text, add the characters of the recognized original audio to the original text; and / or
[0014] If it is determined that there are characters in the original text that are not read enough according to the characters of the recognized original audio, delete the characters that are not read enough in the original text; and / or
[0015] If the characters of the recognized original audio are obtained by misreading the characters in the original text, modify the characters in the original text to the characters of the recognized original audio.
[0016] Optionally, mark the noise segments in the original audio, specifically including:
[0017] Take the audio frames in the original audio that meet the specified conditions as noise segments and mark them; wherein, the specified conditions at least include the unsmooth audio frames in the original audio corresponding to the original text.
[0018] Optionally, train the speech synthesis model to be trained according to the obtained sample text and the marked noise segments, specifically including:
[0019] Input the sample text into the speech synthesis model to be trained, and obtain the synthesized audio output by the speech synthesis model to be trained;
[0020] Determine the loss according to the difference between the synthesized audio and the audio segments other than the noise segments in the original audio;
[0021] Train the speech synthesis model to be trained according to the loss to obtain the trained speech synthesis model.
[0022] Optionally, the speech synthesis model includes a Mel spectrum conversion network and an audio synthesis network.
[0023] Optionally, input the sample text into the speech synthesis model to be trained, and obtain the synthesized audio output by the speech synthesis model to be trained, specifically including:
[0024] Determine the phoneme sequence of the sample text;
[0025] Input the phoneme sequence into the Mel spectrum conversion network in the speech synthesis model to be trained, and obtain the first Mel spectrum output by the Mel spectrum conversion network;
[0026] Input the first Mel spectrum into the audio synthesis network in the speech synthesis model to be trained, and obtain the synthesized audio output by the audio synthesis network.
[0027] Optionally, a loss is determined according to the difference between the synthesized audio and the audio segments other than the noise segment in the original audio, specifically including:
[0028] Determine the second Mel spectrogram of the audio segments other than the noise segment;
[0029] Determine a first loss according to the second Mel spectrogram and the first Mel spectrogram;
[0030] Determine a second loss according to the audio segments other than the noise segment and the synthesized audio.
[0031] Optionally, determining a first loss according to the second Mel spectrogram and the first Mel spectrogram specifically includes:
[0032] Determine a first loss according to at least one of the difference between the second Mel spectrogram and the first Mel spectrogram, the difference between the fundamental frequency of the phoneme in the second Mel spectrogram and the fundamental frequency of the phoneme in the first Mel spectrogram, the difference between the energy of the phoneme in the second Mel spectrogram and the energy of the phoneme in the first Mel spectrogram, and the difference between the frame length of the phoneme in the second Mel spectrogram and the frame length of the phoneme in the first Mel spectrogram.
[0033] Optionally, training the speech synthesis model to be trained according to the loss specifically includes:
[0034] Train the Mel spectrogram conversion network according to the first loss;
[0035] Train the audio synthesis network according to the second loss.
[0036] This specification provides a training device for a speech synthesis model, including:
[0037] An acquisition module, configured to acquire an original text and acquire the original audio corresponding to the original text;
[0038] An identification module, configured to identify characters in the original audio;
[0039] A processing module, configured to replace the characters in the original text with the characters identified from the original audio to obtain a sample text; and mark the noise segments in the original audio;
[0040] A training module, configured to train the speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model.
[0041] Optionally, the processing module is specifically configured to, if the characters of the original audio recognized are obtained by reading more characters in the original text, add the characters of the recognized original audio to the original text; and / or if it is determined that there are characters that are not read enough in the original text according to the characters of the recognized original audio, delete the characters that are not read enough in the original text; and / or if the characters of the recognized original audio are obtained by misreading the characters in the original text, modify the characters in the original text to the characters of the recognized original audio.
[0042] Optionally, the processing module is specifically configured to use the audio frames in the original audio that meet the specified conditions as noise segments and mark them; wherein the specified conditions at least include the unsmooth audio frames in the original audio corresponding to the original text.
[0043] Optionally, the training module is specifically configured to input the sample text into the speech synthesis model to be trained, and obtain the synthesized audio output by the speech synthesis model to be trained; determine the loss according to the difference between the synthesized audio and the audio segments other than the noise segments in the original audio; and train the speech synthesis model to be trained according to the loss to obtain the trained speech synthesis model.
[0044] Optionally, the speech synthesis model includes a Mel spectrum conversion network and an audio synthesis network.
[0045] Optionally, the training module is specifically configured to determine the phoneme sequence of the sample text; input the phoneme sequence into the Mel spectrum conversion network in the speech synthesis model to be trained, and obtain the first Mel spectrum output by the Mel spectrum conversion network; input the first Mel spectrum into the audio synthesis network in the speech synthesis model to be trained, and obtain the synthesized audio output by the audio synthesis network.
[0046] Optionally, the training module is specifically configured to determine the second Mel spectrum of the audio segments other than the noise segments; determine the first loss according to the second Mel spectrum and the first Mel spectrum; and determine the second loss according to the audio segments other than the noise segments and the synthesized audio.
[0047] Optionally, the training module is specifically configured to determine the first loss according to at least one of the difference between the second Mel spectrum and the first Mel spectrum, the difference between the fundamental frequency of the phonemes in the second Mel spectrum and the fundamental frequency of the phonemes in the first Mel spectrum, the difference between the energy of the phonemes in the second Mel spectrum and the energy of the phonemes in the first Mel spectrum, and the difference between the frame length of the phonemes in the second Mel spectrum and the frame length of the phonemes in the first Mel spectrum.
[0048] Optionally, the training module is specifically configured to train the Mel spectrogram conversion network according to the first loss; and train the audio synthesis network according to the second loss.
[0049] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method for training a speech synthesis model.
[0050] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the above method for training a speech synthesis model when executing the program.
[0051] The above at least one technical solution adopted in this specification can achieve the following beneficial effects:
[0052] In the method for training a speech synthesis model provided in this specification, first obtain the original text and the original audio corresponding to the original text. Then, identify the characters in the original audio, replace the characters in the original text with the characters identified from the original audio to obtain a sample text, and mark the noise segments in the original audio. Input the sample text into the speech synthesis model to be trained to obtain a synthesized audio. Finally, train the speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model.
[0053] It can be seen from the above method that this method can use low-quality audio with extra characters, missing characters, misread characters, and sudden noise, and still be able to train a speech synthesis model, reducing the requirements for the quality of the original audio used to train the speech synthesis model, thereby reducing the cost of collecting training data for the speech synthesis model. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and the illustrative embodiments and descriptions thereof of this specification are used to explain this specification, and do not constitute an improper limitation to this specification. In the attached
[0055] In the drawings:
[0056] Figure 1 is a schematic flowchart of a method for training a speech synthesis model in this specification;
[0057] Figure 2 is a schematic diagram of noise marking of an original audio provided in this specification;
[0058] Figure 3 is a schematic diagram of training a speech synthesis model provided in this specification;
[0059] Figure 4 Schematic diagram for training a speech synthesis model provided in this specification;
[0060] Figure 5 Schematic diagram for training a speech synthesis model provided in this specification;
[0061] Figure 6 Schematic diagram of a training device for a speech synthesis model provided in this specification;
[0062] Figure 7 Corresponding to the Figure 1 schematic diagram of an electronic device. Specific embodiments
[0063] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0064] The following will detail the technical solutions provided in each embodiment of this specification in conjunction with the drawings.
[0065] Figure 1 Schematic flowchart of a method for training a speech synthesis model provided in this specification, which may specifically include the following steps:
[0066] S100: Obtain the original text and obtain the original audio corresponding to the original text.
[0067] Generally, for the training of a speech synthesis model, when collecting training data, the speaker has to read aloud according to the given text. Once a mistake is made, the recording has to be redone. Such extremely demanding requirements increase the training cost of the speech synthesis model. Based on this, this application specification provides a method for training a speech synthesis model, which reduces the requirements for the quality of the audio and reduces the cost of collecting training data for the speech synthesis model.
[0068] The execution subject of the embodiments of this specification can be any computing device with computing capabilities (such as: a terminal, a server). Now, the server will be used as the execution subject for illustration.
[0069] The server can obtain the original text and the original audio corresponding to the original text. In one or more embodiments of this specification, when obtaining the original audio, the original audio can be the original audio recorded when the speaker reads the original text, or the original audio corresponding to the original text obtained from other channels, such as generating the original audio corresponding to the original text based on other trained intelligent models. In this specification, there is no limitation on how to obtain the original audio corresponding to the original text.
[0070] S102: Identify the characters in the original audio.
[0071] S104: Replace the characters in the original text with the characters identified from the original audio to obtain a sample text; and mark the noise segments in the original audio.
[0072] After the server obtains the original text and the original audio, the server can identify the characters in the original audio.
[0073] It should be noted that in one or more embodiments of this specification, when identifying the characters in the original audio, computerized phonetics (Praat) can be used to annotate the original audio to obtain the text content corresponding to the original audio, and the characters in the text content are the characters in the obtained original audio. Of course, other methods can also be used to identify the characters in the original audio, such as manually annotating the characters in the original audio. Specifically, this specification does not make any limitations.
[0074] Furthermore, the server can replace the characters in the original text with the characters identified in the original audio to obtain a sample text. That is, the server can modify the original text according to the characters identified in the original audio so that the modified sample text is consistent with the characters in the original audio, and the modified original text is used as the sample text.
[0075] Specifically, if the characters in the identified original audio are obtained by reading more characters in the original text, then add the characters in the identified original audio to the original text, and / or if it is determined that there are characters that are read less in the original text according to the characters in the identified original audio, then delete the characters that are read less in the original text, and / or if the characters in the identified original audio are obtained by misreading the characters in the original text, then modify the characters in the original text to the characters in the identified original audio.
[0076] For example: the original text is "The air is so fresh today". When the speaker reads the original text, the characters corresponding to the original audio obtained are "The air is so fresh today". It can be seen that the character "天" in the original audio is misread as the character "日". Therefore, in order to maintain the consistency between the original text and the original audio, the server can change the character "天" in the original text to the character "日". Similarly, when the characters in the recognized original audio are more than the characters in the original text, that is, the characters in the recognized original audio are more than the characters in the original text, the extra characters can be added to the original text. When the characters in the recognized original audio are less than the characters in the original text, the less-read characters can be deleted in the original text.
[0077] Generally, the original audio corresponding to the original text obtained will have some noise to a greater or lesser extent, such as a person's cough, a car's horn, noise generated by equipment resonance, etc., that is, due to environmental factors and equipment factors, etc., the original audio corresponding to the original text obtained will have noise. Therefore, in order to reduce the noise in the original audio, ensure the consistency between the characters in the original audio and the characters in the original text, and to improve the accuracy of the trained speech synthesis model, in one or more embodiments of this specification, the server may also mark the noise segment in the original audio. Specifically, the server may divide the original audio into audio frames, and when the audio frames in the original audio meet the specified conditions, the audio frames that meet the specified conditions may be marked in the original audio, and the mark is used to indicate that the audio frame is an audio frame with noise. Simply put, the server may treat the audio frames in the original audio that meet the specified conditions as noise segments and mark them.
[0078] In one or more embodiments of the present specification, assuming that the original text is "The air is so fresh today", and there is noise in the audio segment corresponding to the character "sky" in the corresponding original audio, the audio frame corresponding to "sky" can be determined and the audio frame corresponding to "sky" can be marked. Further assuming that one character corresponds to one audio frame, and the frame length corresponding to one character is 50 milliseconds, the original text has 7 audio frames (7 characters), and the total frame length is 350 milliseconds, then the second audio frame and the third audio frame can be used as noise segments (audio frames corresponding to "sky"), or the audio segment from 51 to 150 milliseconds in the original audio can be marked as a noise segment.
[0079] like Figure 2 The figure shows a noise marking schematic diagram of an original audio provided in the specification of this application. The original text is "this you need to write the answer", but the characters in the recognized original audio are "this you you need to write the answer". It can be seen that one more character "you" is read, so the audio frame corresponding to "you" needs to be marked as a noise frame, that is, marked as <noise>, such as Figure 2 the shaded part in Figure 2 , that is, the audio segment of the 0.413596 part.
[0080] Among them, the specified conditions at least include the unsmooth audio frames in the original audio corresponding to the original text, and the unsmoothness at least includes: there are sudden noises in the original audio (such as people's coughs, car honks, equipment resonance noises, etc.), and the audio frames corresponding to the inconsistent characters between the recognized characters in the original audio and the characters in the original text.
[0081] It should be noted that the noise segment in the original audio refers to the set of audio frames with noise in the original audio.
[0082] Moreover, in one or more embodiments of this specification, when marking the noise segment, the audio segments corresponding to the start position and the end position of the noise can be marked with <noise>to mark that the audio between this start position and this end position is noisy audio. Of course, it is also possible to perform respectively at the position where the noise starts and the position where the noise ends <noise>Marks may also be used with other marking methods, which are not specifically limited in this specification.
[0083] S106: Train the speech synthesis model to be trained based on the obtained sample text and the marked noise segments to obtain a trained speech synthesis model.
[0084] Therefore, the server can train the speech synthesis model to be trained based on the obtained sample text and the marked noise segments to obtain a trained speech synthesis model. Specifically, after the server obtains the sample text and the original audio with the marked noise segments, it can input the sample text into the speech synthesis model to be trained to obtain the synthesized audio output by the speech synthesis model to be trained, so as to train the speech synthesis model to be trained based on the original audio and the synthesized audio after the marked noise segments in the subsequent steps.
[0085] In one or more embodiments of this specification, the factor sequence of the sample text may be determined first, and then the speech synthesis model to be trained may be trained based on the factor sequence of the sample text and the marked noise segments to obtain a trained speech synthesis model. As Figure 3 shown, it is a training schematic diagram of a speech synthesis model provided in this specification. The sample text is "We will go on an outing today", and its corresponding factor sequence is "wo3men5jin1tian1qu4jiao1 you2". The audio segment corresponding to the characters "today go" in the original text contains noise, and the noise segment corresponding to the characters "today go" in the original audio corresponding to the original text can be marked. Furthermore, when training the speech synthesis model, the factor sequence of the sample text is input into the speech synthesis model to obtain the synthesized audio output by the speech synthesis model to be trained. The synthesized audio contains a noise segment, and the noise segment is consistent with the noise segment in the marked original audio. Then, the speech synthesis model is trained based on the synthesized audio and the noise segment in the marked original audio.
[0086] Specifically, in this specification, the server can determine the loss based on the difference between the synthesized audio and the audio segments other than the noise segments. Then, based on this loss, the speech synthesis model to be trained is trained to obtain a trained speech synthesis model.
[0087] It should be noted that in one or more embodiments of this specification, when calculating the loss, the loss is calculated based on the difference between the audio segment other than the noise segment and the synthesized audio output by the speech synthesis model to be trained, that is, the noise segment of the audio is not included in the loss calculation, but the sample text input to the speech synthesis model to be trained includes the text corresponding to the noise segment, that is, the sample text must contain the text corresponding to the noise segment in the synthesized audio, so that the speech synthesis model to be trained can learn the contextual information corresponding to the audio in the text. Among them, the text corresponding to the noise segment can be marked in the original text, and can be marked with [noise], or with a symbol, etc. The specific marking method is not limited in this specification.
[0088] based on Figure 1 In the training method of the speech synthesis model provided in the present specification, the original text is modified according to the characters that are read less, read more, or read incorrectly in the original audio to ensure the consistency between the original text and the original audio, and the modified original text is used as a sample text, the noise segment in the original audio is marked, and the audio segment other than the noise segment is used as a label. According to the sample text and the label, the speech synthesis model to be trained is trained to obtain the trained speech synthesis model. That is, the method can still train a speech synthesis model when there are more read characters, less read characters, and wrongly read characters in the original audio corresponding to the original text, or the original audio is a low-quality audio with other sudden noises, without using the high-quality audio in which the characters in the recognized original audio are exactly the same as the original text, thereby reducing the requirements for the quality of the audio, and when the original audio is recorded based on the pronunciation of the speaker, when the speaker reads the text incorrectly, there is no need to re-record the audio, which greatly improves the recording efficiency. According to actual measurements, compared with the prior art, the method provided by the embodiment of the present specification can increase the recording efficiency by more than two times, thereby reducing the cost of collecting training data for the speech synthesis model.
[0089] Furthermore, in order to reduce the noise in the original audio so that the trained speech synthesis model has a better effect, in one or more embodiments of the present specification, in the above step S104, before marking the noise segment in the original audio, each audio frame of the original audio can also be determined, and the audio frames whose signal-to-noise ratio is not greater than the preset threshold in each audio frame can be deleted to remove the audio frames without the speaker's voice. For example: the original text is "Today the air is really fresh.", and there is a long noise segment between the characters "天" and "空" in the original audio, then the audio frame corresponding to the long noise segment can be deleted.
[0090] In addition, the server may further perform noise reduction processing on the audio frames whose signal-to-noise ratio is greater than a preset threshold in each audio frame to suppress the influence of noise.
[0091] It should be noted that after noise reduction processing, the audio segments corresponding to the noise that has not been eliminated in the original audio are also the noise segments described in step S104 above. That is to say, the specified conditions in step S104 above also include: the audio frames corresponding to the noise that has not been eliminated after the original audio has been subjected to noise reduction processing.
[0092] Among them, in one or more embodiments of this specification, a Deep Complex Convolution Recurrent Network (DCCRN) can be used to perform noise reduction processing on audio frames with a signal-to-noise ratio greater than a preset threshold. And the audio after noise reduction processing can also be enhanced for the human voice or the voice of the read characters in the original audio through Adobe Audition software.
[0093] In one or more embodiments of this specification, the speech synthesis model to be trained at least includes: a Mel spectrogram conversion network and an audio synthesis network.
[0094] Among them, in one or more embodiments of this specification, the Mel spectrogram conversion network can be FastSpeech2, and the audio synthesis network can be HifiGan. For FastSpeech2, its training samples are "phoneme sequence-Mel spectrogram pairs", and for HifiGan, its training samples are "Mel spectrogram-audio pairs".
[0095] Then in step S106 above, when inputting the sample text into the speech synthesis model to be trained, the server can first determine the phoneme sequence of the sample text (if the sample text is in Chinese, the phoneme sequence is the pinyin sequence with tones). For example, if the sample text is: "We are going on an outing today", then the phoneme sequence of the sample text is: "wo3men5jin1tian1qu4jiao1you2". Then, input this phoneme sequence into the Mel spectrogram conversion network in the speech synthesis model to be trained to obtain the first Mel spectrogram output by the Mel spectrogram conversion network. Finally, input the first Mel spectrogram into the audio synthesis network in the speech synthesis model to be trained to obtain the synthesized audio output by the audio synthesis network.
[0096] Then in step S108 above, when determining the loss according to the difference between the synthesized audio and the audio segments other than the noise segments in the original audio, the server can determine the second Mel spectrogram of the audio segments other than the noise segments, and determine the first loss according to the second Mel spectrogram and the first Mel spectrogram, and determine the second loss according to the audio segments other than the noise segments and the synthesized audio.
[0097] Furthermore, the server can train the Mel spectrogram conversion network according to the first loss and train the audio synthesis network according to the second loss.
[0098] In addition, in one or more embodiments of this specification, when determining the first loss based on the second Mel spectrogram and the first Mel spectrogram, the first loss can be determined based on the differences in multiple dimensions between the first Mel spectrogram and the second Mel spectrogram. In one or more embodiments of this specification, the differences in the multiple dimensions include the difference in fundamental frequency, the difference in frame length, the energy difference, and the overall difference between the first Mel spectrogram and the second Mel spectrogram itself.
[0099] As Figure 4 shown, it is a training schematic diagram of a speech synthesis model provided in the specification of this application. It can be seen that when training the Mel spectrogram conversion network, i.e., FastSpeech2, in the speech synthesis model, the factor sequence corresponding to the sample text can be input into the Mel spectrogram conversion network, and the Mel spectrogram conversion network outputs the first Mel spectrogram. The first Mel spectrogram contains the Mel spectrogram corresponding to the noise segment. Therefore, the server can use the Mel spectrogram corresponding to the audio segment other than the noise segment as the second Mel spectrogram, that is, the labeled Mel spectrogram, and train the Mel spectrogram conversion network, i.e., FastSpeech2, based on the second Mel spectrogram and the first Mel spectrogram. Specifically, the server can determine the first loss according to at least one of the differences between the second Mel spectrogram and the first Mel spectrogram, the difference in the fundamental frequency of the phonemes in the second Mel spectrogram and the fundamental frequency of the phonemes in the first Mel spectrogram, the difference in the energy of the phonemes in the second Mel spectrogram and the energy of the phonemes in the first Mel spectrogram, and the difference in the frame length of the phonemes in the second Mel spectrogram and the frame length of the phonemes in the first Mel spectrogram.
[0100] That is to say, when training the Mel spectrogram conversion network, the first loss between the first Mel spectrogram and the second Mel spectrogram can be calculated through the differences in multiple dimensions.
[0101] Specifically, the server can use the Mel spectrogram loss function: to determine the difference between the second Mel spectrogram and the first Mel spectrogram, that is, the difference in the waveforms corresponding to the first Mel spectrogram and the second Mel spectrogram. Where T is the number of audio frames in the original audio excluding the noise segment, is the t-th frame of the first Mel spectrogram predicted by the speech synthesis model, is the t-th frame of the second Mel spectrogram corresponding to the original audio excluding the noise segment.
[0102] The server can use the energy loss function: to determine the difference in the energy of the phonemes in the second Mel spectrogram and the energy of the phonemes in the first Mel spectrogram. Where P is the number of all phonemes corresponding to the audio segment in the original audio excluding the noise segment, is the energy of the phoneme p corresponding to the first Mel spectrogram predicted by the speech synthesis model, is the energy of the phoneme p corresponding to the second Mel spectrogram of the original audio excluding the noise segment
[0103] The server can use the frame length loss function: to determine the difference in the frame length of the phoneme in the second Mel spectrogram and the frame length of the phoneme in the first Mel spectrogram. Where P is the number of all phonemes corresponding to the audio segment of the original audio excluding the noise segment is the continuous frame length of the phoneme p corresponding to the first Mel spectrogram predicted by the speech synthesis model is the continuous frame length of the phoneme p corresponding to the second Mel spectrogram of the original audio excluding the noise segment
[0104] The server can use the fundamental frequency loss function: to determine the difference in the fundamental frequency of the phoneme in the second Mel spectrogram and the fundamental frequency of the phoneme in the first Mel spectrogram. Where P is the number of all phonemes corresponding to the audio segment of the original audio excluding the noise segment is the fundamental frequency of the phoneme p corresponding to the first Mel spectrogram predicted by the speech synthesis model is the fundamental frequency of the phoneme p corresponding to the second Mel spectrogram of the original audio excluding the noise segment
[0105] Then, the first loss can be the sum of each loss, that is, L = L mel + L e + L d + L p , where L is the first loss. Of course, each loss can also be weighted and then summed to obtain the first loss
[0106] Among them, the server can obtain the fundamental frequency and energy corresponding to each phoneme in the original text through the WORLD algorithm
[0107] It should be noted that in one or more embodiments of this specification, when determining the loss, the loss functions of multiple dimensions used above are determined based on FastSpeech2. Therefore, when using other networks, the loss functions used may change, and this specification does not limit which network and which loss function to use specifically
[0108] As Figure 5 shown, it is a training schematic diagram of a speech synthesis model provided in the specification of this application. It can be seen that the server can input the noisy Mel spectrogram output by the Mel spectrogram conversion network into the audio conversion network. Then, the audio conversion network can output the synthesized audio. Then, the server can determine the second loss based on the audio segment of the original audio excluding the noise segment and the synthesized audio, and train the audio conversion network based on the second loss
[0109] Moreover, the speech synthesis model to be trained may not include a Mel spectrogram conversion network and an audio synthesis network. That is to say, the speech synthesis model to be trained can be directly an end-to-end machine learning model. That is, when the sample text is input into the speech synthesis model, the output obtained is the audio generated according to the sample text. Then, when training the speech synthesis model to be trained, the above Figure 1 shown method can be used to train the speech synthesis model to be trained. That is, the sample text in step S104 can be input into the speech synthesis model to be trained, and the speech synthesis model to be trained is trained with the goal of minimizing the difference between the audio output by the speech synthesis model to be trained and the audio segment other than the noise segment.
[0110] It should be noted that since the running speed of the speech synthesis model is fast and stable when the speech synthesis model includes a Mel spectrogram conversion network, and the quality of the audio synthesized by the speech synthesis model is high when the speech synthesis model is an end-to-end model, therefore, whether to select an end-to-end model or a model including a Mel spectrogram conversion network for the speech synthesis model can be determined according to the actual situation and different requirements, and the specific description in this specification is not limited.
[0111] Based on the training method of the speech synthesis model described above, the embodiments of this specification also correspondingly provide a schematic diagram of a training device for a speech synthesis model, as Figure 6 shown.
[0112] Figure 6 A schematic diagram of a training device for a speech synthesis model provided by an embodiment of this specification, the device includes:
[0113] An acquisition module 600, configured to acquire an original text and acquire the original audio corresponding to the original text;
[0114] An identification module 602, configured to identify characters in the original audio;
[0115] A processing module 604, configured to replace the characters in the original text with the characters identified from the original audio to obtain a sample text; and mark the noise segments in the original audio;
[0116] A training module 606, configured to train a speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model.
[0117] Optionally, the processing module 604 is specifically configured to, if the characters of the original audio recognized are obtained by reading more characters in the original text, add the characters of the recognized original audio to the original text; and / or if it is determined that there are characters in the original text that are missed according to the characters of the recognized original audio, delete the missed characters in the original text; and / or if the characters of the recognized original audio are obtained by misreading the characters in the original text, modify the characters in the original text to the characters of the recognized original audio.
[0118] Optionally, the processing module 604 is specifically configured to use the audio frames in the original audio that meet the specified conditions as noise segments and mark them; where the specified conditions at least include the unsmooth audio frames in the original audio corresponding to the original text.
[0119] Optionally, the training module 606 is specifically configured to input the sample text into the speech synthesis model to be trained, and obtain the synthesized audio output by the speech synthesis model to be trained; determine the loss according to the difference between the synthesized audio and the audio segments other than the noise segments in the original audio; and train the speech synthesis model to be trained according to the loss to obtain the trained speech synthesis model.
[0120] Optionally, the speech synthesis model includes a Mel spectrum conversion network and an audio synthesis network.
[0121] Optionally, the training module 606 is specifically configured to determine the phoneme sequence of the sample text; input the phoneme sequence into the Mel spectrum conversion network in the speech synthesis model to be trained, and obtain the first Mel spectrum output by the Mel spectrum conversion network; input the first Mel spectrum into the audio synthesis network in the speech synthesis model to be trained, and obtain the synthesized audio output by the audio synthesis network.
[0122] Optionally, the training module 606 is specifically configured to determine the second Mel spectrum of the audio segments other than the noise segments; determine the first loss according to the second Mel spectrum and the first Mel spectrum; and determine the second loss according to the audio segments other than the noise segments and the synthesized audio.
[0123] Optionally, the training module 606 is specifically configured to determine the first loss according to at least one of the difference between the second Mel spectrum and the first Mel spectrum, the difference between the fundamental frequency of the phonemes in the second Mel spectrum and the fundamental frequency of the phonemes in the first Mel spectrum, the difference between the energy of the phonemes in the second Mel spectrum and the energy of the phonemes in the first Mel spectrum, and the difference between the frame length of the phonemes in the second Mel spectrum and the frame length of the phonemes in the first Mel spectrum.
[0124] Optionally, the training module 606 is specifically configured to train the Mel spectrogram conversion network according to the first loss; and train the audio synthesis network according to the second loss.
[0125] An embodiment of this specification also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the training method of the speech synthesis model described above.
[0126] Based on the training method of the speech synthesis model described above, an embodiment of this specification also proposes Figure 7 a schematic structural diagram of the electronic device shown. As Figure 7 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the training method of the speech synthesis model described above.
[0127] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.
[0128] In the 1990s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without asking a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0129] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0130] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0131] For the convenience of description, when describing the above devices, they are described as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0132] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0133] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 or means for implementing the functions specified in one block or more blocks.
[0134] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 or one block or more blocks.
[0135] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 one flow or more flows and / or blocks Figure 1 or one block or more blocks.
[0136] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0137] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.
[0138] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0139] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.
[0140] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0142] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.
[0143] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.< / noise> < / noise> < / noise>
Claims
1. A training method for a speech synthesis model, characterized in that, The method includes: Obtaining the original text and the original audio corresponding to the original text; Identifying the characters in the original audio; Replacing the characters in the original text with the characters identified from the original audio to obtain a sample text; and marking the noise segments in the original audio; Training a speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model; Among them, training the speech synthesis model to be trained according to the obtained sample text and the marked noise segments specifically includes: Inputting the sample text into the speech synthesis model to be trained to obtain a synthesized audio output by the speech synthesis model to be trained; Determining the phoneme sequence of the sample text and determining the first Mel spectrogram according to the phoneme sequence; Determining the second Mel spectrogram of the audio segment in the original audio except the noise segment; Determining a first loss according to the second Mel spectrogram and the first Mel spectrogram; Determining a second loss according to the audio segment except the noise segment and the synthesized audio; Training the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.
2. The method according to claim 1, wherein Replacing the characters in the original text with the characters identified from the original audio specifically includes: If the character of the original audio identified is obtained by reading the character in the original text multiple times, adding the character of the original audio identified in the original text; and / or If it is determined that there are characters read less in the original text according to the character of the original audio identified, deleting the characters read less in the original text; and / or If the character of the original audio identified is obtained by misreading the character in the original text, modifying the character in the original text to the character of the original audio identified.
3. The method according to claim 1, characterized in that, Marking the noise segments in the original audio specifically includes: Taking the audio frames in the original audio that meet the specified conditions as noise segments and marking them; where the specified conditions at least include the unsmooth audio frames in the original audio corresponding to the original text.
4. The method according to claim 1, wherein The speech synthesis model includes a Mel spectrogram conversion network and an audio synthesis network.
5. The method according to claim 4, wherein Inputting the sample text into the speech synthesis model to be trained to obtain a synthesized audio output by the speech synthesis model to be trained specifically includes: Determining the phoneme sequence of the sample text; Inputting the phoneme sequence into the Mel spectrogram conversion network in the speech synthesis model to be trained to obtain the first Mel spectrogram output by the Mel spectrogram conversion network; Inputting the first Mel spectrogram into the audio synthesis network in the speech synthesis model to be trained to obtain the synthesized audio output by the audio synthesis network.
6. The method according to claim 1, characterized in that, Determining a first loss according to the second Mel spectrogram and the first Mel spectrogram specifically includes: Determine a first loss based on at least one of the difference between the second Mel spectrogram and the first Mel spectrogram, the difference between the fundamental frequency of the phoneme in the second Mel spectrogram and the fundamental frequency of the phoneme in the first Mel spectrogram, the difference between the energy of the phoneme in the second Mel spectrogram and the energy of the phoneme in the first Mel spectrogram, and the difference between the frame length of the phoneme in the second Mel spectrogram and the frame length of the phoneme in the first Mel spectrogram.
7. The method according to claim 5, wherein Train the speech synthesis model to be trained according to the loss, specifically including: Train the Mel spectrogram conversion network according to the first loss; Train the audio synthesis network according to the second loss.
8. A training device for a speech synthesis model, characterized in that, The device specifically includes: An acquisition module, configured to acquire an original text and acquire the original audio corresponding to the original text; An identification module, configured to identify characters in the original audio; A processing module, configured to replace the characters in the original text with the characters identified from the original audio to obtain a sample text; and mark the noise segments in the original audio; A training module, configured to train the speech synthesis model to be trained according to the obtained sample text and the marked noise segments to obtain a trained speech synthesis model; The training module is specifically configured to input the sample text into the speech synthesis model to be trained to obtain a synthesized audio output by the speech synthesis model to be trained; determine the phoneme sequence of the sample text and determine a first Mel spectrogram according to the phoneme sequence; determine a second Mel spectrogram of the audio segment other than the noise segment in the original audio; determine a first loss according to the second Mel spectrogram and the first Mel spectrogram; determine a second loss according to the audio segment other than the noise segment and the synthesized audio; and train the speech synthesis model to be trained according to the loss to obtain a trained speech synthesis model.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-7 above is implemented.
10. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method according to any one of claims 1-7 above is implemented.
Citation Information
Patent Citations
Speech annotation method and device
CN109300468A
Specific speaker speech synthesis method and device
CN115101046A