Speech synthesis method, device and equipment and readable storage medium
By fine-tuning the noiseless audio of the target pronunciation person as the label signal in the speech synthesis model, the problem of noise influence is solved and high-quality personalized speech synthesis is achieved.
Patent Information
- Application Number
- CN202410216733.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-08-29
AI Technical Summary
During the fine-tuning process, the existing personalized speech synthesis model uses noisy user recording as the label signal, resulting in poor audio quality and cannot effectively simulate the voice of the target pronunciation person.
By fine-tuning the vocoder in the speech synthesis model using the noiseless audio of the target pronunciation person as the label signal, a high-quality personalized speech synthesis model is generated to avoid the introduction of noise signals.
The audio quality of the personalized speech synthesis model is improved, making the generated audio closer to the voice of the target pronunciation person, and improving the effect of speech synthesis.
Smart Images

Figure CN120564692A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of speech synthesis technology, and in particular to a speech synthesis method, apparatus, device, and readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence (AI), text-to-speech (TTS) technology has also been widely used. Personalized speech synthesis refers to synthesizing audio that sounds like a specific person.
[0003] To achieve personalized speech synthesis, the traditional approach is to pre-train a basic speech synthesis model based on a large sample set. This basic model is used to convert text information into audible natural language, but it cannot convert text information into personalized natural language. The basic model is then fine-tuned based on a small amount of audio recorded from the target speaker, resulting in a personalized speech synthesis model.
[0004] However, the target speaker is usually recorded using a portable device such as a mobile phone. The recorded audio is noisy, resulting in poor quality of the fine-tuned personalized speech synthesis model, and further resulting in poor quality of the audio generated using the personalized speech synthesis model. Summary of the Invention
[0005] The embodiments of the present application provide a speech synthesis method, apparatus, device, and readable storage medium, which generate target audio through a speech synthesis model. When fine-tuning the speech synthesis model, the noise-free audio of the target speaker is used as a label signal to fine-tune the first vocoder, thereby achieving the purpose of improving the quality of the target audio.
[0006] In a first aspect, an embodiment of the present application provides a personalized speech synthesis method, comprising:
[0007] Training an acoustic sub-model and a first vocoder using a sample set comprising a plurality of audio-text pairs, and generating a first speech synthesis model using the acoustic sub-model and the first vocoder;
[0008] Fine-tuning the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model;
[0009] The second speech synthesis model is used to perform speech synthesis on the text to be processed to obtain target audio that simulates the pronunciation of the target speaker.
[0010] In a second aspect, an embodiment of the present application provides a personalized speech synthesis device, including:
[0011] A training module, configured to train an acoustic sub-model and a first vocoder using a sample set comprising a plurality of audio-text pairs, and generate a first speech synthesis model using the acoustic sub-model and the first vocoder;
[0012] a processing module, configured to fine-tune the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model;
[0013] The synthesis module is used to perform speech synthesis on the text to be processed using the second speech synthesis model to obtain target audio that simulates the pronunciation of the target speaker.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the computer program, the electronic device implements the method described in the first aspect or various possible implementation methods of the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed by a processor, they are used to implement the method described in the first aspect or various possible implementation methods of the first aspect.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computing program, which, when executed by a processor, implements the method described in the first aspect or various possible implementations of the first aspect.
[0017] The speech synthesis method, apparatus, device, and readable storage medium provided by the present application are described. An electronic device pre-trains an acoustic sub-model and a first vocoder using a sample set containing multiple audio-text pairs. The acoustic sub-model and the first vocoder are then used to generate a first speech synthesis model. The first vocoder in the first speech synthesis model is then fine-tuned using the noise-free audio of a target speaker as a label signal to obtain a second speech synthesis model. The second speech synthesis model comprises a text front-end, an acoustic sub-model, and the fine-tuned first vocoder. After the electronic device receives the text to be processed, it uses the second speech synthesis model to perform speech synthesis on the text to be processed, thereby obtaining target audio that simulates the pronunciation of the target speaker. With this approach, since the first vocoder fine-tuning process does not use a noisy user recording as a label signal, but instead uses the noise-free audio of the target speaker as the label signal, the label signal and the acoustic features generated by the acoustic sub-model form matching acoustic feature-audio pairs. This avoids the introduction of noise signals during the fine-tuning of the first vocoder, thereby improving the quality of the second speech synthesis model. Therefore, the target audio obtained based on the second speech synthesis model is of high quality, achieving the goal of improving speech synthesis quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 This is a schematic diagram of the fine-tuning process of a traditional personalized speech synthesis model;
[0020] Figure 2A is the spectrum synthesized by the acoustic sub-model;
[0021] Figure 2B It is the spectrum extracted from the user's recording;
[0022] Figure 3 Schematic diagram of an implementation environment applicable to the speech synthesis method provided in an embodiment of the present application;
[0023] Figure 4 is a flowchart of the speech synthesis method provided in an embodiment of the present application;
[0024] Figure 5 Schematic diagram of the fine-tuning process of the speech synthesis model in the speech synthesis method provided by the present application;
[0025] Figure 6 A schematic diagram of a speech synthesis device provided in an embodiment of the present application;
[0026] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] Text-to-speech (TTS) is a hot research topic in the field of artificial intelligence. The purpose of this technology is to enable machines to speak fluently like humans, and make it difficult for listeners to distinguish whether the pronunciation is from a machine or a real human.
[0028] Common speech synthesis models consist of three modules: a text front-end, an acoustic sub-model, and a vocoder. The text front-end processes text into standardized intermediate feature vectors with phonetic characteristics. The acoustic sub-model takes the intermediate feature vectors as input and outputs acoustic feature vectors, such as fundamental frequency and mel-spectrograms. The vocoder converts the acoustic feature vectors into audio.
[0029] Personalized text-to-speech (TTS) is a key application of speech synthesis. To achieve personalized TTS, a speech synthesis model (hereinafter referred to as the general model or base model) is typically trained, consisting of the text frontend, acoustic sub-model, and vocoder modules described above. Next, a certain number of user recordings of the target speaker are recorded to extract speech features. The general model then learns the extracted speech features to produce a personalized speech synthesis model. This personalized speech synthesis model enables the machine to reproduce the target user's voice. Currently, personalized speech synthesis is widely used in scenarios such as audiobooks and virtual humans.
[0030] The personalized speech synthesis model is mainly achieved by fine-tuning the general model based on user recordings. The process of obtaining a personalized speech synthesis model is as follows: First, the acoustic sub-model and vocoder are trained based on a relatively large sample set to obtain an average acoustic sub-model and average vocoder. Among them, the sample set contains multiple audio-text pairs, and the audio is recorded by different speakers. Then, a certain amount of user recordings of the target speaker are used to fine-tune all or part of the parameters in the average acoustic sub-model and average vocoder. After that, the parameters of the fine-tuned average acoustic sub-model and average vocoder are stored to obtain a personalized speech synthesis model, which is subsequently used for personalized speech synthesis.
[0031] Figure 1 This is a diagram of the fine-tuning process of a traditional personalized speech synthesis model. Figure 1 ,During the fine-tuning process, we first use user recordings to extract acoustic ,features. The acoustic features and the corresponding user recording texts ,compose text-acoustic feature pairs, and the text-acoustic feature pairs are ,used to fine-tune the acoustic sub-model.
[0032] After fine-tuning the acoustic sub-model, the vocoder is fine-tuned. Vocoder fine-tuning requires two types of data: acoustic features generated by the acoustic sub-model based on user audio recordings, and the user audio recordings. These two types of data form acoustic feature-audio pairs for vocoder fine-tuning. The vocoder can be, for example, a non-autoregressive vocoder.
[0033] The vocoder consists of a generator and a discriminator. During vocoder fine-tuning, the generator processes the acoustic features generated by the acoustic sub-model to produce synthesized audio. This synthesized audio, along with the user recording, is then fed into the discriminator, which then uses it to make the synthesized audio more similar to the user recording. Both the generator and the discriminator are composed of neural networks.
[0034] As can be seen from the above, to obtain a personalized speech synthesis model, it is necessary to record the target speaker. This user recording is then used to fine-tune the acoustic sub-model and vocoder in the base model. Recording the target speaker is often done using portable devices such as mobile phones. However, real-world recording environments often lack the quality of professional recording studios. User recordings are of poor quality and uncontrollable, often containing certain noise characteristics, which hinders fine-tuning of the base model. This obstacle manifests itself in fine-tuning both the acoustic sub-model and the vocoder.
[0035] Fine-tuning the acoustic submodel essentially learns the mapping from text to acoustic features. Since it's impossible to learn features for every word in every sentence, noise features may or may not be learned. Learning noise features can cause the synthesized audio to be noisy. This problem can be alleviated through noise reduction.
[0036] The problem faced by vocoder fine-tuning is: if the acoustic sub-model learns the noise, then the acoustic features generated by the acoustic sub-model will correspond one-to-one with the user recording, that is, the acoustic features and the user recording will match. If the acoustic sub-model cannot learn the noise, then the acoustic features generated by the acoustic sub-model will not correspond one-to-one with the user recording, that is, the acoustic features and the user recording will not match. For example, please refer to Figure 2A and Figure 2B .
[0037] Figure 2A is the spectrum synthesized by the acoustic sub-model. Since the acoustic sub-model does not learn the noise characteristics, the spectrum is noise-free. Figure 2B The spectrum is extracted from the user recording. The original user recording is noisy. The spectrum is an acoustic feature. If mismatched acoustic features and audio are input to the vocoder for fine-tuning, the discriminator's role is to make the synthesized audio closer to the user recording. This will cause the vocoder to learn the incorrect mapping between acoustic features and audio during fine-tuning. This may generate noise during inference, resulting in poor audio quality generated by the personalized speech synthesis model.
[0038] As can be seen above, the reason the base model's vocoder introduces noise during fine-tuning is essentially because user recordings contain a lot of noise. Consequently, using user recordings as label signals causes the acoustic features generated by the acoustic submodel to mismatch the user recordings, leading to problems with vocoder fine-tuning.
[0039] Based on this, the embodiments of the present application provide a speech synthesis method, apparatus, device and readable storage medium, which fine-tune the first vocoder in the first speech synthesis model using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model, and use the second speech synthesis model to perform speech synthesis on the processed text to generate the target audio. Since the noise-free audio of the target speaker is used as the label signal to fine-tune the first vocoder when fine-tuning the second speech synthesis model, the introduction of noise signals during the fine-tuning process of the first vocoder is avoided, thereby making the quality of the second speech synthesis model higher, thereby achieving the purpose of improving the quality of the target audio.
[0040] Figure 3 This is a schematic diagram of the implementation environment applicable to the speech synthesis method provided in the embodiment of the present application. Figure 3 The implementation environment includes a terminal device 31 and a server 32, where a pre-adjusted personalized second speech synthesis model is deployed. The terminal device 31 and the server 32 communicate with each other via a wired network or a wireless network.
[0041] An application supporting personalized speech synthesis is installed on the terminal device 31; or a website supporting speech synthesis runs on the terminal device 31. The terminal device 31 includes, but is not limited to, smartphones, tablet computers, desktop computers, wearable devices (such as smart watches, augmented reality (AR) devices, virtual reality (VR) devices, etc.), netbooks, e-book readers, audio players, in-vehicle terminals, smart home devices, smart medical devices, smart transportation devices, robots, etc.
[0042] The terminal device 31 indicates the text to be processed to the server 32, which then performs speech synthesis on the text using the second speech synthesis model to generate target audio that simulates the pronunciation of the target speaker. The server 32 then sends the target audio to the terminal device 31, which then plays the target audio.
[0043] The first and second speech synthesis models are deployed on server 32, which provides speech synthesis functionality for terminal device 31. Server 32 can fine-tune the first vocoder in the first speech synthesis model using the noise-free audio of the target speaker as a label signal to obtain the second speech synthesis model. Server 32 can be a standalone physical server, a server cluster consisting of multiple servers, or a cloud server with cloud computing capabilities, etc., and this embodiment of the present application is not limiting.
[0044] It should be understood that Figure 3The number of terminal devices 31 and servers 32 in the embodiment is only for illustration. In actual implementation, any number of terminal devices 31 and servers 32 can be deployed according to actual needs.
[0045] Below, based on Figure 3 The implementation environment shown is used to describe the speech synthesis method provided in the embodiment of the present application in detail. Figure 4 . Figure 4 This is a flow chart of the speech synthesis method provided by the embodiment of the present application. This embodiment is described from the perspective of an electronic device, such as the above-mentioned Figure 3 The server in this embodiment includes:
[0046] 401. Train an acoustic sub-model and a first vocoder using a sample set containing multiple audio-text pairs, and generate a first speech synthesis model using the acoustic sub-model and the first vocoder.
[0047] Exemplarily, the electronic device pre-acquires a sample set containing multiple audio-text pairs, where each audio-text pair includes an audio segment and the corresponding text. The electronic device then uses the sample set containing the multiple audio-text pairs to train an acoustic sub-model and a first vocoder, and uses the acoustic sub-model and the first vocoder to generate a first speech synthesis model. This first speech synthesis model is also referred to as a base model, average model, or the like.
[0048] 402. Fine-tune the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model.
[0049] Exemplarily, the electronic device pre-adjusts a first vocoder in a first speech synthesis model using noise-free audio of a target speaker as a label signal, thereby obtaining a second speech synthesis model. The first speech synthesis model includes an acoustic sub-model and a first vocoder, and the second speech synthesis model includes the acoustic sub-model and the fine-tuned first vocoder.
[0050] 403. Perform speech synthesis on the text to be processed using the second speech synthesis model to obtain target audio that simulates the pronunciation of the target speaker.
[0051] Exemplarily, the electronic device obtains the text to be processed and processes the text to be processed using the second speech synthesis model, thereby obtaining the target audio that simulates the target speaker, that is, obtaining personalized audio. The electronic device can flexibly obtain the text to be processed. For example, the electronic device is a server, and the terminal device sends the text to be processed to the electronic device; for another example, the user enters the name of the text to be processed on the terminal device and sends it to the electronic device, and the electronic device performs a network search based on the name to obtain the text to be processed. For another example, the terminal device obtains the recording of the live broadcast staff in real time and converts it into text to be processed, and sends the text to be processed to the electronic device.
[0052] After obtaining the second speech synthesis model, the electronic device uses the second speech synthesis model to perform personalized speech synthesis on the text to be processed, thereby obtaining audio that simulates the voice of a specific person. The sample set includes multiple audio-text pairs, each of which contains a sample audio segment and the sample text corresponding to the sample audio segment. The sample audio can be a sample of the voice of any user and does not need to be limited to the audio of the voice of a specific person. Therefore, the text to be processed and the sample text can be the same or different. For example, for the same text A, user A reads text A aloud to obtain audio A, and user B reads text A aloud to obtain audio B. Audio A and text A constitute the audio-text pair in the sample set, and audio B and text A constitute the audio-text pair in the sample set. After obtaining the second speech synthesis model, the electronic device uses the second speech synthesis model to process text B to obtain audio that simulates the voice of user C. Text B and text A can be the same text or different texts.
[0053] In the embodiment of the present application, the second speech synthesis model includes a text front-end, an acoustic sub-model, and a fine-tuned first vocoder. For example, the first vocoder is obtained by fine-tuning the first vocoder in the general model using the noise-free audio of the target speaker as the label signal.
[0054] The universal model is also known as the base model, average model, etc. When training the universal model, an acoustic sub-model and vocoder are trained based on a relatively large sample set, resulting in an average acoustic sub-model and average vocoder. The universal model is then derived from the text front-end, the average acoustic sub-model, and the average vocoder. The average vocoder is also known as the first vocoder.
[0055] After the basic model is trained, the acoustic sub-model is fine-tuned, and the first vocoder is fine-tuned to obtain a personalized second speech synthesis model for the target speaker. In order to avoid the problem of noise introduced when fine-tuning the traditional vocoder, in an embodiment of the present application, when fine-tuning the first vocoder, the user recording, that is, the reference audio obtained by recording the target speaker is no longer used as the label signal, but the noise-free audio of the target speaker is used as the label signal, so that the first acoustic feature obtained by the acoustic sub-model based on the noise-free audio and the acoustic feature-audio pair composed of the noise-free audio are matched, which can avoid the problem of mismatch between acoustic features and audio when fine-tuning the traditional vocoder. Among them, noise-free audio is a relative concept, which refers to an audio with less noise content than the audio obtained by the user recording, or even negligible.
[0056] When the electronic device uses the second speech synthesis model to synthesize speech from the text being processed, the text front end processes the text into a standardized intermediate feature vector with phonetic features. The acoustic sub-model then uses the intermediate feature vector as input and outputs an acoustic feature vector, such as fundamental frequency and mel-spectrogram. Finally, the fine-tuned first vocoder processes the acoustic feature vector into the target audio.
[0057] In the speech synthesis method provided by an embodiment of the present application, an electronic device pre-trains an acoustic sub-model and a first vocoder using a sample set containing multiple audio-text pairs. The acoustic sub-model and the first vocoder are used to generate a first speech synthesis model. The first vocoder in the first speech synthesis model is then fine-tuned using the noise-free audio of a target speaker as a label signal to obtain a second speech synthesis model. The second speech synthesis model includes a text front-end, an acoustic sub-model, and the fine-tuned first vocoder. After the electronic device obtains the text to be processed, it uses the second speech synthesis model to perform speech synthesis on the text to be processed, thereby obtaining target audio that simulates the pronunciation of the target speaker. With this solution, since the first vocoder fine-tuning process does not use the user recording with high noise as the label signal, but instead uses the noise-free audio of the target speaker as the label signal, the label signal and the acoustic features generated by the acoustic sub-model form matching acoustic feature-audio pairs, avoiding the introduction of noise signals during the fine-tuning of the first vocoder, thereby improving the quality of the second speech synthesis model. Therefore, the target audio quality obtained based on the second speech synthesis model is high, achieving the purpose of improving speech synthesis quality.
[0058] Optionally, in the above embodiment, in the process of fine-tuning the first vocoder in the first speech synthesis model using the noise-free audio of the target speaker as a label signal to obtain the second speech synthesis model, first, a second vocoder is trained while training the basic model, the basic model includes a text front end, an acoustic sub-model and a first vocoder, and the second vocoder is a vocoder with better performance than the first vocoder, that is, the quality of the audio generated by the second vocoder is higher than the quality of the audio generated by the first vocoder. During the training process, the electronic device obtains a large-scale sample set, and the samples in the sample set are individual audio-text pairs. The electronic device uses the sample set to train the acoustic sub-model, the first vocoder and the second vocoder.
[0059] The electronic device then fine-tunes the acoustic sub-model and the first vocoder in the base model. The fine-tuning of the first vocoder relies on the second vocoder. This means that rather than using the target speaker's reference audio as the label signal, the electronic device uses the acoustic sub-model to generate a first acoustic feature. This feature is then fed into the second vocoder to generate a second audio signal, which is then used as the label signal.
[0060] Based on the above, it can be seen that: during the fine-tuning process of the first vocoder, the user recording with a lot of noise is not used as the label signal. Instead, the second audio generated by the second vocoder is used as the label signal. The first audio and the second audio are input into the discriminator of the first vocoder together, and the quality of the generator of the first vocoder is improved through fine-tuning. Because the acoustic features generated by the label signal and the acoustic sub-model form matching acoustic feature-audio pairs, the quality of the second speech synthesis model is higher. Among them, the user recording is the reference audio of the target speaker below, which can be obtained by recording the target speaker using a mobile phone or the like.
[0061] Using this solution, the electronic device pre-trains a basic model that includes at least an acoustic sub-model and a first vocoder, and at the same time trains a second vocoder. The second vocoder is then used to fine-tune the first vocoder to improve the quality of the second speech synthesis model.
[0062] Below, the training and fine-tuning of the basic model are explained in detail.
[0063] First, the basic model is trained.
[0064] A. Pre-trained acoustic sub-model.
[0065] The data required for training the acoustic sub-model is the text corresponding to the samples in the sample set and the second acoustic features of the audio samples. During acoustic sub-model training, for each audio-text pair in the sample set, the electronic device determines the second acoustic features of the audio sample in the audio-text pair and the factor sequence of the text sample corresponding to the audio sample. The electronic device then models the phoneme sequence and the corresponding second acoustic features to generate the acoustic sub-model.
[0066] Exemplarily, the electronic device converts the sample text into a phoneme sequence, models the relationship between the phoneme sequence and the second acoustic feature, and uses a mean square error function as the loss function.
[0067] Using this solution, the electronic device trains an acoustic sub-model by modeling the relationship between the second acoustic feature of the sample audio in the sample set and the phoneme sequence of the corresponding sample text, thereby achieving the purpose of training a high-quality acoustic sub-model.
[0068] Optionally, in the above embodiment, for each audio-text pair in the sample set, in the process of the electronic device determining the second acoustic feature of the sample audio in the audio-text pair, the electronic device first determines a first audio-text pair from the sample set, wherein the length of the sample text in the first audio-text pair is longer than the length of the sample text in the second audio-text pair, and the second audio-text pair is included in the sample set. Thereafter, the electronic device supplements the length of the sample text in the second audio-text pair based on the length of the sample text in the first audio-text pair.
[0069] Exemplary, the first audio-text pair is the audio-text pair with the longest text length in the sample set. That is to say, the length of the text in the first audio-text pair is longer than the length of the text in any one of the second audio-text pairs in the sample set. After determining the first audio-text pair, the electronic device fills the text in the second audio-text pair with zeros or other predetermined values so that the length of the text in the second audio-text pair is the same as the length of the text in the first audio-text pair. Simultaneously, the electronic device adds a section of silence to the audio in the second audio-text pair so that the length of the audio in the second audio-text pair text is the same as the length of the text in the first audio-text pair.
[0070] In addition, the audio-text pairs in the sample set can be divided into multiple batches (batches), and the text and audio acoustic features of each audio-text pair in a batch of audio-text pairs are input into the network together. For a batch, in order to adapt to the calculation of the processor, the electronic device performs zero padding on the second audio-text pair, that is, the audio-text pair with shorter text length and shorter audio, so that the audio length and text length of the audio-text pairs in the same batch are the same. In this way, what is input to the network is a square matrix, which is convenient for processing by the electronic device.
[0071] By adopting this solution, for the same batch, the electronic device increases the length of the text and audio in each second audio-text pair so that the length of the sample audio in the same batch is the same and the length of the sample text in the same batch is the same, which is convenient for the electronic device to process and achieves the purpose of improving the training speed of the acoustic sub-model.
[0072] B. Pre-train the first vocoder and the second vocoder.
[0073] In an embodiment of the present application, the second vocoder is a vocoder with better performance than the first vocoder. For example, the first vocoder is a non-autoregressive vocoder, and the second vocoder is an autoregressive vocoder. In the process of generating audio, the autoregressive vocoder adopts a point-by-point sampling method, and the generation of the sampling point at the current moment depends on the sampling point at the previous moment. This is a serial sampling method, which is often robust and has high generation quality, but is time-consuming and slow. The non-autoregressive vocoder adopts a parallel generation method of sampling points. There is no temporal dependency between the sampling points at different moments. While the inference speed is extremely fast, the generated audio quality is high.
[0074] In this embodiment of the present application, since non-autoregressive vocoders have extremely fast inference speed and generate high-quality audio, after the user enters the text to be processed, the first vocoder can obtain personalized target audio in a relatively short period of time. Therefore, the first vocoder used in the second speech synthesis model is a non-autoregressive vocoder, such as a GAN network-based vocoder.
[0075] During pre-training of the first vocoder, for each audio-text pair in the sample set, the electronic device extracts a first audio segment from the sample audio in the audio-text pair. Then, the electronic device extracts a third acoustic feature from the first audio segment, and trains the first vocoder based on the third acoustic feature and the first audio segment.
[0076] Exemplarily, during the training process of the first vocoder, the first audio segment input into the network each time is of equal length, and the third acoustic feature is also of fixed length. Therefore, for each sample audio, the electronic device extracts an audio segment from the sample audio, namely the first audio segment. The length of the first audio segment is, for example, 8192 sampling points, and the length of the third acoustic feature is, for example, 32 frames, etc., which are not limited in the embodiments of the present application. The first vocoder includes a generator, a discriminator, etc. The electronic device inputs each third acoustic feature into the generator to complete upsampling and generate audio. The first audio segment corresponding to the audio and the third acoustic feature is sent to the discriminator together, and the parameters are updated by generating an adversarial loss function, a mean square error loss function of the acoustic feature, etc., so as to train the first vocoder.
[0077] By adopting this solution, a first audio segment is captured from the sample audio, a third acoustic feature is extracted from the first audio segment, and a first vocoder is trained based on the third acoustic feature and the first audio segment, thereby achieving the purpose of quickly training a high-quality first vocoder.
[0078] The training process for the second vocoder is similar to that for the first vocoder. During training, for each audio-text pair in the sample set, the electronic device extracts a second audio segment from the sample audio in the audio-text pair and extracts a fourth acoustic feature from the second audio segment. The electronic device then trains the second vocoder based on the fourth acoustic feature and the second audio segment.
[0079] The difference between the training process of the second vocoder and the training process of the first vocoder is that the length of the second audio segment is less than the length of the first audio segment. For example, the length of the first audio segment is 8192 sampling points and the length of the third acoustic feature is 32 frames; the length of the first audio segment is 1280 sampling points and the length of the third acoustic feature is 5 frames. The length of the audio segment and the length of the acoustic feature are determined according to the frame shift. For example, if the frame shift is 256 and the length of the third acoustic feature is 32 frames, the length of the first audio segment is 256×32. The value of the frame shift can be set according to actual needs and is not limited in the embodiments of the present application.
[0080] In addition, during the training process of the second vocoder, a loss function used is, for example, a cross entropy loss function.
[0081] Using this solution, the electronic device intercepts a second audio segment from the sample audio, extracts a fourth acoustic feature from the second audio segment, and trains a second vocoder based on the fourth acoustic feature and the second audio segment, thereby achieving the purpose of quickly training a high-quality second vocoder.
[0082] In the above embodiment, the second acoustic feature, the third acoustic feature and the fourth acoustic feature can be extracted by digital signal processing, such as constructing an acoustic feature extraction network using Fourier transform and filters, and extracting acoustic features based on the network.
[0083] Second, fine-tuning of the base model.
[0084] C. Fine-tuning of the acoustic sub-model.
[0085] After completing the pre-training of the acoustic sub-model, the first vocoder, and the second vocoder to obtain the basic model, the basic model can be fine-tuned to obtain a second speech synthesis model for synthesizing the pronunciation of the model target speaker.
[0086] During the fine-tuning of the acoustic sub-model, you can choose to freeze the weights of some network layers. That is, during the fine-tuning process, the weights of some network layers remain unchanged. In this embodiment of the application, the weights of some network layers can be frozen, or none of them can be frozen.
[0087] Figure 5 This is a schematic diagram of the fine-tuning process of the first speech synthesis model in the speech synthesis method provided by this application. Figure 5 The electronic device uses digital signal processing to extract acoustic features from the reference audio, combines the acoustic features with the corresponding text to form a text-acoustic feature pair, and uses the text-acoustic feature pair to adjust the acoustic sub-model.
[0088] Please refer to Figure 5 The second vocoder is used to synthesize high-quality second audio. It can be an autoregressive vocoder such as a wave recurrent neural network (RNN) or WaveGlow, or other autoregressive vocoders, and the embodiments of the present application are not limited thereto. The acoustic submodel is a non-autoregressive acoustic submodel, such as FastSpeech, FastSpeech2, or other non-autoregressive acoustic submodels. The first vocoder is, for example, a GAN-type vocoder, a MelGAN vocoder, a HiFiGAN vocoder, or the like.
[0089] Please refer to Figure 5 In order to overcome the problem of mismatch between the first acoustic features generated by the acoustic sub-model and the reference audio, a second vocoder is trained while training the base model. The second vocoder is, for example, an autoregressive neural network vocoder that can synthesize high-quality audio. When fine-tuning the first vocoder, the first acoustic features generated by the acoustic sub-model are input into the second vocoder to generate the second audio. Due to the high robustness of the second vocoder, the quality of the second audio is higher than that of the first audio generated by the unadjusted first vocoder. Therefore, the second audio generated by the second vocoder is used instead of the reference audio as the label signal and is input into the discriminator of the first vocoder together with the first audio to improve the generation quality of the generator through fine-tuning.
[0090] D. Fine-tuning of the first vocoder.
[0091] After fine-tuning the acoustic sub-model, the first vocoder needs to be fine-tuned using a reference audio of the target speaker. This reference audio can be recorded using a portable device such as a mobile phone.
[0092] During the fine-tuning process, the electronic device inputs reference audio of the target speaker into the acoustic sub-model to obtain a first acoustic feature, inputs the first acoustic feature into the first vocoder to obtain a first audio, and inputs the first acoustic feature into the second vocoder to obtain a second audio. The electronic device then fine-tunes the first vocoder based on the first and second audio, and generates the second speech synthesis model based on the adjusted first vocoder and the acoustic sub-model.
[0093] In the embodiment of the present application, the second vocoder has high robustness, and the second audio generated by the second vocoder is of higher quality than the first audio generated by the untuned first vocoder. Therefore, the second audio generated by the second vocoder can be used as the label signal of the first vocoder. In other words, in the process of fine-tuning the first vocoder, the first acoustic features generated by the acoustic sub-model and the second audio generated by the second vocoder are combined into an acoustic feature-audio pair for fine-tuning the first vocoder, that is, the second audio is used as the label signal. In this way, there will be no problem of mismatch between acoustic features and audio.
[0094] Based on the above, it can be seen that: in the embodiment of the present application, in order to avoid introducing noise during the fine-tuning stage of the first vocoder, when fine-tuning the first vocoder, an acoustic sub-model is used to generate a first acoustic feature, and a highly robust and high-quality second vocoder is used to generate a second audio with the first acoustic feature as input, thereby replacing the user recording as a reference signal, solving the problem of mismatch between the first acoustic feature generated by the acoustic sub-model and the user recording, so that the fine-tuning of the first vocoder no longer depends on the actual recording of the target speaker, reducing the dependence on the actual recording.
[0095] By adopting this solution, the electronic device uses the first acoustic feature and the second audio to form an acoustic feature-audio pair as a fine-tuning data set for fine-tuning the first vocoder, and fine-tunes the first vocoder required for the personalized second speech synthesis model. The fine-tuning effect is good, and the purpose of effectively improving the sound quality of the target audio generated by the second speech synthesis model is achieved.
[0096] The reason why noise is introduced in the traditional vocoder fine-tuning process is that in personalized TTS, the user recording, that is, the above-mentioned reference audio, contains a lot of noise. Using the reference audio as the label signal will cause the first acoustic feature generated by the acoustic sub-model to not match the reference audio. In order to overcome this problem, in the above embodiment, the fine-tuning of the first vocoder is described by taking the second audio generated by the second vocoder as the label signal as an example. However, the embodiment of the present application is not limited. In other feasible implementation methods, a classifier can also be trained, and the noise classifier can be used to determine the noise-free audio segment from the reference audio to obtain the noise-free audio. Afterwards, the second speech synthesis model is trained based on the noise-free audio.
[0097] For example, during fine-tuning of the first vocoder, only a third audio segment of limited length is used as a label signal for fine-tuning the first vocoder. This allows for training a noise classifier. During fine-tuning, noisy segments of the reference audio are discarded, and only noise-free segments are used as label signals. This label signal and the first acoustic feature generated by the acoustic sub-model form an acoustic feature-audio pair for fine-tuning the first vocoder, resolving the mismatch between acoustic features and audio.
[0098] This solution uses a noise classifier to select noise-free audio segments from the reference audio as label signals to fine-tune the first vocoder. This avoids the introduction of noise during the fine-tuning stage of the first vocoder, solves the problem of mismatch between acoustic features and audio, and achieves the purpose of accurately fine-tuning the first vocoder.
[0099] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0100] Figure 6 Schematic diagram of a speech synthesis device provided in an embodiment of the present application. The speech synthesis device 600 includes: a training module 61, a processing module 62 and a synthesis module 63.
[0101] A training module 61 is configured to train an acoustic sub-model and a first vocoder using a sample set comprising a plurality of audio-text pairs, and generate a first speech synthesis model using the acoustic sub-model and the first vocoder;
[0102] A processing module 62 is configured to fine-tune the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model;
[0103] The synthesis module 63 is configured to perform speech synthesis on the text to be processed using the second speech synthesis model to obtain target audio that simulates the pronunciation of a target speaker.
[0104] In one feasible implementation, the processing module 62 is configured to train a second vocoder using the sample set, wherein the quality of the audio generated by the second vocoder is higher than the quality of the audio generated by the first vocoder; generate the noise-free audio using the second vocoder to fine-tune the first vocoder; and generate the second speech synthesis model based on the fine-tuned first vocoder and the acoustic sub-model.
[0105] In a feasible implementation, when the processing module 62 uses the second vocoder to generate the noise-free audio to fine-tune the first vocoder, it is used to input the reference audio of the target speaker into the acoustic sub-model to obtain a first acoustic feature; input the first acoustic feature into the first vocoder to obtain a first audio, input the first acoustic feature into the second vocoder to obtain a second audio, and use the second audio as the noise-free audio; and fine-tune the first vocoder according to the first audio and the second audio.
[0106] In a feasible implementation, when the training module 61 uses a sample set containing multiple audio-text pairs to train the acoustic sub-model, it is used to determine the second acoustic feature of the sample audio in the audio-text pair for each audio-text pair in the sample set; determine the phoneme sequence of the sample text in the audio-text pair, and model the phoneme sequence and the corresponding second acoustic feature to obtain the acoustic sub-model.
[0107] In a feasible implementation, the training module 61 is further used to determine a first audio-text pair from the sample set before determining the second acoustic feature of the sample audio in the audio-text pair for each audio-text pair in the sample set, wherein the length of the sample text in the first audio-text pair is longer than the length of the sample text in the second audio-text pair, and the second audio-text pair is included in the sample set; and supplement the length of the sample text in the second audio text according to the length of the sample text in the first audio-text pair.
[0108] In one feasible implementation, when the training module 61 uses a sample set containing multiple audio-text pairs to train the first vocoder, it is used to, for each audio-text pair in the sample set, intercept a first audio segment of the sample audio in the audio-text pair; determine a third acoustic feature of the first audio segment; and train the first vocoder based on the third acoustic feature and the first audio segment.
[0109] In one feasible implementation, when the training module 61 uses the sample set to train the second vocoder, it is used to, for each audio-text pair in the sample set, intercept a second audio segment from the sample audio in the audio-text pair, where the length of the second audio segment is less than the length of the first audio segment; determine a fourth acoustic feature of the second audio segment; and train the second vocoder based on the fourth acoustic feature and the second audio segment.
[0110] In a feasible implementation, the training module 61 is also used to train a noise classifier before the processing module 62 fine-tunes the first vocoder using the noise-free audio of the target speaker as a label signal to obtain the second speech synthesis model; and the noise classifier is used to determine a noise-free audio segment from the reference audio to obtain the noise-free audio.
[0111] The speech synthesis device provided in the embodiment of the present application can perform the actions of the electronic device in the above embodiment. Its implementation principle and technical effects are similar and will not be repeated here.
[0112] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 7 The electronic device 700 described in the embodiment of the present application includes: at least one processor 71, at least one communication bus 72, a user interface 73, at least one network interface 74 and a memory 75.
[0113] The communication bus 72 is used to realize the connection and communication between these components.
[0114] The user interface 73 may include a display screen and a camera. Optionally, the user interface 73 may also include a standard wired interface and a wireless interface. The display screen is used to display the second image that renders the algorithm result.
[0115] The network interface 74 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0116] The processor 71 may include one or more processing cores. The processor 71 utilizes various interfaces and circuits to connect the various components within the entire electronic device 70. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 75, and by accessing data stored in the memory 75, the processor 71 performs various functions of the electronic device 70 and processes data. Optionally, the processor 71 may be implemented in the form of at least one hardware component selected from the group consisting of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 71 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 71 and may be implemented separately on a separate chip.
[0117] Among them, the memory 75 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 75 includes a non-transitory computer-readable storage medium. The memory 75 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 75 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 75 may also be optionally at least one storage device located away from the aforementioned processor 71. As Figure 7 As shown, the memory 75 as a computer storage medium may include an operating system, a network communication module, a user interface module, and operating applications of the electronic device.
[0118] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0119] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0120] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0122] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0123] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0124] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0125] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0126] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A speech synthesis method, characterized in that: Applied to electronic equipment, the method includes: Training an acoustic sub-model and a first vocoder using a sample set comprising a plurality of audio-text pairs, and generating a first speech synthesis model using the acoustic sub-model and the first vocoder; Fine-tuning the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model; The second speech synthesis model is used to perform speech synthesis on the text to be processed to obtain target audio that simulates the pronunciation of the target speaker.
2. The method according to claim 1, characterized in that The method of fine-tuning the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model includes: training a second vocoder using the sample set, wherein the quality of audio generated by the second vocoder is higher than the quality of audio generated by the first vocoder; generating the noise-free audio using the second vocoder to fine-tune the first vocoder; The second speech synthesis model is generated according to the fine-tuned first vocoder and the acoustic sub-model.
3. The method according to claim 2, characterized in that The generating the noise-free audio using the second vocoder to fine-tune the first vocoder comprises: Inputting the reference audio of the target speaker into the acoustic sub-model to obtain a first acoustic feature; Inputting the first acoustic feature to the first vocoder to obtain a first audio, inputting the first acoustic feature to the second vocoder to obtain a second audio, and using the second audio as the noise-free audio; The first vocoder is fine-tuned based on the first audio and the second audio.
4. The method according to any one of claims 1 to 3, characterized in that The acoustic sub-model is trained using a sample set containing multiple audio-text pairs, including: For each audio-text pair in the sample set, determining a second acoustic feature of the sample audio in the audio-text pair; Determining a phoneme sequence of a sample text in the audio-text pair; The phoneme sequence and the corresponding second acoustic feature are modeled to obtain the acoustic sub-model.
5. The method according to claim 4, characterized in that Before determining the second acoustic feature of the sample audio in each audio-text pair in the sample set, the method further includes: Determining a first audio-text pair from the sample set, wherein the length of the sample text in the first audio-text pair is longer than the length of the sample text in the second audio-text pair, and the second audio-text pair is included in the sample set; The length of the sample text in the second audio text is supplemented according to the length of the sample text in the first audio text pair.
6. The method according to claim 2, characterized in that The method of training the first vocoder using a sample set comprising a plurality of audio-text pairs comprises: For each audio-text pair in the sample set, extract a first audio segment from the sample audio in the audio-text pair; determining a third acoustic characteristic of the first audio segment; The first vocoder is trained based on the third acoustic feature and the first audio segment.
7. The method according to claim 6, characterized in that The step of training the second vocoder by using the sample set includes: For each audio-text pair in the sample set, extract a second audio segment from the sample audio in the audio-text pair, where the length of the second audio segment is shorter than the length of the first audio segment; determining a fourth acoustic characteristic of the second audio segment; The second vocoder is trained based on the fourth acoustic feature and the second audio segment.
8. The method according to claim 1, characterized in that Before fine-tuning the first vocoder using the noise-free audio of the target speaker as a label signal to obtain the second speech synthesis model, the method further includes: Train a noise classifier; The noise classifier is used to determine a noise-free audio segment from the reference audio to obtain the noise-free audio.
9. A speech synthesis device, characterized in that: include: A training module, configured to train an acoustic sub-model and a first vocoder using a sample set comprising a plurality of audio-text pairs, and generate a first speech synthesis model using the acoustic sub-model and the first vocoder; a processing module, configured to fine-tune the first vocoder using the noise-free audio of the target speaker as a label signal to obtain a second speech synthesis model; The synthesis module is used to perform speech synthesis on the text to be processed using the second speech synthesis model to obtain target audio that simulates the pronunciation of the target speaker.
10. An electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Speech model training and synthesizing method with few corpora
CN112992118A
Corpus generation method and device, computer readable storage medium and terminal equipment
CN114637820A
Model training and speech synthesis method and device, equipment and medium
CN115700871A
Voice generation method and device, storage medium and electronic equipment
CN117095670A
Speech synthesis processing method and device, equipment and medium
CN117456979A