Speech synthesis method and related devices and equipment
By obtaining phonemes, pitch and phoneme duration, and combining Mel's spectrum and deep learning network for encoding and decoding, the problem of low speech synthesis effect is solved, and high-precision and diversified speech synthesis is achieved.
Patent Information
- Application Number
- CN202111280665.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the prior art, the synthesis effect is poor, especially when using the recurrent neural network model for speech synthesis, the synthesis effect is not high.
By obtaining the phoneme, pitch and phoneme duration of the target object, the Mel spectrum of the object to be synthesized is determined, and the tone feature matrix is extracted based on the Mel spectrum, the deep learning network is used for encoding and decoding, and the speech synthesis is achieved by combining the attention network and the decoder.
It improves the accuracy and accuracy of speech synthesis, can restore the tone of the object to be synthesized, and match the phonemes, pitch and phoneme duration of the synthesized speech with the target object, achieving diversified speech synthesis.
Smart Images

Figure CN114220414B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech synthesis, and particularly to a speech synthesis method and related devices and equipment. Background Art
[0002] With the continuous development of electronic information processing technology, speech, as an important carrier for people to obtain information, has been widely used in daily life and work. In application scenarios involving speech, there is usually a process of speech synthesis. Speech synthesis refers to synthesizing specified text or audio into target audio that meets requirements.
[0003] Currently, a recurrent neural network model is often used for speech synthesis, but the method of using a recurrent neural network model for speech synthesis has the problem of low synthesis effect.
[0004] Based on this, how to improve the effect of speech synthesis has become a technical problem that needs to be solved currently. Summary of the Invention
[0005] This application provides a speech synthesis method and related devices and equipment, which solve the problem of poor speech synthesis effect in the prior art.
[0006] This application provides a speech synthesis method, which includes: obtaining the phonemes, pitch, and phoneme duration of a target object; and obtaining a to-be-synthesized object, determining the Mel spectrogram of the to-be-synthesized object, and extracting the timbre feature matrix of the to-be-synthesized object based on the Mel spectrogram; encoding the phonemes, pitch, and phoneme duration of the target object through a speech synthesis model to obtain encoded data; and decoding the timbre feature matrix and the encoded data through the speech synthesis model to obtain the synthesized speech of the to-be-synthesized object.
[0007] Among them, the step of obtaining the to-be-synthesized object and determining the Mel spectrogram of the to-be-synthesized object includes: obtaining the to-be-synthesized object, performing frame addition and windowing and Fourier transform on the to-be-synthesized object to obtain the linear spectrogram of the to-be-synthesized object; and inputting the linear spectrogram into a Mel filter bank for filtering processing to obtain the Mel spectrogram.
[0008] Among them, the step of extracting the timbre feature matrix of the to-be-synthesized object based on the Mel spectrogram includes: inputting the Mel spectrogram into a deep learning network, and using the deep learning network to extract the timbre feature matrix of the to-be-synthesized object.
[0009] Among them, the speech synthesis model includes an encoder, an attention network, and a decoder that are cascaded with each other; the steps of encoding the phonemes, pitch, and phoneme duration of the target object through the speech synthesis model to obtain encoded data include: encoding the phonemes, pitch, and phoneme duration of the target object through the encoder in the speech synthesis model to obtain encoded data; the steps of decoding the timbre feature matrix and the encoded data through the speech synthesis model to obtain the synthesized speech of the object to be synthesized include: sequentially decoding the timbre feature matrix and the encoded data through the attention network and the decoder in the speech synthesis model to obtain the synthesized speech.
[0010] Among them, the attention network includes a position-based attention mechanism.
[0011] Among them, the speech synthesis model further includes a sentence-breaking model; the steps of decoding the timbre feature matrix and the encoded data through the decoder in the speech synthesis model to obtain the synthesized speech further include: performing sentence-breaking on the decoded data through the sentence-breaking model in the speech synthesis model to obtain the synthesized speech.
[0012] Among them, before the step of obtaining the phonemes, pitch, and phoneme duration of the target object, it includes: obtaining a sample audio, and extracting the phonemes, pitch, and phoneme duration of the sample audio, where the phonemes, pitch, and phoneme duration of the sample audio respectively cover a preset phoneme range, a preset pitch range, and a preset phoneme duration range; inputting the phonemes, pitch, and phoneme duration of the sample audio into the encoder in the initial model for encoding to obtain sample encoded data; determining the sample Mel spectrogram of the sample audio, and extracting the timbre feature matrix of the sample audio based on the sample Mel spectrogram; sequentially inputting the timbre feature matrix and the sample encoded data of the sample audio into the attention network and the decoder in the initial model for decoding to obtain the sample synthesized speech; in response to the similarity between the sample synthesized speech and the sample audio meeting the preset requirements, the speech synthesis model is obtained.
[0013] Among them, the steps of obtaining the phonemes, pitch, and phoneme duration of the target object include: obtaining a target audio, and extracting the corresponding phonemes, pitch, and phoneme duration from the target audio; or obtaining a target text, and determining the corresponding phonemes, pitch, and phoneme duration based on the target text.
[0014] The present application also provides a voice synthesis device, including: an acquisition module, configured to acquire the phonemes, pitch, and phoneme duration of a target object; an extraction module, configured to acquire a to-be-synthesized object, determine the Mel spectrum of the to-be-synthesized object, and extract the timbre feature matrix of the to-be-synthesized object based on the Mel spectrum; an encoding module, configured to encode the phonemes, pitch, and phoneme duration of the target object through a voice synthesis model to obtain encoded data; and a decoding module, configured to decode the timbre feature matrix and the encoded data through the voice synthesis model to obtain the synthesized voice of the to-be-synthesized object.
[0015] The present application also provides an electronic device, including a memory and a processor coupled to each other, where the processor is configured to execute program instructions stored in the memory to implement any one of the above voice synthesis methods.
[0016] The present application also provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, any one of the above voice synthesis methods is implemented.
[0017] In the above solution, the present application acquires the phonemes, pitch, and phoneme duration of a target object; acquires a to-be-synthesized object, determines the Mel spectrum of the to-be-synthesized object, and extracts the timbre feature matrix of the to-be-synthesized object based on the Mel spectrum; encodes the phonemes, pitch, and phoneme duration of the target object to obtain encoded data; decodes the timbre feature matrix and the encoded data to obtain the synthesized voice of the to-be-synthesized object. The present application encodes and decodes based on the phonemes, pitch, and phoneme duration of the target object, and can refine the voice features on the basis of realizing voice synthesis, improve the accuracy and precision of voice synthesis, thereby improving the voice quality of the synthesized voice. And by encoding the timbre feature matrix, voice synthesis can be fully performed based on the timbre of the to-be-synthesized object, so that the finally generated synthesized voice can restore the timbre of the to-be-synthesized object, thereby improving the effect of voice synthesis. Moreover, the present application encodes the timbre feature matrix and the encoded data at the same time, which can make the timbre of the synthesized voice be restored based on the timbre of the to-be-synthesized object, and can also make the phonemes, pitch, and phoneme duration of the synthesized voice match the phonemes, pitch, and phoneme duration of the target object. Furthermore, by adjusting the phonemes, pitch, and phoneme duration of the target object, a variety of synthesized voices can be synthesized based on a small number of to-be-synthesized objects. Description of the Drawings
[0018] Figure 1 is a schematic flowchart of an embodiment of the voice synthesis method of the present application;
[0019] Figure 2 is a schematic flowchart of another embodiment of the voice synthesis method of the present application;
[0020] Figure 3 isFigure 2 Schematic diagram of the framework of a corresponding implementation of the distributed system in the embodiment;
[0021] Figure 4 Schematic diagram of the framework of an embodiment of the voice synthesis device of the present application;
[0022] Figure 5 Schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0023] Figure 6 Schematic diagram of the framework of an embodiment of the computer-readable storage medium of the present application. Specific implementation manners
[0024] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0025] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0026] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein is merely a description of the association relationship of associated objects, and there can be three relationships. For example, A and / or B can be: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" herein means two or more than two.
[0027] Please refer to Figure 1 , Figure 1 Schematic diagram of the process of an embodiment of the voice synthesis method of the present application.
[0028] Step S11: Obtain the phonemes, pitch, and phoneme duration of the target object.
[0029] First, obtain the phonemes, pitch, and phoneme duration of the target object. Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Pitch refers to various sounds with different pitch levels, that is, the height of the sound, which is one of the basic characteristics of the sound. The phoneme duration is the pronunciation duration corresponding to the phoneme. Among them, the target object may include target audio such as songs, operas, broadcasts, or recitations, and the specific type of the target object is not limited herein. The phonemes, pitch, and phoneme duration can reflect the speech characteristics of the target object as much as possible.
[0030] In a specific application scenario, when obtaining the phonemes, pitch, and phoneme duration of a target object, it is possible to input the target object into a pre-trained deep learning model for prediction to obtain the phonemes, pitch, and phoneme duration of the target object. In another specific application scenario, when obtaining the phonemes, pitch, and phoneme duration of a target object, it is also possible to analyze the phonemes of the target object based on the corresponding dictionary of the target object, analyze the pitch of the target object based on the international standard pitch, and predict the target object through a pre-trained deep learning model to obtain the phoneme duration of the target object.
[0031] Among them, the voice synthesis in this embodiment needs to improve the effect of voice synthesis. Therefore, the phonemes, pitch, and phoneme duration of the target object obtained in this step are all indispensable to refine the voice characteristics of the target object to the greatest extent. Thus, when performing voice synthesis subsequently, based on various characteristics of the phonemes, pitch, and phoneme duration of the target object, the feature accuracy of the synthesized voice is improved, thereby improving the accuracy of the synthesized voice.
[0032] Step S12: Obtain the object to be synthesized, determine the Mel spectrum of the object to be synthesized, and extract the timbre feature matrix of the object to be synthesized based on the Mel spectrum.
[0033] The object to be synthesized in this embodiment may include the audio to be synthesized or other voice data containing pronunciation data. Among them, the voice synthesis in this embodiment can be applied to scenarios such as audio production in music software and audio pitch correction in KTV software. For example: the object to be synthesized can be the voice audio input by the user, can also be the music score, opera score or any other text including pronunciation data made by the user. The audio to be synthesized can also include audio data emitted manually or audio data output by a deep neural network. Among them, the specific type of the object to be synthesized can be set based on the actual situation and is not limited here.
[0034] After obtaining the object to be synthesized, determine the Mel spectrum of the object to be synthesized based on the object to be synthesized, and extract the timbre feature matrix of the object to be synthesized based on the Mel spectrum. Among them, the Mel spectrum, also known as the mel spectrogram, can reflect audio features.
[0035] In a specific application scenario, it is possible to first obtain the linear spectrum of the object to be synthesized, and then perform weighted summation of the Mel scale on the linear spectrum to obtain the Mel spectrum. In another specific application scenario, it is also possible to first obtain the linear spectrum of the object to be synthesized, and then input the linear spectrum into a Mel filter bank for filtering processing to obtain the Mel spectrum. The specific method for obtaining the Mel spectrum is not limited here.
[0036] The features in the timbre feature matrix may include one or more of features such as zero-crossing rate, spectral centroid, spectral attenuation, chroma frequency, etc., and are not limited specifically here.
[0037] In a specific application scenario, the timbre feature matrix of the object to be synthesized can be extracted based on the Mel spectrum through the three-level center clipping cross-correlation function method. In a specific application scenario, the timbre feature matrix of the object to be synthesized can be extracted based on the Mel spectrum through the Newton-Gauss type nonlinear estimation method. In another specific application scenario, the timbre feature matrix of the object to be synthesized can also be extracted based on the Mel spectrum through a deep learning network. Among them, the specific method for extracting the timbre feature matrix is not limited herein.
[0038] Among them, the execution order of this step and step S11 is not sequential. Step S11 can be executed first and then step S12, or step S12 can be executed first and then step S11, or step S12 and step S11 can be executed simultaneously. Specifically, step S11 only needs to be executed before step S13, and step S12 only needs to be executed before step S14.
[0039] Step S13: Encode the phonemes, pitch, and phoneme duration of the target object through a speech synthesis model to obtain encoded data.
[0040] After obtaining the phonemes, pitch, and phoneme duration of the target object, encode the phonemes, pitch, and phoneme duration of the target object through a speech synthesis model to obtain encoded data.
[0041] Among them, the speech synthesis model in this step is a pre-trained speech synthesis model.
[0042] In this step of other embodiments, the phonemes, pitch, and phoneme duration of the target object can also be encoded through an independent encoder to obtain encoded data. Among them, the encoder can include one or more of encoders such as AutoEncoder, Sparse AutoEncoder, Denoising AutoEncoders, waveform encoder, parameter encoder, and hybrid encoder. In another specific application scenario, the phonemes, pitch, and phoneme duration of the target object can also be encoded through encoding methods such as hard coding, one-hot coding, or target variable coding to obtain encoded data.
[0043] Step S14: Decode the timbre feature matrix and the encoded data through a speech synthesis model to obtain the synthesized speech of the object to be synthesized.
[0044] Taking the timbre feature matrix and the encoded data together as the object of decoding, input them into the speech synthesis model to decode the timbre feature matrix and the encoded data to obtain the synthesized speech of the object to be synthesized.
[0045] In a specific application scenario, before decoding the timbre feature matrix, each running step of the speech synthesis model is frozen respectively, and then the unfrozen running steps are fine-tuned with mini-batch data to determine the influence of each running step in the speech synthesis model on the timbre. Specifically, 1. Freeze the encoding step of the speech synthesis model, and then perform timbre training on the decoding step to determine the importance of the decoding step in the timbre synthesis of the entire model; 2. Freeze the decoding step, and then perform timbre training on the encoding step to determine the importance of the encoding step in the timbre synthesis of the entire model. Finally, it is obtained that the decoding step has a greater influence on speech synthesis and timbre restoration. Therefore, in this embodiment, only the timbre feature matrix is decoded, without first encoding and then decoding the timbre feature matrix or performing other additional operations, so that on the basis of fully considering the specific influence of the model on the timbre, the ineffective operations of the model on the timbre feature matrix can be reduced, and on the basis of ensuring speech synthesis based on the timbre feature matrix, the synthesis efficiency of the speech synthesis model can be improved.
[0046] After obtaining the encoded data, the timbre feature matrix and the encoded data are jointly decoded to obtain the synthesized speech of the object to be synthesized.
[0047] In a specific application scenario, the output of the encoder can be used as the input of the decoder, and the input of the encoder can be used as the label of the decoder to train the corresponding decoder, so that the decoder decodes the timbre feature matrix and the encoded data to obtain the synthesized speech of the object to be synthesized.
[0048] In this step of other embodiments, one or more decoding methods such as the greedy decoding algorithm, the beam search decoding algorithm, and the sampling-based decoding algorithm can also be used to decode the timbre feature matrix and the encoded data to obtain the synthesized speech of the object to be synthesized. Among them, the decoding step of this step can correspond to the encoding in step S13 for the convenience of speech synthesis.
[0049] Among them, since at this time not only the encoded data is decoded but also the timbre feature matrix is decoded, it can make the timbre of the synthesized speech be restored based on the timbre of the object to be synthesized, and at the same time make the phonemes, pitch, and phoneme duration of the synthesized speech match those of the target object. As a result, the finally generated synthesized speech can restore the timbre of the object to be synthesized while achieving the phonemes, pitch, and phoneme duration of the target object, thereby improving the effect of speech synthesis. Therefore, in the speech synthesis method of this embodiment, only by adjusting the phonemes, pitch, and phoneme duration of the target object can the modification of the phonemes, pitch, and phoneme duration of the synthesized speech of the object to be synthesized be achieved. Furthermore, the speech synthesis method of this embodiment can synthesize a variety of synthesized speeches based on a small number of objects to be synthesized. That is, by encoding the phonemes, pitch, and phoneme duration of the target object, the synthesized speech can be synthesized based on the phonemes, pitch, and phoneme duration of the target object, and a synthesized speech with the phonemes, pitch, and phoneme duration of the target object can be obtained. Then, by decoding the timbre feature matrix and the encoded data, the synthesized speech of the object to be synthesized can be obtained, which can further make the synthesized speech with the phonemes, pitch, and phoneme duration of the target object have the timbre of the object to be synthesized.
[0050] Through the above steps, the speech synthesis method of this embodiment obtains the phonemes, pitch, and phoneme duration of the target object; and obtains the object to be synthesized, determines the Mel spectrum of the object to be synthesized, and extracts the timbre feature matrix of the object to be synthesized based on the Mel spectrum; encodes the phonemes, pitch, and phoneme duration of the target object to obtain encoded data; decodes the timbre feature matrix and the encoded data to obtain the synthesized speech of the object to be synthesized. This embodiment encodes and decodes based on the phonemes, pitch, and phoneme duration of the target object, which can refine the speech features while realizing speech synthesis, improve the accuracy and precision of speech synthesis, thereby improving the speech quality of the synthesized speech. Moreover, by encoding the timbre feature matrix, speech synthesis can be fully based on the timbre of the object to be synthesized, so that the finally generated synthesized speech can restore the timbre of the object to be synthesized, thereby improving the effect of speech synthesis. And this embodiment also encodes the timbre feature matrix and the encoded data at the same time, which can make the timbre of the synthesized speech not only be restored based on the timbre of the object to be synthesized, but also make the phonemes, pitch, and phoneme duration of the synthesized speech match the phonemes, pitch, and phoneme duration of the target object. Furthermore, by adjusting the phonemes, pitch, and phoneme duration of the target object, a variety of synthesized speeches can be synthesized based on a small number of objects to be synthesized. And this embodiment only decodes the timbre feature matrix, without first encoding and then decoding the timbre feature matrix or other additional operations, so that on the basis of fully considering the specific influence of the model on the timbre, the invalid operations of the model on the timbre feature matrix can be reduced, and on the basis of ensuring speech synthesis based on the timbre feature matrix, the synthesis efficiency of the speech synthesis model can be improved.
[0051] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the speech synthesis method of this application.
[0052] Step S21: Obtain the target audio, and extract the corresponding phonemes, pitch, and phoneme duration from the target audio.
[0053] Obtain the target audio, extract the phonemes and pitch of the audio to be synthesized from the target audio based on a preset standard, and input the target audio into a duration prediction model for duration prediction to obtain the phoneme duration of the target audio.
[0054] Among them, the preset standard can be selected based on the language type of the target audio. In a specific application scenario, when the target object is a Chinese audio, the phonemes and pitch of the audio to be synthesized can be extracted from the target object based on the Chinese dictionary standard and the international phonetic standard, and then the target object is input into the duration prediction model for duration prediction to obtain the phoneme duration of the target object. Among them, the duration prediction model in this application scenario is a duration prediction model that has been pre-trained based on Chinese audio.
[0055] In a specific application scenario, when the target object is an English audio, the phonemes and pitch of the audio to be synthesized can be extracted from the target object based on the English dictionary standard and the international pitch standard. Then, the target object is input into a duration prediction model for duration prediction to obtain the phoneme duration of the target object. Among them, the duration prediction model in this application scenario is a duration prediction model that has been pre-trained based on English audio. That is, the duration prediction model corresponds to the language type of the target object.
[0056] Among them, the method of extracting the phonemes and pitch of the audio to be synthesized from the target object based on a preset standard can be carried out by using a trained deep learning network, manually, or by using audio software. Among them, the preset standard can be based on the language type of the target audio, and specific details are not limited here.
[0057] In other embodiments, the steps of obtaining phonemes, pitch, and phoneme duration may further include: obtaining the target text, and determining the corresponding phonemes, pitch, and phoneme duration based on the target text.
[0058] In a specific application scenario, the obtained phonemes, pitch, and phoneme duration can be subjected to feature extraction to obtain a feature matrix containing the feature information of phonemes, pitch, and phoneme duration. Then, the feature matrix containing the feature information of phonemes, pitch, and phoneme duration is input into a subsequent speech synthesis model to adapt to the input format of the model and improve the model synthesis efficiency.
[0059] Step S22: Perform frame windowing and Fourier transform on the object to be synthesized to obtain the linear spectrum of the object to be synthesized. Input the linear spectrum into a Mel filter bank for filtering to obtain a Mel spectrum. Input the Mel spectrum into a deep learning network, and use the deep learning network to extract the timbre feature matrix of the object to be synthesized.
[0060] Since the object to be synthesized in this embodiment is the audio to be synthesized, first, perform frame windowing and Fourier transform on the object to be synthesized to obtain the linear spectrum of the object to be synthesized. Then, input the linear spectrum into a Mel filter bank for filtering to obtain a Mel spectrum. Input the Mel spectrum into a deep learning network, and use the deep learning network to extract the timbre feature matrix of the object to be synthesized. Among them, the Fourier transform includes the short-time Fourier transform (STFT) or other Fourier transforms. Specifically, after performing frame windowing and Fourier transform on the object to be synthesized, the time-frequency domain signal of the object to be synthesized can be obtained, and then the linear spectrum of the object to be synthesized is generated based on the time-frequency domain signal of the object to be synthesized.
[0061] In a specific application scenario, after performing frame windowing and Fourier transform on the object to be synthesized, the time-frequency domain signal of the object to be synthesized can be obtained first, and then the amplitude of the time-frequency domain signal is taken to generate the linear spectrum of the object to be synthesized, so as to refine certain features in the linear spectrum, thereby improving the extraction speed. Then, the linear spectrum is input into the Mel filter bank for filtering to obtain the Mel spectrum. Finally, the Mel spectrum is input into the deep learning network to extract the timbre information of the audio to be synthesized, so as to extract the timbre feature matrix of the object to be synthesized by using the deep learning network.
[0062] Among them, the deep learning network in this step can include a combined network of a bidirectional gated recurrent unit (GRU) network + a fully connected layer (FC) + a Bayesian network (BN), a long short-term memory model, a convolutional neural network, or a recurrent neural network and other deep learning networks. The specific type can be set based on the actual situation and is not limited here. And the deep learning network in this step is a pre-trained network that can be applied for extracting the timbre feature matrix based on the Mel spectrum.
[0063] Among them, the execution order of this step and step S21 is not sequential. Step S21 can be executed first and then step S22, or step S22 can be executed first and then step S21, or step S22 and step S21 can be executed simultaneously. Specifically, step S21 only needs to be executed before step S23, and step S22 only needs to be executed before step S24.
[0064] Step S23: Encode the phonemes, pitch, and phoneme duration of the target object through the encoder in the speech synthesis model to obtain encoded data.
[0065] After obtaining the phonemes, pitch, and phoneme duration of the target object, the phonemes, pitch, and phoneme duration of the target object are input into the encoder in the speech synthesis model for encoding to obtain encoded data. The speech synthesis model in this embodiment is based on the tacotron1 model and is improved. Specifically, the speech synthesis model in this embodiment includes an encoder, an attention network, and a decoder that are cascaded with each other. Among them, since the feature extraction of phonemes, prosody, and pitch does not need to consider context information, the bidirectional GRU module in the tacotron1 model is removed in this embodiment to simplify the initial model structure, reduce ineffective model operations, and thereby improve the speech synthesis efficiency of the initial model and the final speech synthesis model.
[0066] In a specific application scenario, the input of the speech synthesis model can be (iou C3 10), where iou is phoneme information, C3 is pitch information, and 10 is the phoneme duration corresponding to the phoneme information.
[0067] Step S24: Decode the timbre feature matrix and the encoded data through the attention network and the decoder in the speech synthesis model in sequence to obtain the synthesized speech.
[0068] Input the encoded data output by the encoder and the timbre feature matrix extracted in step S22 into the attention network and the decoder in the speech synthesis model for decoding. After the decoding is completed, the decoder of the speech synthesis model outputs the synthesized speech.
[0069] Furthermore, in this embodiment, the attention mechanism Bahdanau Attention in the original Tacotron1 model is replaced with a location-based attention mechanism (Location Based Attention), that is, the initial model adopts a location-based attention mechanism. Among them, the location-based attention mechanism will ignore or reduce the silence in the speech data, so as to reduce the ineffective analysis and consumption of silence in speech synthesis, thereby improving the efficiency and quality of speech synthesis. Further, the initial model also adopts Forward Location Based Attention as the attention mechanism to further reduce the ineffective analysis and consumption of silence, thereby improving the efficiency and quality of speech synthesis.
[0070] In a specific application scenario, a sentence-breaking model can also be added to the speech synthesis model. The sentence-breaking model breaks the decoded data into sentences, thereby separating the entire synthesized speech to make it conform to the pronunciation rules, and further improving the quality of speech synthesis.
[0071] In a specific application scenario, the decoded data can be further input into the sentence-breaking model in the speech synthesis model for sentence-breaking to obtain the synthesized speech. Among them, the sentence-breaking model can include a Stop Token network. The trained sentence-breaking model breaks the decoded data into sentences to determine the breakpoints of the decoded data, thereby separating the entire synthesized speech to make it conform to the pronunciation rules, and further improving the quality of speech synthesis.
[0072] In another specific application scenario, the sentence-breaking model can also be set in the decoder of the speech synthesis model. While the decoder is decoding, the sentence-breaking model judges the decoding breakpoints when ending the decoding, thereby separating the decoded data in each segment, and then being able to separate the entire synthesized speech based on the decoded data in each separated segment to make it conform to the pronunciation rules, and further improving the quality of speech synthesis.
[0073] In a specific application scenario, the sentence segmentation rules of the target object can be manually input into the sentence segmentation model so that the sentence segmentation model can segment the decoded data or the data being decoded based on the sentence segmentation rules of the target object. In another specific application scenario, the sentence segmentation model can be trained in advance based on the sentence segmentation of the sample audio until it can segment sentences based on the phonemes, pitch, and phoneme duration of the target object.
[0074] Therefore, through the above improvements, the speech synthesis efficiency and effect of the speech synthesis model of this embodiment are superior to those of the traditional Tacotron1 model.
[0075] Among them, inputting the timbre feature matrix into the decoder of the speech synthesis model for decoding can adjust the timbre of the synthesized speech based on the timbre feature matrix, thereby controlling the timbre of the synthesized speech. That is, by modifying the timbre features of the Mel spectrogram, a timbre feature matrix with different features can be obtained, and thus the synthesized speech with different timbres can be actually output. The above speech synthesis method can be applied to the voice correction work of karaoke software or similar software, so that when correcting the voices of different users, speech synthesis can be performed based on the timbre of each user, and finally a synthesized speech that has been corrected and restores the user's own timbre can be output, thereby improving the quality and effect of speech synthesis. And this embodiment also generates synthesized speech based on the phonemes, pitch, and phoneme duration of the target object, so that the synthesized speech not only restores the user's own timbre but also can match the phonemes, pitch, and phoneme duration of the target object. Thus, on the basis of correcting the user's voice, the generation of synthesized speech based on the phonemes, pitch, and phoneme duration of the target object is further realized, so that the synthesized speech is no longer limited by the phonemes, pitch, and phoneme duration of the object to be synthesized itself, but can achieve diverse synthesis of the object to be synthesized by adjusting the phonemes, pitch, and phoneme duration of the target object. This improves the freedom and diversity of speech synthesis.
[0076] In a specific application scenario, when the user inputs the audio to be synthesized of a song, the speech synthesis method of this embodiment can learn the timbre of the user based on the audio to be synthesized and, by adjusting the phonemes, pitch, and phoneme duration of the target object, realize the speech synthesis of multiple songs based on the timbre of the user.
[0077] Please refer to Figure 3 , Figure 3 is Figure 2 the schematic flowchart of one embodiment of generating a speech synthesis model in the embodiment.
[0078] Step S31: Obtain the sample audio and extract the phonemes, pitch, and phoneme duration of the sample audio.
[0079] Obtain the sample audio and extract the phonemes, pitch, and phoneme duration of the sample audio.
[0080] In a specific application scenario, professional recorders can record sample audio based on the requirements of the desired phonemes, pitch, and phoneme duration. Among them, the phonemes, pitch, and phoneme duration of the sample audio respectively cover a preset phoneme range, a preset pitch range, and a preset phoneme duration range. Specifically, the preset phoneme range needs to cover all the phonemes of the language type corresponding to the sample audio; the preset pitch range needs to cover most or all of the pitch ranges; the preset phoneme duration range needs to cover most of the phoneme durations in the pronunciation rules related to the object to be synthesized. By setting the above preset phoneme range, preset pitch range, and preset phoneme duration range, the sample audio in this embodiment can be made to have a certain comprehensiveness and extensiveness, and the target objects that can be supported for adjustment in the speech synthesis model trained by the above sample audio also have comprehensiveness and extensiveness. Furthermore, the finally trained speech synthesis model can generate a large number of diverse voices containing different phonemes, pitch, and phoneme durations based on a small number of objects to be synthesized.
[0081] After obtaining the sample audio, all the phonemes, pitch, and phoneme duration of the sample audio are extracted from the sample audio.
[0082] Among them, the specific method of extracting all the phonemes, pitch, and phoneme duration of the sample audio from the sample audio is the same as that in the previous embodiment. Please refer to the previous text and will not be elaborated here.
[0083] Step S32: Input the phonemes, pitch, and phoneme duration of the sample audio into the encoder in the initial model for encoding to obtain sample encoded data.
[0084] Input the phonemes, pitch, and phoneme duration of the sample audio into the encoder and the attention network in the initial model in sequence for encoding to obtain sample encoded data.
[0085] Among them, the initial model in this embodiment is based on the tacotron1 model and is improved. Specifically, the initial model in this embodiment includes an encoder, an attention network, and a decoder. Among them, since the feature extraction of phonemes, pitch, and phoneme duration does not need to consider context information, the bidirectional GRU module in the tacotron1 model is removed in this embodiment to simplify the initial model structure, reduce invalid model operations, and thereby improve the speech synthesis efficiency of the initial model and the final speech synthesis model. Similarly, the bidirectional GRU module is also removed in the finally trained speech synthesis model.
[0086] Further, in this embodiment, the attention mechanism Bahdanau Attention in the original Tacotron1 model is replaced with a location-based attention mechanism (Location Based Attention), that is, the initial model adopts a location-based attention mechanism. Among them, the location-based attention mechanism will ignore or reduce the silence in the speech data, so as to reduce the ineffective analysis and consumption of the silence in speech synthesis, thereby improving the efficiency and quality of speech synthesis. Further, the initial model also adopts Forward Location Based Attention as the attention mechanism to further reduce the ineffective analysis and consumption of the silence, thereby improving the efficiency and quality of speech synthesis.
[0087] In a specific application scenario, the initial model can also add a Stop Token network to segment the decoded data through the StopToken network, so as to separate the entire synthesized speech, making it conform to the pronunciation rules, and further improving the quality of speech synthesis. Among them, the structure of the speech classification model obtained after training is the same as that of the initial model.
[0088] Step S33: Determine the sample Mel spectrogram of the sample audio, and extract the timbre feature matrix of the sample audio based on the sample Mel spectrogram.
[0089] Determine the sample Mel spectrogram of the sample audio based on the phonemes, pitch, and phoneme duration of the sample audio, and extract the timbre feature matrix of the sample audio based on the sample Mel spectrogram.
[0090] Among them, the specific methods for extracting the Mel spectrogram based on phonemes, pitch, and phoneme duration and extracting the timbre feature matrix based on the Mel spectrogram are the same as those in the foregoing embodiments. Please refer to the previous text and will not be elaborated here.
[0091] Step S34: Input the timbre feature matrix of the sample audio and the sample encoded data into the attention network and decoder in the initial model for decoding to obtain the sample synthesized speech.
[0092] Add the timbre feature matrix of the sample audio to the input of the decoder and input it into the decoder in the initial model together with the sample encoded data for decoding to obtain the sample synthesized speech.
[0093] Among them, inputting the timbre feature matrix of the sample audio into the decoder of the initial model for decoding can adjust the timbre of the sample synthesized speech based on the timbre feature matrix, thereby controlling the timbre of the sample synthesized speech. That is, the timbre features of the Mel spectrogram can be modified to obtain different timbre feature matrices, so as to achieve the purpose of controlling the sample synthesized speech to output different timbres. Thus, the timbre of the synthesized speech can be changed by adjusting the timbre feature matrix in the background. Thus, in the speech synthesis process of the speech synthesis model, it is possible to directly implement the speech synthesis of a specific timbre by modifying the timbre feature matrix without obtaining the object to be synthesized, thereby reducing the conditions for speech synthesis and improving the efficiency of speech synthesis.
[0094] Step S35: In response to the similarity between the sample synthesized speech and the sample audio meeting the preset requirements, obtain the speech synthesis model.
[0095] After obtaining the sample synthesized speech predicted by the initial model, compare the sample synthesized speech with the sample audio recorded by the voice actor. In response to the similarity between the sample synthesized speech and the sample audio meeting the preset requirements, obtain the speech synthesis model. Among them, the structure of the speech synthesis model is the same as that of the initial model.
[0096] In a specific application scenario, the similarity between the sample synthesized speech and the sample audio can be compared through a loss function. When the loss function converges, the similarity between the sample synthesized speech and the sample audio meets the preset requirements. Among them. The loss function can include cross-entropy loss function, absolute value loss function, square loss function, etc., which are not specifically limited here.
[0097] Among them, the preset requirements of this embodiment can include a similarity threshold or the convergence of the loss function. The similarity threshold can be set based on specific situations and is not limited here.
[0098] Since the sample audio of this embodiment covers most or all of the phonemes, pitch ranges, and phoneme duration ranges, therefore, after training is completed and the speech synthesis model is obtained, the speech synthesis model only needs a small number of objects to be synthesized to achieve the speech synthesis or even speech generation of most or all of the phonemes, pitch ranges, and phoneme duration ranges possessed by the sample audio. And the setting based on the timbre feature matrix enables the speech synthesis model to learn different timbre information, thereby improving the quality of speech synthesis. Moreover, the speech synthesis model of this embodiment can also improve the training quality of the model based on the improvement of the sample audio, thereby reducing the duration requirements and professional requirements of the objects to be synthesized, lowering the threshold of speech synthesis, and expanding the application scope of speech synthesis.
[0099] In a specific application scenario, after the model training is completed, the parameters of the encoder, the attention network, and the decoder can be frozen respectively, and then the unfrozen modules can be fine-tuned with mini-batch data. Specifically, based on the trained model, the parameters of the encoder are frozen, and then the decoder and the attention network are trained to determine the importance of the decoder in the entire model; based on the trained model, the parameters of the encoder and the attention network are frozen, and then the decoder is trained to determine the importance of the decoder in the entire model; then, based on the trained model, the decoder is frozen, and then the encoder and the attention network are trained... and so on for timbre impact analysis. Finally, it is obtained that the decoder has a greater impact on speech synthesis and timbre restoration. In a specific application scenario, the entire speech synthesis model can be migrated based on the influence degree of the decoder. Among them, since the fine-tuning has been fully utilized in this application scenario to analyze the influence of each part of the entire model, during the migration process, the stability of each parameter in the decoder is ensured to guarantee the migration training of new features and the stability of the model.
[0100] Through the above method, the speech synthesis method of this embodiment first extracts the phonemes and pitch of the target audio from the target audio based on a preset standard, inputs the object to be synthesized into the duration prediction model for duration prediction to obtain the phoneme duration of the object to be synthesized; and determines the Mel spectrogram of the object to be synthesized, and extracts the timbre feature matrix of the object to be synthesized based on the Mel spectrogram; encodes the phonemes, pitch, and phoneme duration of the target audio to obtain encoded data; decodes the timbre feature matrix and the encoded data to obtain the synthesized speech of the object to be synthesized. This embodiment encodes and decodes based on the phonemes, pitch, and phoneme duration of the object to be synthesized, which can refine the speech features on the basis of realizing speech synthesis, improve the accuracy and precision of speech synthesis, thereby improving the speech quality of the synthesized speech, and through encoding the timbre feature matrix, it can fully perform speech synthesis based on the timbre of the object to be synthesized, so that the finally generated synthesized speech can restore the timbre of the object to be synthesized, and further improve the effect of speech synthesis. And this embodiment uses the position-based attention mechanism to ignore or reduce the silence in the speech data, so that it can reduce the ineffective analysis and consumption of silence in speech synthesis, thereby improving the efficiency and quality of speech synthesis. And this embodiment also encodes the timbre feature matrix and the encoded data simultaneously, which can make the timbre of the synthesized speech be restored based on the timbre of the object to be synthesized, and at the same time make the phonemes, pitch, and phoneme duration of the synthesized speech match the phonemes, pitch, and phoneme duration of the target object, and then realize the synthesis of a variety of synthesized speeches based on a small number of objects to be synthesized by adjusting the phonemes, pitch, and phoneme duration of the target object.
[0101] Please refer toFigure 4 , Figure 4 is a schematic framework diagram of an embodiment of the voice synthesis device of the present application. The voice synthesis device 40 includes an acquisition module 41, an extraction module 42, an encoding module 43, and a decoding module 44. The acquisition module 41 is configured to acquire the phonemes, pitch, and phoneme duration of the target object; the extraction module 42 is configured to acquire the object to be synthesized, determine the Mel spectrum of the object to be synthesized, and extract the timbre feature matrix of the object to be synthesized based on the Mel spectrum; the encoding module 43 is configured to encode the phonemes, pitch, and phoneme duration of the target object through a voice synthesis model to obtain encoded data; the decoding module 44 is configured to decode the timbre feature matrix and the encoded data through a voice synthesis model to obtain the synthesized voice of the object to be synthesized.
[0102] The extraction module 42 is further configured to acquire the object to be synthesized, perform frame addition and windowing and Fourier transform on the object to be synthesized to obtain the linear spectrum of the object to be synthesized; input the linear spectrum into a Mel filter bank for filtering processing to obtain the Mel spectrum.
[0103] The extraction module 42 is further configured to input the Mel spectrum into a deep learning network and use the deep learning network to extract the timbre feature matrix of the object to be synthesized.
[0104] The encoding module 43 is further configured to encode the phonemes, pitch, and phoneme duration of the target object through an encoder in the voice synthesis model to obtain encoded data.
[0105] The decoding module 44 is further configured to decode the timbre feature matrix and the encoded data through an attention network and a decoder in the voice synthesis model to obtain the synthesized voice.
[0106] The decoding module 44 is further configured to perform sentence segmentation on the decoded data through a sentence segmentation model in the voice synthesis model to obtain the synthesized voice.
[0107] The acquisition module 41 is further configured to acquire a sample audio, and extract the phonemes, pitch, and phoneme duration of the sample audio, wherein the phonemes, pitch, and phoneme duration of the sample audio respectively cover a preset phoneme range, a preset pitch range, and a preset phoneme duration range; input the phonemes, pitch, and phoneme duration of the sample audio into an encoder in an initial model for encoding to obtain sample encoded data; determine the sample Mel spectrum of the sample audio, and extract the timbre feature matrix of the sample audio based on the sample Mel spectrum; input the timbre feature matrix and the sample encoded data of the sample audio into an attention network and a decoder in the initial model in sequence for decoding to obtain a sample synthesized voice; in response to the similarity between the sample synthesized voice and the sample audio meeting a preset requirement, the voice synthesis model is obtained.
[0108] The obtaining module 41 is further configured to obtain the target audio and extract the corresponding phonemes, pitches, and phoneme durations from the target audio; or obtain the target text and determine the corresponding phonemes, pitches, and phoneme durations based on the target text.
[0109] In the above solution, by performing voice synthesis on the object to be synthesized based on the phonemes, pitches, and phoneme durations of the target object, it is possible to implement the phonemes, pitches, and phoneme durations of the target object in the synthesized voice and restore the timbre of the object to be synthesized, thereby improving the effect of voice synthesis.
[0110] Please refer to Figure 5 , Figure 5 , which is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 50 includes a memory 51 and a processor 52 that are coupled to each other. The processor 52 is configured to execute program instructions stored in the memory 51 to implement the steps of the voice synthesis method in any of the above embodiments. In a specific implementation scenario, the electronic device 50 may include, but is not limited to: a microcomputer, a server. In addition, the electronic device 50 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.
[0111] Specifically, the processor 52 is configured to control itself and the memory 51 to implement the steps of any of the above voice synthesis method embodiments. The processor 52 may also be referred to as a CPU (Central Processing Unit). The processor 52 may be an integrated circuit chip with signal processing capabilities. The processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 52 may be implemented jointly by integrated circuit chips.
[0112] In the above solution, by performing voice synthesis on the object to be synthesized based on the phonemes, pitches, and phoneme durations of the target object, it is possible to implement the phonemes, pitches, and phoneme durations of the target object in the synthesized voice and restore the timbre of the object to be synthesized, thereby improving the effect of voice synthesis.
[0113] Please refer to Figure 6 , Figure 6This is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 60 stores program instructions 601 that can be run by a processor, and the program instructions 601 are used to implement the steps of the voice synthesis method in any of the above embodiments.
[0114] The above solution can implement voice synthesis based on phonemes, pitch, and phoneme duration, and can restore the timbre of the object to be synthesized in the synthesized voice, improving the effect of voice synthesis.
[0115] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0116] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0118] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
Claims
1. A voice synthesis method, characterized in that, The voice synthesis method includes: Obtaining the phonemes, pitch, and phoneme duration of a target object, where the target object is a target audio; and Obtaining a to-be-synthesized object, and determining the Mel spectrogram of the to-be-synthesized object. Specifically, frame addition and windowing, as well as Fourier transform, are performed on the to-be-synthesized object to obtain the time-frequency domain signal of the to-be-synthesized object. The amplitude of the time-frequency domain signal is taken to generate a linear spectrogram, and the linear spectrogram is input into a Mel filter bank for filtering processing to obtain the Mel spectrogram; the to-be-synthesized object is a to-be-synthesized audio; Extracting the timbre feature matrix of the to-be-synthesized object based on the Mel spectrogram; Encoding the phonemes, pitch, and phoneme duration of the target object through a voice synthesis model to obtain encoded data; the voice synthesis model includes a cascaded encoder, attention network, and decoder; the voice synthesis model is a tacotron1 model without a bidirectional GRU module, and a position-based attention mechanism is set in the attention network; Decoding the timbre feature matrix and the encoded data through the voice synthesis model to obtain the synthesized voice of the to-be-synthesized object.
2. The speech synthesis method according to claim 1, wherein The step of extracting the timbre feature matrix of the to-be-synthesized object based on the Mel spectrogram includes: Inputting the Mel spectrogram into a deep learning network, and using the deep learning network to extract the timbre feature matrix of the to-be-synthesized object.
3. The voice synthesis method according to claim 1, wherein The step of encoding the phonemes, pitch, and phoneme duration of the target object through a voice synthesis model to obtain encoded data includes: Encoding the phonemes, pitch, and phoneme duration of the target object through the encoder in the voice synthesis model to obtain encoded data; The step of decoding the timbre feature matrix and the encoded data through the voice synthesis model to obtain the synthesized voice of the to-be-synthesized object includes: Sequentially decoding the timbre feature matrix and the encoded data through the attention network and decoder in the voice synthesis model to obtain the synthesized voice.
4. The voice synthesis method according to claim 3, wherein The attention network includes a position-based attention mechanism.
5. The speech synthesis method according to claim 3, wherein The voice synthesis model further includes a sentence-breaking model; The step of decoding the timbre feature matrix and the encoded data through the decoder in the voice synthesis model to obtain the synthesized voice further includes: Performing sentence-breaking on the decoded data through the sentence-breaking model in the voice synthesis model to obtain the synthesized voice.
6. The speech synthesis method according to any one of claims 3-5, characterized in that Before the step of obtaining the phonemes, pitch, and phoneme duration of the target object, it includes: Obtaining a sample audio, and extracting the phonemes, pitch, and phoneme duration of the sample audio, where the phonemes, pitch, and phoneme duration of the sample audio respectively cover a preset phoneme range, a preset pitch range, and a preset phoneme duration range; Inputting the phonemes, pitch, and phoneme duration of the sample audio into the encoder in the initial model for encoding to obtain sample encoded data; Determine the sample Mel spectrogram of the sample audio, and extract the timbre feature matrix of the sample audio based on the sample Mel spectrogram; Input the timbre feature matrix of the sample audio and the sample coding data into the attention network and decoder in the initial model in sequence for decoding to obtain the sample synthesized speech; In response to the similarity between the sample synthesized speech and the sample audio meeting the preset requirements, the speech synthesis model is obtained.
7. The speech synthesis method according to claim 1, wherein The step of obtaining the phonemes, pitch and phoneme duration of the target object includes: Obtain the target audio, and extract the corresponding phonemes, pitch and phoneme duration from the target audio; or Obtain the target text, and determine the corresponding phonemes, pitch and phoneme duration based on the target text.
8. A voice synthesis device, characterized in that, The speech synthesis device includes: An acquisition module, configured to acquire the phonemes, pitch and phoneme duration of a target object, where the target object is the target audio; An extraction module, configured to obtain an object to be synthesized, and determine the Mel spectrogram of the object to be synthesized. Specifically, perform frame addition and windowing and Fourier transform on the object to be synthesized to obtain the time-frequency domain signal of the object to be synthesized, take the amplitude of the time-frequency domain signal to generate a linear spectrum, and input the linear spectrum into a Mel filter bank for filtering processing to obtain the Mel spectrogram; the object to be synthesized is the audio to be synthesized; Extract the timbre feature matrix of the object to be synthesized based on the Mel spectrogram; An encoding module, configured to encode the phonemes, pitch and phoneme duration of the target object through a speech synthesis model to obtain encoding data; the speech synthesis model includes a cascaded encoder, attention network and decoder; the speech synthesis model is the tacotron1 model without a bidirectional GRU module, and a position-based attention mechanism is set in the attention network; A decoding module, configured to decode the timbre feature matrix and the encoding data through the speech synthesis model to obtain the synthesized speech of the object to be synthesized.
9. An electronic device, characterized in that, It includes a memory and a processor coupled to each other, and the processor is configured to execute program instructions stored in the memory to implement the speech synthesis method according to any one of claims 1 to 7.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the speech synthesis method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Speech synthesis model training method and device, computer equipment and storage medium
CN111133506A
Speech synthesis method and system for new tone generation
CN112802448A
Speech synthesis method and device, electronic equipment and storage medium
CN113380222A