Method, apparatus, and storage medium for speech synthesis
By introducing vocalprint feature vectors and speaker recognition models in speech synthesis, the problems of poor voice quality and large recording volume requirements in the prior art are solved, and high-quality speech synthesis and the ability to quickly adapt to new speakers are achieved.
Patent Information
- Application Number
- CN202010988010.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-09-18
AI Technical Summary
In the prior art, when using end-to-end network structures directly for speech synthesis, the network structure of the speaker cannot be effectively introduced, resulting in poor voice quality and huge demands on recording volume during the training process.
By obtaining the target text information and target speech information, the vocal print feature vector is determined, and combined with the pre-trained speech synthesis model, Mel spectral features are generated to improve the quality of speech synthesis. In addition, by introducing a speaker recognition model, the speech synthesis model of new speakers is quickly trained to reduce the recording volume requirement.
The quality of speech synthesis is significantly improved and the need for recording volume during training is greatly reduced, achieving a speech synthesis model that quickly adapts to the new speaker.
Smart Images

Figure CN114283777B_ABST
Abstract
Claims
1. A method for speech synthesis, characterized in that, it includes: obtaining target text information for speech synthesis and target speech information corresponding to the target text information; determining a voiceprint feature vector corresponding to the target speech information; and generating Mel spectrum features according to the target text information, the voiceprint feature vector, and a pre-trained speech synthesis model, and determining a target synthesized speech according to the Mel spectrum features; the pre-trained speech synthesis model includes an encoder and a decoder, and the operation of generating Mel spectrum features according to the target text information, the voiceprint feature vector, and the pre-trained speech synthesis model includes: generating a pronunciation feature vector according to the target text information and the encoder; concatenating the voiceprint feature vector and the pronunciation feature vector; and using the decoder to calculate the concatenated voiceprint feature vector and pronunciation feature vector to generate the Mel spectrum features; the operation of generating a pronunciation feature vector according to the target text information and the encoder includes: determining first pronunciation feature information corresponding to the pronunciation features of the target text information; and using the encoder to calculate the first pronunciation feature information to determine the pronunciation feature vector; the operation of determining first pronunciation feature information corresponding to the pronunciation features of the target text information includes: determining the character sequence included in the target text information; sequentially determining the pinyin information corresponding to the characters in the character sequence, where the pinyin information includes the letters of the pinyin and the tones of the characters; using a preset sorting rule to determine the sequence numbers of the letters and tones included in the pinyin information, where the sorting rule is used to describe the order of the twenty-six English letters and the five tones; and respectively determining second pronunciation feature information corresponding to the characters in the character sequence according to the sequence numbers of the letters and tones, and sequentially concatenating the second pronunciation feature information corresponding to the characters in the character sequence to determine the first pronunciation feature information.
2. The method according to claim 1, characterized in that, the operation of determining a voiceprint feature vector corresponding to the target speech information includes: determining the MFCC features of the target speech information; and generating the voiceprint feature vector according to the MFCC features and a pre-trained speaker recognition model.
3. The method according to claim 1, characterized in that, it further includes: training the speech synthesis model according to the following steps: obtaining first sample speech information that has completed quality detection and first sample text information corresponding to the first sample speech information; determining the actual Mel spectrum features of the first sample speech information; using a pre-trained speaker recognition model to generate an output voiceprint feature vector corresponding to the first sample speech information; using the encoder in the speech synthesis model to generate an output pronunciation feature vector corresponding to the first sample text information; Using the decoder in the voice synthesis model, generate an output Mel spectrogram feature corresponding to the spliced output pronunciation feature vector and the output voiceprint feature vector; and Compare the actual Mel spectrogram feature with the output Mel spectrogram feature, and adjust the voice synthesis model according to the comparison result.
4. The method according to claim 3, wherein, further comprising: Obtain second sample voice information and second sample text information corresponding to the second sample voice information, wherein the second sample voice information is voice information that has not undergone quality detection, and the data volume of the second sample voice information is smaller than the data volume of the first sample voice information; and Use the second sample voice information and the second sample text information to perform transfer learning training on the voice synthesis model.
5. The method according to claim 4, wherein, further comprising: Train the speaker recognition model according to the following steps: Obtain a plurality of audio data, wherein the plurality of audio data are respectively audio data of a plurality of different speakers; Determine a plurality of MFCC features corresponding to the plurality of audio data; Use the speaker recognition model and the plurality of MFCC features to determine a plurality of output voiceprint feature vectors respectively corresponding to the plurality of audio data; Compare the plurality of output voiceprint feature vectors pairwise, and determine whether the audio data corresponding to the output voiceprint feature vectors being compared are audio data of the same speaker according to the comparison result; and Adjust the speaker recognition model according to the determined output result and a preset actual result.
6. A storage medium, wherein, The storage medium includes a stored program, wherein, when the program runs, the method according to any one of claims 1 to 5 is executed by a processor.
7. A voice synthesis device, wherein, comprising: An acquisition module for acquiring target text information for voice synthesis and target voice information corresponding to the target text information; A determination module for determining a voiceprint feature vector corresponding to the target voice information; and A generation module for generating a Mel spectrogram feature according to the target text information, the voiceprint feature vector, and a pre-trained voice synthesis model, and determining a target synthesized voice according to the Mel spectrogram feature; The pre-trained voice synthesis model includes an encoder and a decoder, and the operation of generating a Mel spectrogram feature according to the target text information, the voiceprint feature vector, and the pre-trained voice synthesis model includes: Generate a pronunciation feature vector according to the target text information and the encoder; Splice the voiceprint feature vector and the pronunciation feature vector; and Use the decoder to calculate the spliced voiceprint feature vector and pronunciation feature vector to generate the Mel spectrogram feature; The operation of generating a pronunciation feature vector according to the target text information and the encoder includes: Determine first pronunciation feature information corresponding to the pronunciation feature of the target text information; and Calculate the first pronunciation feature information using the encoder to determine the pronunciation feature vector; The operation of determining the first pronunciation feature information corresponding to the pronunciation feature of the target text information includes: Determine the character sequence included in the target text information; Successively determine the pinyin information corresponding to the characters in the character sequence, where the pinyin information includes the letters of the pinyin and the tones of the characters; Use a preset sorting rule to determine the sequence numbers of the letters and the tones included in the pinyin information, where the sorting rule is used to describe the order of the twenty-six English letters and the five tones; and Respectively determine the second pronunciation feature information corresponding to the characters in the character sequence according to the sequence numbers of the letters and the tones, and successively splice the second pronunciation feature information corresponding to the characters in the character sequence to determine the first pronunciation feature information.
Citation Information
Patent Citations
Text-independent voiceprint recognition method
CN108648759A
Speech synthesis method and device for text, as well as computer equipment
CN109754778A