Method and apparatus for training timbre feature extraction model and audio synthesis
By extracting and training the sample audio of different characters, the problem of low training efficiency of the tone feature extraction model in the prior art is solved, and the efficiency and accuracy of extracting tone features for any tone is achieved.
Patent Information
- Application Number
- CN202211485541.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-11-24
AI Technical Summary
In the prior art, the training efficiency of the tone feature extraction model is low, and it is necessary to assign new ID and sample audio to each specified tone, resulting in low model training and application efficiency.
By obtaining sample audio of different characters, using the tone feature extraction model to be trained to extract tone features for these audios, and training the tone feature extraction model is carried out to minimize the differences in similar tone features and maximize the differences in different tone features.
The ability to extract timbre features for any tone without additional training is achieved, improving the efficiency and accuracy of timbre feature extraction.
Smart Images

Figure CN115862586B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for training a timbre feature extraction model and audio synthesis. Background Art
[0002] In recent years, artificial intelligence has been applied to the TTS (text to speech) technology. The TTS technology generates corresponding speech audio according to the input text and has wide applications in scenarios such as voice assistants, chatbots, audiobooks, and virtual humans. In the TTS technology, it is generally necessary to synthesize audio of a specific timbre. Therefore, a timbre feature extraction model is necessary.
[0003] Generally, before training a timbre feature extraction model, it is necessary to record sample audio of different people. When recording audio, each person is assigned an ID (identity document). The training process of the timbre feature extraction model is as follows: First, obtain the ID, input the ID into the timbre feature extraction model to be trained to obtain the timbre feature, and at the same time obtain the text pronunciation feature of the target text; then, input the extracted timbre feature and the text pronunciation feature into the audio synthesis model to be trained to obtain the predicted audio; finally, train the timbre feature extraction model with the aim of minimizing the difference between the predicted audio and the sample audio corresponding to the ID (the sample audio is the audio obtained by reading the target text with the timbre corresponding to the ID).
[0004] The training samples used in the above training process are sample audio of certain specified timbres. The IDs corresponding to the sample audio of the same timbre are the same. The trained timbre feature extraction model can only output the timbre features of these specified timbres. When a new specified timbre needs to be added, a new ID needs to be assigned, and the sample audio of the new timbre and the new ID are used to retrain the timbre feature extraction model. Then, the timbre feature extraction model can output the timbre features corresponding to the timbre when the ID is input. This will affect the efficiency of timbre feature extraction. Summary of the Invention
[0005] Embodiments of this application provide a method and device for training a timbre feature extraction model and audio synthesis, which can solve the problem of low efficiency of timbre feature extraction. The technical solutions are as follows:
[0006] In a first aspect, a method for training a timbre feature extraction model is provided. The method includes:
[0007] Obtain the first sample audio of the first person, the second sample audio of the first person, and the third sample audio of the second person;
[0008] Extract the first timbre feature from the first sample audio according to the timbre feature extraction model to be trained, extract the second timbre feature from the second sample audio according to the timbre feature extraction model to be trained, and extract the third timbre feature from the third sample audio according to the timbre feature extraction model to be trained;
[0009] Train the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature; if the training end condition is met, determine the timbre feature extraction model that meets the training end condition as the target timbre feature extraction model.
[0010] In a possible implementation, the training of the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature includes:
[0011] Train the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, and maximizing the difference between the second timbre feature and the third timbre feature.
[0012] In a possible implementation, the first timbre feature is a first timbre feature vector, the second timbre feature is a second timbre feature vector, and the third timbre feature is a third timbre feature vector;
[0013] The training of the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, and maximizing the difference between the second timbre feature and the third timbre feature includes:
[0014] Determine a first vector angle between the first timbre feature vector and the second timbre feature vector, determine a second vector angle between the first timbre feature vector and the third timbre feature vector, and determine a third vector angle between the second timbre feature vector and the third timbre feature vector;
[0015] Determine a first loss value according to the first vector angle, the second vector angle, and the third vector angle, where the first loss value is positively correlated with the first vector angle, negatively correlated with the second vector angle, and negatively correlated with the third vector angle;
[0016] Train the timbre feature extraction model to be trained according to the first loss value.
[0017] In a possible implementation, determining the first loss value according to the first vector angle, the second vector angle, and the third vector angle includes:
[0018] Determining a first cosine value of the first vector angle, a second cosine value of the second vector angle, and a third cosine value of the third vector angle;
[0019] Determining the first loss value according to the first cosine value, the second cosine value, and the third cosine value, where the first loss value is negatively correlated with the first cosine value, positively correlated with the second cosine value, and positively correlated with the third cosine value.
[0020] In a possible implementation, determining the first loss value according to the first cosine value, the second cosine value, and the third cosine value includes:
[0021] Determining a first sub-loss value according to the first cosine value and the second cosine value, determining a second sub-loss value according to the second cosine value, determining a third sub-loss value according to the third cosine value, and determining a fourth sub-loss value according to the first cosine value, where the first sub-loss value is negatively correlated with the first cosine value and positively correlated with the second cosine value, the second sub-loss value is positively correlated with the second cosine value, the third sub-loss value is positively correlated with the third cosine value, and the fourth sub-loss value is negatively correlated with the first cosine value;
[0022] Determining the first loss value according to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value.
[0023] In a possible implementation, determining the first sub-loss value according to the first cosine value and the second cosine value; determining the second sub-loss value according to the second cosine value; determining the third sub-loss value according to the third cosine value; determining the fourth sub-loss value according to the first cosine value includes:
[0024] Determining the first sub-loss value L according to the formula 1 ;
[0025] Determining the second sub-loss value L 2 = cos(y a , y n ) according to the formula; 2 ;
[0026] Determining the third sub-loss value L 3 = cos(y p , y n ) according to the formula; 3 ;
[0027] According to the formula L 4 = -cos(y a , y p ), determine the fourth sub - loss value L 4 ;
[0028] where y a is the first timbre feature vector, y p is the second timbre feature vector, y n is the third timbre feature vector, cos(y a , y p ) is the first cosine value, cos(y a , y n ) is the second cosine value, cos(y p , y n ) is the third cosine value.
[0029] In one possible implementation, determining the first loss value according to the first sub - loss value, the second sub - loss value, the third sub - loss value and the fourth sub - loss value includes:
[0030] Determine the weighted average of the first sub - loss value, the second sub - loss value, the third sub - loss value and the fourth sub - loss value according to the first weight, the second weight, the third weight and the fourth weight, and obtain the first loss value.
[0031] In one possible implementation, the first sample audio and the third sample audio are reading audios corresponding to the same text, and the first sample audio and the second sample audio are reading audios corresponding to different texts.
[0032] In one possible implementation, before determining the timbre feature extraction model that meets the training end condition as the target timbre feature extraction model if the training end condition is met, the method further includes:
[0033] Determine the phoneme sequence corresponding to the sample text, input the phoneme sequence into the encoder to be trained, and obtain the text pronunciation feature corresponding to the sample text; and input the text pronunciation feature and the first timbre feature into the audio synthesis model to be trained to obtain the predicted audio;
[0034] Determine the second loss value according to the predicted audio and the first sample audio; and train the encoder to be trained and the audio synthesis model to be trained according to the second loss value;
[0035] Training the to-be-trained timbre feature extraction model with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature includes: training the to-be-trained timbre feature extraction model with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and according to the second loss value.
[0036] If the training end condition is satisfied, the method further includes:
[0037] Determining the encoder that satisfies the training end condition as the target encoder and determining the audio synthesis model that satisfies the training end condition as the target audio synthesis model.
[0038] In a possible implementation manner, after inputting the text pronunciation feature and the first timbre feature into the to-be-trained audio synthesis model to obtain a predicted audio, the method further includes:
[0039] Obtaining a fourth sample audio of the first person and a fifth sample audio of the second person;
[0040] Extracting a fourth timbre feature from the predicted audio according to the to-be-trained timbre feature extraction model, extracting a fifth timbre feature from the fourth sample audio according to the to-be-trained timbre feature extraction model, and extracting a sixth timbre feature from the fifth sample audio according to the to-be-trained timbre feature extraction model;
[0041] The training of the to-be-trained encoder and the to-be-trained audio synthesis model according to the second loss value includes: training the to-be-trained encoder and the to-be-trained audio synthesis model with the aim of minimizing the difference between the fourth timbre feature and the fifth timbre feature and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and according to the second loss value.
[0042] Training the to-be-trained timbre feature extraction model with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and according to the second loss value, includes: training the to-be-trained timbre feature extraction model with the aim of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, minimizing the difference between the fourth timbre feature and the fifth timbre feature, and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and according to the second loss value.
[0043] Second aspect, a method for timbre synthesis is provided, the method comprising:
[0044] extracting timbre features from a target audio according to the target timbre feature extraction model as described in the first aspect;
[0045] determining a target phoneme sequence corresponding to the target text, inputting the target phoneme sequence into a target encoder, and obtaining a target text pronunciation feature corresponding to the target text;
[0046] inputting the target text pronunciation feature and the timbre feature of the target audio into a target audio synthesis model to obtain a synthesized audio.
[0047] Third aspect, a training device for a timbre feature extraction model is provided, the device comprising:
[0048] an acquisition module, configured to acquire a first sample audio of a first person, a second sample audio of the first person, and a third sample audio of a second person;
[0049] an extraction module, configured to extract a first timbre feature from the first sample audio according to a timbre feature extraction model to be trained, extract a second timbre feature from the second sample audio according to the timbre feature extraction model to be trained, and extract a third timbre feature from the third sample audio according to the timbre feature extraction model to be trained;
[0050] a training module, configured to train the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature; if a training end condition is satisfied, determining the timbre feature extraction model that satisfies the training end condition as the target timbre feature extraction model.
[0051] In a possible implementation manner, the training module is configured to:
[0052] train the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, and maximizing the difference between the second timbre feature and the third timbre feature.
[0053] In a possible implementation manner, the first timbre feature is a first timbre feature vector, the second timbre feature is a second timbre feature vector, and the third timbre feature is a third timbre feature vector;
[0054] the training module is configured to:
[0055] Determine a first vector angle between the first timbre feature vector and the second timbre feature vector, determine a second vector angle between the first timbre feature vector and the third timbre feature vector, and determine a third vector angle between the second timbre feature vector and the third timbre feature vector;
[0056] Determine a first loss value according to the first vector angle, the second vector angle, and the third vector angle, where the first loss value is positively correlated with the first vector angle, negatively correlated with the second vector angle, and negatively correlated with the third vector angle;
[0057] Train the timbre feature extraction model to be trained according to the first loss value.
[0058] In a possible implementation, the training module is configured to:
[0059] Determine a first cosine value of the first vector angle, a second cosine value of the second vector angle, and a third cosine value of the third vector angle;
[0060] Determine a first loss value according to the first cosine value, the second cosine value, and the third cosine value, where the first loss value is negatively correlated with the first cosine value, positively correlated with the second cosine value, and positively correlated with the third cosine value.
[0061] In a possible implementation, the training module is configured to:
[0062] Determine a first sub-loss value according to the first cosine value and the second cosine value, determine a second sub-loss value according to the second cosine value, determine a third sub-loss value according to the third cosine value, and determine a fourth sub-loss value according to the first cosine value, where the first sub-loss value is negatively correlated with the first cosine value and positively correlated with the second cosine value, the second sub-loss value is positively correlated with the second cosine value, the third sub-loss value is positively correlated with the third cosine value, and the fourth sub-loss value is negatively correlated with the first cosine value;
[0063] Determine a first loss value according to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value.
[0064] In a possible implementation, the training module is configured to:
[0065] According to the formula Determine the first sub-loss value L 1 ;
[0066] According to the formula L 2 = cos(y a ,y n), determine the second sub-loss value L 2 ;
[0067] According to the formula L 3 = cos(y p , y n ), determine the third sub-loss value L 3 ;
[0068] According to the formula L 4 = -cos(y a , y p ), determine the fourth sub-loss value L 4 ;
[0069] where y a is the first timbre feature vector, y p is the second timbre feature vector, y n is the third timbre feature vector, cos(y a , y p ) is the first cosine value, cos(y a , y n ) is the second cosine value, cos(y p , y n ) is the third cosine value.
[0070] In a possible implementation, the training module is configured to:
[0071] According to the first weight, the second weight, the third weight, and the fourth weight, determine the weighted average of the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value, and obtain the first loss value.
[0072] In a possible implementation, the first sample audio and the third sample audio are reading audios corresponding to the same text, and the first sample audio and the second sample audio are reading audios corresponding to different texts.
[0073] In a possible implementation, the device further includes a synthesis module;
[0074] The synthesis module is configured to determine the phoneme sequence corresponding to the sample text, input the phoneme sequence into the encoder to be trained, and obtain the text pronunciation feature corresponding to the sample text; input the text pronunciation feature and the first timbre feature into the audio synthesis model to be trained, and obtain the predicted audio;
[0075] The training module is further configured to determine a second loss value according to the predicted audio and the first sample audio; and train the encoder to be trained and the audio synthesis model to be trained according to the second loss value; with the purpose of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and train the timbre feature extraction model to be trained according to the second loss value; if the training end condition is satisfied, determine the encoder that satisfies the training end condition as the target encoder, and determine the audio synthesis model that satisfies the training end condition as the target audio synthesis model.
[0076] In a possible implementation manner, the obtaining module is further configured to obtain a fourth sample audio of the first person and a fifth sample audio of the second person;
[0077] The extraction module is further configured to extract a fourth timbre feature from the predicted audio according to the timbre feature extraction model to be trained, extract a fifth timbre feature from the fourth sample audio according to the timbre feature extraction model to be trained, and extract a sixth timbre feature from the fifth sample audio according to the timbre feature extraction model to be trained;
[0078] The training module is further configured to take minimizing the difference between the fourth timbre feature and the fifth timbre feature and maximizing the difference between the fourth timbre feature and the sixth timbre feature as the training purpose, and train the encoder to be trained and the audio synthesis model to be trained according to the second loss value; with the purpose of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, minimizing the difference between the fourth timbre feature and the fifth timbre feature, and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and train the timbre feature extraction model to be trained according to the second loss value.
[0079] In a fourth aspect, a timbre extraction device is provided, and the device includes:
[0080] An extraction module, configured to extract a timbre feature from a target audio;
[0081] A determination module, configured to determine a target phoneme sequence corresponding to a target text, input the target phoneme sequence into a target encoder, and obtain a target text pronunciation feature corresponding to the target text;
[0082] A synthesis module, configured to input the target text pronunciation feature and the timbre feature of the target audio into a target audio synthesis model to obtain a synthesized audio.
[0083] Fifth aspect, a computer device is provided. The computer device includes a memory and a processor. The memory is used to store computer instructions. The processor executes the computer instructions stored in the memory so that the computer device executes the method according to the first aspect and its possible implementation manners.
[0084] Sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program code. When the computer program code is executed by a computer device, the computer device executes the method according to the first aspect and its possible implementation manners.
[0085] Seventh aspect, a computer program product is provided. The computer program product includes computer program code. When the computer program code is executed by a computer device, the computer device executes the method according to the first aspect and its possible implementation manners.
[0086] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0087] Through the method provided by the embodiments of the present application, it is necessary to obtain the first sample audio and the second sample audio of the first person and the third sample audio of the second person, and respectively extract the timbre features of these three sample audios through a timbre feature extraction model. The extracted audio features are the first timbre feature, the second timbre feature, and the third timbre feature respectively. The timbre feature extraction model is trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature. After the training end condition is reached, the timbre feature extraction model that meets the training end condition can be determined as the target timbre feature extraction model. In this way, for the target timbre feature extraction model trained by this method, when extracting the timbre features of an audio with any timbre, there is no need to perform additional training on the model. The audio with this timbre can be directly input into the timbre feature extraction model, and the timbre features can be obtained, thereby improving the efficiency of timbre feature extraction. Description of the Drawings
[0088] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0089] Figure 1 It is a schematic flowchart of a method for training a timbre feature extraction model provided by an embodiment of the present application;
[0090] Figure 2It is a schematic flowchart of a process for calculating a first loss value provided by an embodiment of the present application;
[0091] Figure 3 It is a schematic flowchart of a training method for a timbre feature extraction model provided by an embodiment of the present application;
[0092] Figure 4 It is a schematic flowchart of a method for calculating a second loss value provided by an embodiment of the present application;
[0093] Figure 5 It is a schematic flowchart of a training method for a timbre feature extraction model provided by an embodiment of the present application;
[0094] Figure 6 It is a schematic flowchart of a method for calculating a third loss value provided by an embodiment of the present application;
[0095] Figure 7 It is a schematic flowchart of a usage method for a timbre feature extraction model provided by an embodiment of the present application;
[0096] Figure 8 It is a schematic structural diagram of a training device for a timbre feature extraction model provided by an embodiment of the present application;
[0097] Figure 9 It is a schematic structural diagram of a device for timbre synthesis provided by an embodiment of the present application;
[0098] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0099] An embodiment of the present application provides a training method for a timbre feature extraction model. The timbre feature extraction model is used in the process of converting a piece of text into an audio with a specific timbre and is applied to an application program with a speech synthesis function.
[0100] This method can be implemented by a computer device, which can be a terminal or a server. The terminal can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc. An application program with a speech synthesis function can be installed in the terminal. The server can be a background server of an application program with a speech synthesis function. The server can be a single server or a group of devices composed of multiple devices.
[0101] Viewed from the hardware composition, the structure of the computer device can include a processor and a memory.
[0102] The processor can be a CPU (central processing unit) or an SoC (system on chip), etc. The processor can be used to execute various instructions involved in this method, etc.
[0103] The memory can include various volatile memories or non-volatile memories, such as SSD (solid state disk), DRAM (dynamic random access memory), etc. The memory can be used to store pre-stored data, intermediate data, and result data during the training process, for example, sample audio, etc.
[0104] In addition to the processor and the memory, the terminal can also include a display component, a communication component, an audio acquisition component, an audio output component, etc.
[0105] The display component can be an independent screen, or a screen integrated with the terminal body, a projector, etc. The screen can be a touch screen or a non-touch screen. The display component is used to display the system interface, application interface, etc., for example, the interface for recording audio, etc.
[0106] The communication component can be a wired network connector, a WiFi (wireless fidelity) module, a Bluetooth module, a cellular network communication module, etc. The communication component can be used to transmit data with other devices, and the other devices can be a server or other terminals, etc.
[0107] The audio acquisition component can be a microphone, which is used to acquire the user's voice. The audio output component can be a speaker, headphones, etc., which are used to play audio.
[0108] From the perspective of the hardware composition, the structure of the server can include a processor and a memory.
[0109] The processor can be a CPU or an SoC, etc. The processor can be used to execute various instructions involved in this method, etc.
[0110] The memory can include various volatile memories or non-volatile memories, such as SSD, DRAM memory, etc. The memory can be used to store pre-stored data, intermediate data, and result data during the training process, for example, sample audio, audio features, predicted audio, etc.
[0111] Next, several terms involved in this embodiment will be introduced:
[0112] TTS (text to speech): Generate the corresponding voice audio according to the input text.
[0113] MFC (mel-frequency spectrum): hereinafter referred to as the Mel spectrum. The Mel spectrum is a spectrum that can be used to represent short-term audio. The Mel-frequency cepstral coefficients are the coefficients that make up the Mel spectrum and are usually used as feature observations in a speech recognition system.
[0114] Phoneme: The smallest speech unit divided according to the natural attributes of speech. Analyzed according to the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes are divided into two major categories: vowels and consonants. For example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, "dài" has three phonemes, etc.
[0115] In a speech synthesis assistant, when a user needs to synthesize a corresponding audio for a piece of text, the user can enter a text content in the text box, click on the avatar of a person in the person list to select the voice color of the person for synthesizing the audio, and finally click the synthesis button to obtain the audio with the corresponding voice color. The synthesized audio can be used for dubbing or audio-visual editing, etc. In an intelligent navigation, when the user selects a destination location, after the terminal obtains this operation, it generates route prompt text in real time based on the located position and route. The terminal generates navigation audio in real time according to this prompt text and the voice color preset by the user. In an e-book, when the user needs to convert a certain text into audio, the user can click on voice playback. When the terminal obtains the user's operation, it synthesizes the text into audio with the preset voice color for playback.
[0116] The embodiment of the present application provides a method for training a voice color feature extraction model for the above application scenarios. The processing flow of this method can be as Figure 1 shown, including the following processing steps:
[0117] 101. Obtain the first sample audio of the first person, the second sample audio of the first person, and the third sample audio of the second person.
[0118] Among them, the sample audio can be a Mel spectrum, or spectral data, or time-domain audio data. The embodiment of the present application takes the Mel spectrum as an example for illustration.
[0119] In implementation, before training the voice color feature extraction model, it is necessary to record the audio of different people. Then, perform short-time Fourier transform on the recorded audio respectively to obtain spectral data. Then convert the spectral data into Mel spectra as sample audio respectively. The sample audio can be grouped. Three sample audios can be divided into a group. The three sample audios in each group include two sample audios of the same person and one sample audio of a different person. For example, a group of three sample audios includes the first sample audio and the second sample audio of Zhang San, and also includes the third sample audio of Li Si.
[0120] 102. Extract the first timbre feature from the first sample audio according to the timbre feature extraction model to be trained, extract the second timbre feature from the second sample audio according to the timbre feature extraction model to be trained, and extract the third timbre feature from the third sample audio according to the timbre feature extraction model to be trained.
[0121] Among them, the first timbre feature, the second timbre feature, and the third timbre feature can all be in mathematical forms such as vectors or matrices. In the embodiments of the present application, vectors are used as examples for illustration.
[0122] The timbre feature extraction model can be a machine learning model. The network structure used in the timbre feature extraction model can be ECAPA-TDNN (emphasized channel attention propagation and aggregation in time delay neural network based speaker verification, a mainstream voiceprint model for speaker verification based on time delay neural network with emphasized channel attention, propagation, and aggregation).
[0123] In implementation, the computer device can input the first sample audio into the timbre feature extraction model to be trained, and then the timbre feature extraction model to be trained can output the predicted first timbre feature. And it can input the second sample audio into the timbre feature extraction model to be trained, and then the timbre feature extraction model to be trained can output the predicted second timbre feature. And it can input the third sample audio into the timbre feature extraction model to be trained, and then the timbre feature extraction model to be trained can output the predicted third timbre feature. When the timbre feature extraction model has not been trained yet, each of the above-obtained timbre features is only a predicted value and may not be accurate enough, which is used to train the timbre feature extraction model.
[0124] 103. Take minimizing the first difference between the first timbre feature and the second timbre feature and maximizing the second difference between the first timbre feature and the third timbre feature as the training objective, and train the timbre feature extraction model to be trained.
[0125] Since the first sample audio and the second sample audio are from the same speaker, the timbres of these two sample audios should be the same. Therefore, if the timbre feature extraction model is accurate enough, the first timbre feature and the second timbre feature should have extremely small differences. Since the first sample audio and the third sample audio are from different speakers, the timbres of these two sample audios should be quite different. Therefore, if the timbre feature extraction model is accurate enough, the first timbre feature and the third timbre feature should have large differences. Therefore, with the training objective of minimizing the first difference between the first timbre feature and the second timbre feature and maximizing the second difference between the first timbre feature and the third timbre feature, training the timbre feature extraction model to be trained can enable the timbre feature extraction model to extract more accurate timbre features.
[0126] In addition, since the second sample audio and the third sample audio are also from different speakers, the second timbre feature and the third timbre feature should have large differences. Therefore, optionally, when training the timbre feature extraction model to be trained, in addition to the above two training objectives, the training objective of maximizing the difference between the second timbre feature and the third timbre feature can also be adopted simultaneously.
[0127] Due to the different environments when people record audio, different types of noise, etc. may also be recorded in the audio. Or, for the same person, in different situations, the recorded audio may vary due to the person's emotional fluctuations. Information such as noise, emotions, and different speech contents may be included when extracting the Mel spectrogram from the audio, affecting the extraction of timbre features in the subsequent process. Therefore, when extracting timbre features, a positive sample (the second sample audio) is needed for contrast learning and a negative sample (the third sample audio) is needed for adversarial training.
[0128] Since the first audio sample and the second audio sample are Mel spectrograms of different contents spoken by the same person, during the training process, if the first timbre feature and the second timbre feature carry information other than timbre such as noise and emotions, then the difference between the first timbre feature and the second timbre feature will increase. At this time, training with the training objective of minimizing the first difference between the first timbre feature and the second timbre feature can make the information such as noise and emotions carried in the timbre feature less and less.
[0129] During the training process, the above differences can be obtained by calculating the loss value. The flowchart for calculating the first loss value can be as Figure 2 shown, and the specific processing method can be:
[0130] Step 1: Determine the first vector angle between the first timbre feature vector and the second timbre feature vector, determine the second vector angle between the first timbre feature vector and the third timbre feature vector, and determine the third vector angle between the second timbre feature vector and the third timbre feature vector.
[0131] In implementation, timbre features are represented by vectors, and the differences between different timbre features can be represented by the angles between the vectors. When the angle between two vectors is extremely small, it indicates that the differences between the timbre features represented by the two vectors are small. When the angle between two vectors is large, it indicates that the differences between the timbre features represented by the two vectors are large.
[0132] Step 2: Determine the first loss value according to the first vector angle, the second vector angle, and the third vector angle.
[0133] Among them, the first loss value is negatively correlated with the first vector angle, negatively correlated with the second vector angle, and negatively correlated with the third vector angle.
[0134] The following gives two specific processing methods for Step 2:
[0135] Method 1: Trigonometric functions can be selected as the basic functions for constructing the loss function. In this embodiment of the application, the cosine function is taken as an example for illustration. Correspondingly, the specific processing method of Step 2 can be:
[0136] Determine the first cosine value of the first vector angle, the second cosine value of the second vector angle, and the third cosine value of the third vector angle. Determine the first sub-loss value according to the first cosine value and the second cosine value; determine the second sub-loss value according to the second cosine value; determine the third sub-loss value according to the third cosine value; determine the fourth sub-loss value according to the first cosine value. Determine the first loss value according to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value.
[0137] Among them, the first sub-loss value is negatively correlated with the first cosine value and positively correlated with the second cosine value, the second sub-loss value is positively correlated with the second cosine value, the third sub-loss value is positively correlated with the third cosine value, and the fourth sub-loss value is negatively correlated with the first cosine value.
[0138] For a clearer representation, the following gives the function expressions corresponding to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value, as follows:
[0139]
[0140] L 2 =cos(y a ,y n )
[0141] L 3 =cos(y p,y n )
[0142] L 4 =-cos(y a ,y p )
[0143] where y a is the first timbre feature vector, y p is the second timbre feature vector, y n is the third timbre feature vector, cos(y a ,y p ) is the first cosine value, cos(y a ,y n ) is the second cosine value, cos(y p ,y n ) is the third cosine value. The above functions are the contrastive learning loss function (L 1 ), the discriminator loss function in adversarial training (L 2 and L 3 ), and the synthesizer loss function (L 4 ).
[0144] For the four sub-loss values, in addition to using the above function expressions, other function expressions can also be used. Another function expression is given below:
[0145]
[0146] L 2 =cos 2 (y a ,y n )
[0147] L 3 =cos 3 (y p ,y n )
[0148] L 4 =log cos(y a ,y p )
[0149] where k and b are constants.
[0150] After determining each sub-loss value, the weighted average of the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value can be calculated as the first loss value. Among them, the weight values of the first sub-loss value, the second sub-loss value, and the third sub-loss value can be slightly larger, and the fourth sub-loss value can be slightly smaller. For example, the first loss value L = 0.3L 1 +0.3L 2+0.3L 3 +0.1L 4 。
[0151] Method 2: A linear function can also be selected as the basis function of the loss function. For example:
[0152]
[0153]
[0154]
[0155] where α 1 is the first vector angle, α 1 is the second vector angle, α 3 is the third vector angle; Μ 1 is the first sub-loss value, Μ 2 is the second sub-loss value, Μ 3 is the third sub-loss value. After determining each sub-loss value, a weighted average of each sub-loss value is taken to obtain the first loss value.
[0156] Step 3: Train the tone feature extraction model to be trained according to the first loss value.
[0157] 104. If the training end condition is satisfied, the tone feature extraction model that satisfies the training end condition is determined as the target tone feature extraction model.
[0158] If the training end condition is not satisfied, a new set of sample audio can be obtained and the above process can be repeated.
[0159] There are many choices for the training end condition. The following are several examples:
[0160] Condition 1: Reaching the specified number of training times. Condition 2: The loss value is less than the specified value. Condition 3: The loss value no longer shows a decreasing trend. Condition 4: Using the features extracted by the tone feature extraction model for audio synthesis experiments, comparing the synthesized audio with the audio used to extract the tone features, and the matching degree reaches the specified value, or the matching degree of consecutive multiple experiments reaches the specified value.
[0161] For each sample audio in the above step 101, optionally, in order to improve the accuracy of the tone feature extraction model in extracting tone features, each sample audio can meet the following conditions: The first sample audio and the third sample audio are the reading audio corresponding to the same text, and the first sample audio and the second sample audio are the reading audio corresponding to different texts.
[0162] There are significant differences in the timbre characteristics between the first sample audio and the third sample audio, while the differences in text information are relatively small. There are relatively small differences in the timbre characteristics between the first sample audio and the second sample audio, while the differences in text information are significant. Training with the aim of minimizing the first difference between the first timbre characteristic and the second timbre characteristic and maximizing the second difference between the first timbre characteristic and the third timbre characteristic can reduce the text information carried in the timbre characteristics, thereby improving the accuracy of the timbre characteristic extraction model in extracting timbre characteristics.
[0163] In addition to being trained independently, the timbre characteristic extraction model can also be trained in combination with an audio synthesis model. The processing flow of the corresponding training method is as Figure 3 shown, including the following processing steps:
[0164] 301. Determine the phoneme sequence corresponding to the sample text, and input the phoneme sequence into the encoder to be trained to obtain the text pronunciation feature corresponding to the sample text.
[0165] Among them, the sample text can be the same as or different from the speech content corresponding to the first audio sample. The text pronunciation feature can also be called a phoneme vector or text encoding. The encoder is a machine learning model, which can specifically adopt a linear regression model, a logistic regression model, a neural network model, etc.
[0166] In implementation, when performing speech synthesis, it is necessary to input the text content and determine the phoneme sequence corresponding to the text content. For example, the phoneme sequence corresponding to "today" is "j, i, n, t, i, a, n". After inputting the phoneme sequence into the encoder, the encoder can output the phoneme vector corresponding to the phoneme sequence, that is, the text pronunciation feature corresponding to the sample text.
[0167] 302. Input the text pronunciation feature and the first timbre characteristic into the audio synthesis model to be trained to obtain the predicted audio.
[0168] Among them, the first timbre characteristic is obtained by extracting the timbre characteristic of the first sample audio in the above process. The predicted audio can be a Mel spectrogram or time-domain audio data.
[0169] 303. Determine the second loss value according to the predicted audio and the first sample audio.
[0170] The flowchart for calculating the second loss value can be as Figure 4 shown.
[0171] Among them, the second loss value can be the difference between the predicted audio and the first sample audio.
[0172] 304. With the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and based on the second loss value, train the timbre feature extraction model to be trained (this process can be considered a refinement of step 103); according to the second loss value, train the encoder to be trained and the audio synthesis model to be trained.
[0173] Among them, the second timbre feature is obtained by extracting the timbre feature of the second sample audio in the above process, and the third timbre feature is obtained by extracting the timbre feature of the third sample audio in the above process.
[0174] In implementation, the smaller the difference between the predicted audio and the first sample audio, the higher the accuracy of the predicted audio. Because the accuracy of the predicted audio is jointly affected by the three models: the timbre feature extraction model to be trained, the encoder to be trained, and the audio synthesis model to be trained. Therefore, according to the second loss value, with the aim of minimizing the difference between the predicted audio and the first sample audio, train the encoder to be trained and the audio synthesis model to be trained. And in addition to tuning the parameters of the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, it is also necessary to retune the parameters of the timbre feature extraction model to be trained according to the second loss value. In step 304, it can be considered that the parameters of the timbre feature extraction model to be trained are tuned twice, and the parameters of the encoder to be trained and the audio synthesis model to be trained are tuned once respectively.
[0175] 305. If the training end condition is met, in addition to performing the processing of step 104, the encoder that meets the training end condition can be determined as the target encoder, and the audio synthesis model that meets the training end condition can be determined as the target audio synthesis model.
[0176] If the training end condition is not met, a new set of sample audio can be obtained, and the above process can be repeated.
[0177] There are many choices for the training end condition. The following are several examples:
[0178] Condition 1: Reach the specified number of training times. Condition 2: The loss value is less than the specified value. Condition 3: The loss value no longer shows a decreasing trend. Condition 4: Use the encoder to extract the text pronunciation feature, and use the feature extracted by the timbre feature extraction model for audio synthesis experiment. Compare the synthesized audio with the audio used for extracting the timbre feature, and the matching degree reaches the specified value, or the matching degree of consecutive multiple experiments reaches the specified value.
[0179] The timbre feature extraction model can also be used as a timbre supervision model. After the audio synthesis model outputs the predicted audio, before determining the timbre feature extraction model that meets the training end condition as the target timbre feature extraction model, determining the encoder that meets the training end condition as the target encoder, and determining the audio synthesis model that meets the training end condition as the target audio synthesis model, the predicted audio and other sample audios can be input into the timbre feature extraction model again for training. The corresponding training method process can be as Figure 5 shown, including the following processing steps:
[0180] 501. Obtain the fourth sample audio of the first person and the fifth sample audio of the second person.
[0181] Among them, the fourth sample audio and the fifth sample audio can be Mel spectrograms, can also be spectral data, or can also be time-domain audio data.
[0182] The specific processing of step 501 is similar to that of step 101. For the relevant description content of step 101, please refer to it. It will not be elaborated here.
[0183] 502. Extract the fourth timbre feature from the predicted audio according to the timbre feature extraction model to be trained, extract the fifth timbre feature from the fourth sample audio according to the timbre feature extraction model to be trained, and extract the sixth timbre feature from the fifth sample audio according to the timbre feature extraction model to be trained.
[0184] Among them, the fourth timbre feature, the fifth timbre feature, and the sixth timbre feature can all be in mathematical forms such as vectors or matrices. The timbre feature extraction model to be trained can be a model that has not been trained in step 103, or can also be a model that has been trained in step 103.
[0185] The specific processing of step 502 is similar to that of step 102. For the relevant description content of step 102, please refer to it. It will not be elaborated here.
[0186] 503. With the purpose of minimizing the difference between the fourth timbre feature and the fifth timbre feature and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and according to the second loss value, train the encoder to be trained and the audio synthesis model to be trained; with the purpose of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, minimizing the difference between the fourth timbre feature and the fifth timbre feature, and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and according to the second loss value, train the timbre feature extraction model to be trained. (This step can be considered a refinement of step 304)
[0187] Among them, the first timbre feature is obtained by extracting the timbre feature of the first sample audio in the above process, the second timbre feature is obtained by extracting the timbre feature of the second sample audio in the above process, and the third timbre feature is obtained by extracting the timbre feature of the third sample audio in the above process.
[0188] First, the parameter adjustment values of each model can be determined in three parts.
[0189] In the first part, according to the first loss value (aiming to minimize the difference between the first timbre feature and the second timbre feature and maximize the difference between the first timbre feature and the third timbre feature), the parameter adjustment value of the timbre feature extraction model to be trained is determined.
[0190] In the second part, according to the second loss value, the parameter adjustment values of the timbre feature extraction model to be trained, the encoder to be trained, and the audio synthesis model to be trained are determined.
[0191] In the third part, calculate the third loss value. According to the third loss value (aiming to minimize the difference between the fourth timbre feature and the fifth timbre feature and maximize the difference between the fourth timbre feature and the sixth timbre feature for training), the parameter adjustment values of the timbre feature extraction model to be trained, the encoder to be trained, and the audio synthesis model to be trained are determined.
[0192] Then, according to all the determined parameter adjustment values, the parameters of the timbre feature extraction model to be trained, the encoder to be trained, and the audio synthesis model to be trained are adjusted.
[0193] In step 503, it can be considered that the timbre feature extraction model to be trained is adjusted three times, and the encoder to be trained and the audio synthesis model to be trained are adjusted twice respectively.
[0194] Among them, the first loss value and the second loss value have been given in the above process.
[0195] The specific processing method for calculating the third loss value can be:
[0196] Step 1, determine the fourth vector angle between the fourth timbre feature and the fifth timbre feature, determine the fifth vector angle between the fourth timbre feature and the sixth timbre feature, and determine the sixth vector angle between the fifth timbre feature and the sixth timbre feature.
[0197] Step 2, according to the fourth vector angle, the fifth vector angle, and the sixth vector angle, determine the third loss value.
[0198] The calculation of the third loss value is similar to the calculation method of the first loss value. For relevant description content, refer to step 103, which will not be elaborated here.
[0199] The flowchart for calculating the third loss value can be asFigure 6 as shown
[0200] After the tone feature extraction model, audio synthesis model, encoder, etc. are trained, the tone feature extraction model can be used alone. The server can obtain a large number of audios of different people, extract the tone features of each person's audio respectively, and the extracted tone features can be saved in the tone library of the server or terminal. When synthesizing audio, the tone features can be directly called. For example, in intelligent navigation, the user can select a certain tone in the tone library, and during subsequent navigation, the intelligent navigation can synthesize the real-time generated prompt text into the navigation audio of the selected tone. The tone feature extraction model can also be used simultaneously with the audio synthesis model and encoder. This usage method is usually the case where the tone features of a new person need to be added. For example, in a related application, the user can input a text content and upload an audio. After the terminal obtains the audio and content, it extracts the tone features through the tone feature extraction model, extracts the text pronunciation features through the encoder, then synthesizes the audio of a specific tone and specific text content through the audio synthesis model, and the extracted tone features can be saved in the tone library.
[0201] The embodiment of this application provides a method for audio synthesis for the above application scenarios. The processing flow of this method can be as Figure 7 shown, including the following processing steps:
[0202] 701. Extract the tone features of the target audio using the target tone feature extraction model.
[0203] 702. Determine the target phoneme sequence corresponding to the target text, and input the target phoneme sequence into the target encoder to obtain the target text pronunciation features corresponding to the target text.
[0204] 703. Input the target text pronunciation features and the tone features of the target audio into the target audio synthesis model to obtain the synthesized audio.
[0205] Among them, the target audio can be the audio selected by the user that they want to use the tone in. The synthesized audio is the synthesized reading audio of the target text with the tone of the target audio.
[0206] In implementation, when the user needs to synthesize a reading audio of the target text with the tone features corresponding to the target audio, they can input the target text and target audio on the interface corresponding to the application of the voice synthesis assistant class. The application determines the target factor sequence corresponding to the target text, extracts the target text pronunciation features from the target factor sequence through the target encoder, the target tone feature extraction model extracts the tone features of the target audio, and synthesizes the target text pronunciation features and the tone features of the target audio through the target audio synthesis model to obtain the synthesized audio.
[0207] Through the method provided by the embodiments of the present application, it is necessary to obtain the first sample audio and the second sample audio of the first person, as well as the third sample audio of the second person, and extract the timbre features of these three sample audios respectively through a timbre feature extraction model. The extracted audio features are the first timbre feature, the second timbre feature, and the third timbre feature respectively. The timbre feature extraction model is trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature. After the condition for ending the training is reached, the timbre feature extraction model that meets the training end condition can be determined as the target timbre feature extraction model. In this way, the target timbre feature extraction model trained by this method can directly input the audio with this timbre into the timbre feature extraction model to obtain the timbre feature without additional training of the model when extracting the timbre feature of any audio, thereby improving the efficiency of timbre feature extraction.
[0208] In addition, maximizing the difference between the first timbre feature and the third timbre feature can well improve the proportion of individual features in the timbre features extracted by the timbre feature extraction model, and can well suppress the proportion of collective features in the timbre features, so that the timbre features can better reflect the different characteristics of the voices of different people.
[0209] Based on the same technical concept, the embodiments of the present application also provide a training device for a timbre feature extraction model. This device can be applied to the terminal or server in the above embodiments, such as Figure 8 As shown, the device includes:
[0210] An acquisition module 810, configured to acquire the first sample audio of the first person, the second sample audio of the first person, and the third sample audio of the second person;
[0211] An extraction module 820, configured to extract the first timbre feature from the first sample audio according to the timbre feature extraction model to be trained, extract the second timbre feature from the second sample audio according to the timbre feature extraction model to be trained, and extract the third timbre feature from the third sample audio according to the timbre feature extraction model to be trained;
[0212] A training module 830, configured to train the timbre feature extraction model to be trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature; if the training end condition is met, determine the timbre feature extraction model that meets the training end condition as the target timbre feature extraction model.
[0213] In a possible implementation manner, the training module 830 is configured to:
[0214] Training the tone feature extraction model to be trained with the aim of minimizing the difference between the first tone feature and the second tone feature, maximizing the difference between the first tone feature and the third tone feature, and maximizing the difference between the second tone feature and the third tone feature.
[0215] In a possible implementation, the first tone feature is a first tone feature vector, the second tone feature is a second tone feature vector, and the third tone feature is a third tone feature vector;
[0216] The training module 830 is used for:
[0217] Determining a first vector angle between the first tone feature vector and the second tone feature vector, determining a second vector angle between the first tone feature vector and the third tone feature vector, and determining a third vector angle between the second tone feature vector and the third tone feature vector;
[0218] Determining a first loss value according to the first vector angle, the second vector angle, and the third vector angle, where the first loss value is positively correlated with the first vector angle, negatively correlated with the second vector angle, and negatively correlated with the third vector angle;
[0219] Training the tone feature extraction model to be trained according to the first loss value.
[0220] In a possible implementation, the training module 830 is used for:
[0221] Determining a first cosine value of the first vector angle, a second cosine value of the second vector angle, and a third cosine value of the third vector angle;
[0222] Determining a first loss value according to the first cosine value, the second cosine value, and the third cosine value, where the first loss value is negatively correlated with the first cosine value, positively correlated with the second cosine value, and positively correlated with the third cosine value.
[0223] In a possible implementation, the training module 830 is used for:
[0224] Determining a first sub-loss value according to the first cosine value and the second cosine value, determining a second sub-loss value according to the second cosine value, determining a third sub-loss value according to the third cosine value, and determining a fourth sub-loss value according to the first cosine value, where the first sub-loss value is negatively correlated with the first cosine value and positively correlated with the second cosine value, the second sub-loss value is positively correlated with the second cosine value, the third sub-loss value is positively correlated with the third cosine value, and the fourth sub-loss value is negatively correlated with the first cosine value;
[0225] Determining a first loss value according to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value.
[0226] According to the formula Determine the first sub-loss value L 1 ;
[0227] According to the formula L 2 = cos(y a , y n ), determine the second sub-loss value L 2 ;
[0228] According to the formula L 3 = cos(y p , y n ), determine the third sub-loss value L 3 ;
[0229] According to the formula L 4 = -cos(y a , y p ), determine the fourth sub-loss value L 4 ;
[0230] where y a is the first timbre feature vector, y p is the second timbre feature vector, y n is the third timbre feature vector, cos(y a , y p ) is the first cosine value, cos(y a , y n ) is the second cosine value, cos(y p , y n ) is the third cosine value.
[0231] In a possible implementation, the parameter tuning module 830 is used to:
[0232] Determine the weighted average of the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value according to the first weight, the second weight, the third weight, and the fourth weight, and obtain the first loss value.
[0233] In a possible implementation, the first sample audio and the third sample audio are the reading audios corresponding to the same text, and the first sample audio and the second sample audio are the reading audios corresponding to different texts.
[0234] In a possible implementation, the device further includes a synthesis module;
[0235] The synthesis module is used to determine the phoneme sequence corresponding to the sample text, input the phoneme sequence into the encoder to be trained, and obtain the text pronunciation feature corresponding to the sample text; input the text pronunciation feature and the first timbre feature into the audio synthesis model to be trained, and obtain the predicted audio;
[0236] The training module 830 is further configured to determine a second loss value according to the predicted audio and the first sample audio; and train the encoder to be trained and the audio synthesis model to be trained according to the second loss value; with the purpose of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and train the timbre feature extraction model to be trained according to the second loss value; if the training end condition is satisfied, determine the encoder that satisfies the training end condition as the target encoder, and determine the audio synthesis model that satisfies the training end condition as the target audio synthesis model.
[0237] In a possible implementation manner, the obtaining module 810 is further configured to obtain a fourth sample audio of a first person and a fifth sample audio of a second person;
[0238] The extraction module 820 is further configured to extract a fourth timbre feature from the predicted audio according to the timbre feature extraction model to be trained, extract a fifth timbre feature from the fourth sample audio according to the timbre feature extraction model to be trained, and extract a sixth timbre feature from the fifth sample audio according to the timbre feature extraction model to be trained;
[0239] The training module 830 is further configured to use as the training purpose minimizing the difference between the fourth timbre feature and the fifth timbre feature and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and train the encoder to be trained and the audio synthesis model to be trained according to the second loss value; with the purpose of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, minimizing the difference between the fourth timbre feature and the fifth timbre feature, and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and train the timbre feature extraction model to be trained according to the second loss value.
[0240] Based on the same technical concept, an embodiment of the present application provides an audio synthesis device, which can be applied to the terminal or server in the above embodiments, such as Figure 9 As shown, the device includes:
[0241] An extraction module 910, configured to extract a timbre feature from a target audio;
[0242] A determination module 920, configured to determine a target phoneme sequence corresponding to a target text, input the target phoneme sequence into a target encoder, and obtain a target text pronunciation feature corresponding to the target text;
[0243] A synthesis module 930, configured to input the target text pronunciation feature and the timbre feature of the target audio into a target audio synthesis model to obtain a synthesized audio.
[0244] Through the device provided by the embodiments of the present application, it is necessary to obtain the first sample audio and the second sample audio of the first person, as well as the third sample audio of the second person, and extract the timbre features of these three sample audios respectively through the timbre feature extraction model. The extracted audio features are the first timbre feature, the second timbre feature, and the third timbre feature respectively. The timbre feature extraction model is trained with the aim of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature. After the condition for ending the training is reached, the timbre feature extraction model that meets the training end condition can be determined as the target timbre feature extraction model. In this way, the target timbre feature extraction model trained by this method can directly input the audio with a certain timbre into the timbre feature extraction model to obtain the timbre feature without additional training when extracting the timbre feature of the audio with any timbre, thereby improving the efficiency of timbre feature extraction.
[0245] It should be noted that when the training device of the timbre feature extraction model provided in the above embodiments extracts the timbre feature, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the training device of the timbre feature extraction model provided in the above embodiments and the embodiments of the training method of the timbre feature extraction model belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0246] Figure 10 The structural block diagram of the electronic device 1000 provided by the embodiments of the present application is shown. This electronic device can be the computer device in the above embodiments. The electronic device 1000 can be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (moving picture experts group audio layer III), an MP4 (moving picture experts group audio layer IV) player, a notebook computer or a desktop computer. The electronic device 1000 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0247] Generally, the electronic device 1000 includes a processor 1001 and a memory 1002.
[0248] The processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1001 may be implemented in at least one of the following hardware forms: DSP (digital signal processing), FPGA (field-programmable gate array), and PLA (programmable logic array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (central processing unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (graphics processing unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may further include an AI (artificial intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0249] The memory 1002 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1001 to implement the method provided in the embodiments of the present application.
[0250] In some embodiments, the electronic device 1000 may further optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1003 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, a positioning component 1008, and a power supply 1009.
[0251] The peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0252] The radio frequency circuit 1004 is used to receive and transmit RF (radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1004 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (wireless fidelity) network. In some embodiments, the radio frequency circuit 1004 may further include a circuit related to NFC (near field communication), and this application does not limit this.
[0253] The display screen 1005 is used to display the UI (user interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 also has the ability to collect touch signals on or above the surface of the display screen 1005. The touch signals can be input to the processor 1001 as control signals for processing. At this time, the display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, which is provided on the front panel of the electronic device 1000; in other embodiments, there may be at least two display screens 1005, which are respectively provided on different surfaces of the electronic device 1000 or are in a foldable design; in other embodiments, the display screen 1005 may be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 1000. Even further, the display screen 1005 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1005 can be prepared using materials such as LCD (liquid crystal display) and OLED (organic light-emitting diode).
[0254] The camera module 1006 is used to capture images or videos. Optionally, the camera module 1006 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, to achieve functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (virtual reality) shooting functions or other combined shooting functions. In some embodiments, the camera module 1006 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0255] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1001 for processing, or input to the radio frequency circuit 1004 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 1000. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1007 may also include a headphone jack.
[0256] The positioning component 1008 is used to locate the current geographical location of the electronic device 1000 to achieve navigation or LBS (location based service). The positioning component 1008 may be a positioning component based on GPS (global positioning system), Beidou system or Galileo system.
[0257] The power supply 1009 is used to supply power to each component in the electronic device 1000. The power supply 1009 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1009 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0258] In some embodiments, the electronic device 1000 further includes one or more sensors 1010. The one or more sensors 1010 include but are not limited to: an acceleration sensor 1011, a gyroscope sensor 1012, a pressure sensor 1013, a fingerprint sensor 1014, an optical sensor 1015 and a proximity sensor 1016.
[0259] The acceleration sensor 1011 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 1000. For example, the acceleration sensor 1011 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1001 can control the display screen 1005 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1011. The acceleration sensor 1011 can also be used for games or the collection of the user's motion data.
[0260] The gyroscope sensor 1012 can detect the body direction and rotation angle of the electronic device 1000. The gyroscope sensor 1012 can cooperate with the acceleration sensor 1011 to collect the 3D actions of the user on the electronic device 1000. Based on the data collected by the gyroscope sensor 1012, the processor 1001 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0261] The pressure sensor 1013 can be disposed on the side frame of the electronic device 1000 and / or the lower layer of the display screen 1005. When the pressure sensor 1013 is disposed on the side frame of the electronic device 1000, it can detect the holding signal of the user on the electronic device 1000, and the processor 1001 can identify the left or right hand or perform a quick operation according to the holding signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1005. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0262] The fingerprint sensor 1014 is used to collect the fingerprint of the user. The processor 1001 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 1014, or the fingerprint sensor 1014 can identify the user's identity according to the collected fingerprint. When the identified user identity is a trusted identity, the processor 1001 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 1014 can be disposed on the front, back, or side of the electronic device 1000. When there are physical buttons or a manufacturer logo on the electronic device 1000, the fingerprint sensor 1014 can be integrated with the physical button or the manufacturer logo.
[0263] The optical sensor 1015 is used to collect the ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 according to the ambient light intensity collected by the optical sensor 1015. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera module 1006 according to the ambient light intensity collected by the optical sensor 1015.
[0264] The proximity sensor 1016, also known as the distance sensor, is usually disposed on the front panel of the electronic device 1000. The proximity sensor 1016 is used to collect the distance between the user and the front of the electronic device 1000. In one embodiment, when the proximity sensor 1016 detects that the distance between the user and the front of the electronic device 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from the lit state to the off state; when the proximity sensor 1016 detects that the distance between the user and the front of the electronic device 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from the off state to the lit state.
[0265] Those skilled in the art can understand that Figure 10 the structure shown in does not constitute a limitation on the electronic device 1000, and it may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0266] In an embodiment of the present application, there is also provided a computer-readable storage medium, such as a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the method of performing an interactive operation in the above embodiment. The computer-readable storage medium may be non-transitory. For example, the computer-readable storage medium may be a ROM (read-only memory), a RAM (random access memory), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0267] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between the user terminal and other devices, etc.) involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0268] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0269] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for a timbre feature extraction model, characterized in that, the method includes: Obtain the first sample audio of the first person, the second sample audio of the first person, and the third sample audio of the second person. The first sample audio and the third sample audio are reading audios corresponding to the same text, and the first sample audio and the second sample audio are reading audios corresponding to different texts; Extract the first timbre feature from the first sample audio according to the timbre feature extraction model to be trained, extract the second timbre feature from the second sample audio according to the timbre feature extraction model to be trained, and extract the third timbre feature from the third sample audio according to the timbre feature extraction model to be trained; Taking minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature as the training objective, train the timbre feature extraction model to be trained; if the training end condition is satisfied, determine the timbre feature extraction model that satisfies the training end condition as the target timbre feature extraction model; wherein, the training end condition is condition one, condition two or condition three. Condition one is reaching the specified number of training times. Condition two is that the loss value calculated using the timbre feature extraction model obtained by training is less than the specified value. Condition three is to conduct an experiment on audio synthesis using the timbre features extracted by the timbre feature extraction model obtained by training to obtain the synthesized audio, compare the synthesized audio with the audio used for extracting the timbre features, and the matching degree reaches the specified value, or the matching degree of consecutive multiple experiments reaches the specified value.
2. The method according to claim 1, characterized in that, the training of the timbre feature extraction model to be trained with the objective of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature includes: Training the timbre feature extraction model to be trained with the objective of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, and maximizing the difference between the second timbre feature and the third timbre feature.
3. The method according to claim 2, characterized in that, the first timbre feature is a first timbre feature vector, the second timbre feature is a second timbre feature vector, and the third timbre feature is a third timbre feature vector; the training of the timbre feature extraction model to be trained with the objective of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, and maximizing the difference between the second timbre feature and the third timbre feature includes: Determine the first vector angle between the first timbre feature vector and the second timbre feature vector, determine the second vector angle between the first timbre feature vector and the third timbre feature vector, and determine the third vector angle between the second timbre feature vector and the third timbre feature vector; Determine a first loss value according to the first vector angle, the second vector angle, and the third vector angle, where the first loss value is positively correlated with the first vector angle, negatively correlated with the second vector angle, and negatively correlated with the third vector angle; Train the to-be-trained timbre feature extraction model according to the first loss value.
4. The method according to claim 3, wherein, the determining the first loss value according to the first vector angle, the second vector angle, and the third vector angle includes: Determine a first cosine value of the first vector angle, a second cosine value of the second vector angle, and a third cosine value of the third vector angle; Determine the first loss value according to the first cosine value, the second cosine value, and the third cosine value, where the first loss value is negatively correlated with the first cosine value, positively correlated with the second cosine value, and positively correlated with the third cosine value.
5. The method according to claim 4, wherein, the determining the first loss value according to the first cosine value, the second cosine value, and the third cosine value includes: Determine a first sub-loss value according to the first cosine value and the second cosine value; determine a second sub-loss value according to the second cosine value; determine a third sub-loss value according to the third cosine value; determine a fourth sub-loss value according to the first cosine value; wherein, the first sub-loss value is negatively correlated with the first cosine value and positively correlated with the second cosine value, the second sub-loss value is positively correlated with the second cosine value, the third sub-loss value is positively correlated with the third cosine value, and the fourth sub-loss value is negatively correlated with the first cosine value; Determine the first loss value according to the first sub-loss value, the second sub-loss value, the third sub-loss value, and the fourth sub-loss value.
6. The method according to claim 5, wherein, the determining the first sub-loss value according to the first cosine value and the second cosine value; determining the second sub-loss value according to the second cosine value; determining the third sub-loss value according to the third cosine value; determining the fourth sub-loss value according to the first cosine value includes: According to the formula determine the first sub-loss value L 1 ; According to the formula L 2 = cos(y a , y n ), determine the second sub-loss value L 2 ; According to the formula L 3 = cos(y p , y n ), determine the third sub-loss value L 3 ; According to the formula L 4 = -cos(y a , y p ), determine the fourth sub-loss value L 4 ; Among them, y a is the first timbre feature vector, y p is the second timbre feature vector, y n is the third timbre feature vector, cos(y a , y p ) is the first cosine value, cos(y a , y n ) is the second cosine value, cos(y p , y n ) is the third cosine value.
7. The method according to any one of claims 1-6, wherein, before the determining that if the training end condition is satisfied, the timbre feature extraction model that satisfies the training end condition is determined as the target timbre feature extraction model, the method further includes: Determine the phoneme sequence corresponding to the sample text, input the phoneme sequence into the to-be-trained encoder to obtain the text pronunciation feature corresponding to the sample text; and input the text pronunciation feature and the first timbre feature into the to-be-trained audio synthesis model to obtain a predicted audio; Determine a second loss value according to the predicted audio and the first sample audio; and train the to-be-trained encoder and the to-be-trained audio synthesis model according to the second loss value. Training the to-be-trained timbre feature extraction model for the purpose of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature includes: Training the to-be-trained timbre feature extraction model for the purpose of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and training the to-be-trained timbre feature extraction model according to the second loss value; If the training end condition is satisfied, the method further includes: Determining the encoder that satisfies the training end condition as the target encoder, and determining the audio synthesis model that satisfies the training end condition as the target audio synthesis model.
8. The method according to claim 7, wherein, After inputting the text pronunciation feature and the first timbre feature into the to-be-trained audio synthesis model to obtain a predicted audio, the method further includes: Obtaining a fourth sample audio of the first person and a fifth sample audio of the second person; Extracting a fourth timbre feature from the predicted audio according to the to-be-trained timbre feature extraction model, extracting a fifth timbre feature from the fourth sample audio according to the to-be-trained timbre feature extraction model, and extracting a sixth timbre feature from the fifth sample audio according to the to-be-trained timbre feature extraction model; The training of the to-be-trained encoder and the to-be-trained audio synthesis model according to the second loss value includes: Training the to-be-trained encoder and the to-be-trained audio synthesis model for the purpose of minimizing the difference between the fourth timbre feature and the fifth timbre feature and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and training the to-be-trained encoder and the to-be-trained audio synthesis model according to the second loss value; The training of the to-be-trained timbre feature extraction model for the purpose of minimizing the difference between the first timbre feature and the second timbre feature and maximizing the difference between the first timbre feature and the third timbre feature, and training the to-be-trained timbre feature extraction model according to the second loss value includes: Training the to-be-trained timbre feature extraction model for the purpose of minimizing the difference between the first timbre feature and the second timbre feature, maximizing the difference between the first timbre feature and the third timbre feature, minimizing the difference between the fourth timbre feature and the fifth timbre feature, and maximizing the difference between the fourth timbre feature and the sixth timbre feature, and training the to-be-trained timbre feature extraction model according to the second loss value.
9. A method for audio synthesis, wherein, The method further includes: Extracting timbre features from the target audio according to the target timbre feature extraction model according to any one of claims 1-8; Determining a target phoneme sequence corresponding to the target text, inputting the target phoneme sequence into the target encoder to obtain a target text pronunciation feature corresponding to the target text; Inputting the target text pronunciation feature and the timbre feature of the target audio into the target audio synthesis model to obtain a synthesized audio.
10. A computer device, wherein, The computer device includes a memory and a processor, and the memory is used for storing computer instructions; The processor executes the computer instructions stored in the memory, so that the computer device executes the method described in any one of claims 1-9 above.
11. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer program code, and when the computer program code is executed by a computer device, the computer device executes the method described in any one of claims 1-9 above.
Citation Information
Patent Citations
Timbre conversion method and device
CN112164407A
Timbre feature extraction method and device, computer equipment and storage medium
CN113870875A