Audio style vector training method and audio style vector generation method
By adopting the style coding layer and speaker classifier, combined with unsupervised and supervised speech synthesis model, a model that can accurately extract audio style vectors is trained, which solves the style vector deviation problem caused by the existing TTS model relying on speaker information, and improves the generalization ability and synthesis effect of the model.
Patent Information
- Application Number
- CN202411591238.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing TTS model relies on speaker information when extracting audio style vectors, resulting in deviations in style vectors extracted in unclear or incorrect situations, affecting the synthesis effect, and poor model flexibility and generalization ability.
By obtaining multiple sample audio and their corresponding text and features, a style encoding layer and a speaker classifier are used, combined with an unsupervised and supervised speech synthesis model, a model that can accurately extract audio style vectors is trained and integrated into the target speech synthesis model.
It improves the generalization ability and synthesis effect of the model, enhances the robustness and flexibility of the model, and can accurately extract audio style vectors without relying on the speaker's information, improving the accuracy of speech synthesis.
Smart Images

Figure CN119479614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech technology, and in particular to an audio style vector training method and an audio style vector generation method. Background Art
[0002] With the rapid development of artificial intelligence technology, text-to-speech (TTS) models have been widely used in human-computer interaction, virtual assistants, audiobooks, voiceprint recognition, speech classification and other fields. High-quality speech synthesis requires not only natural and fluent voices, but also the ability to accurately express the speaker's style. Currently, the autoregressive model in the TTS model is usually used, such as the Tacotron model. The reference encoder in the Tacotron model is used to extract the corresponding audio style vector according to the speaker.
[0003] However, the above-mentioned Tacotron model relies on speaker information when extracting audio style vectors. When the speaker information is unclear or incorrect, the extracted audio style vector may be biased, affecting the accuracy of the audio style vector. As a result, the synthesis effect of the model trained based on the audio style vector is poor, and the model has poor flexibility and generalization ability. Summary of the invention
[0004] In view of this, the present invention provides an audio style vector training method and an audio style vector generation method to solve the problems that the existing model has deviations when extracting audio style vectors, the synthesis effect of the model trained based on the audio style vector is poor, and the flexibility and generalization ability of the model are poor.
[0005] In a first aspect, the present invention provides an audio style vector training method, the method comprising:
[0006] Obtain multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features, where any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio;
[0007] For any sample audio, a style coding layer is used to determine a sample style vector of the sample audio based on the sample audio features of the sample audio;
[0008] Using a speaker classifier to classify the sample audio, and obtaining a sample classification result;
[0009] Using a first speech synthesis model and a second speech synthesis model, speech synthesis is performed based on a sample style vector and a sample audio feature of the sample audio to obtain a first synthesized audio and a second synthesized audio, respectively, the first speech synthesis model is an unsupervised model of a conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of a conditional variational autoencoding structure;
[0010] Determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on a sample classification result of the sample audio and a speaker label;
[0011] Based on the first synthesis loss, the second synthesis loss and the classification loss, the first speech synthesis model, the second speech synthesis model and the style coding layer are trained, and the trained style coding layer is integrated into the trained first speech synthesis model and the trained second speech synthesis model to obtain the first target speech synthesis model and the second target speech synthesis model.
[0012] The audio style vector training method provided by the embodiment of the present invention can cover a variety of different styles and speakers by using multiple sample audios, thereby improving the generalization ability of the model. The sample text provides additional contextual information, which helps the model to better understand the audio content and style. The style coding layer can extract style-related information from audio features without relying on speaker information. The speaker classifier is used for classification, which helps the model learn the characteristics of different speakers. The unsupervised model and the supervised model of the conditional variational automatic encoding structure are used for speech synthesis respectively to obtain the first synthesized audio and the second synthesized audio, providing two different synthesis paths, thereby enhancing the robustness and flexibility of the model. Without relying on speaker information, the style vector in the audio can be accurately extracted, so that training can be performed based on the style vector. By considering the synthesis loss and the classification loss at the same time, the model can be optimized from multiple angles. Finally, the trained style coding layer is respectively integrated into the two trained speech synthesis models, so that the two speech synthesis models can accurately extract the style vector in the audio, and can perform speech synthesis more accurately according to the style vector of the audio.
[0013] In an optional implementation, before acquiring a plurality of sample audios, sample texts corresponding to the plurality of sample audios, and sample audio features, the method further includes:
[0014] Obtain multiple initial audios and initial texts corresponding to the multiple initial audios;
[0015] For any initial audio, adjust the sampling rate of the initial audio to obtain a sample audio corresponding to the initial audio;
[0016] The initial text corresponding to the initial audio is regularized to obtain a sample text corresponding to the initial text.
[0017] The audio style vector training method provided in the embodiment of the present invention ensures the consistency of all input data by preprocessing the initial audio and initial text, reduces the errors caused by inconsistent data formats, and improves the input quality of the model through unified sampling rate and standardized text, thereby improving the overall performance and robustness of the model and improving the overall computing efficiency.
[0018] In an optional implementation, the sample audio features include phonemes, pitch features, orthograms, and mel-spectrograms, and obtaining the sample audio features corresponding to the plurality of sample audios includes:
[0019] Phonemize the sample text of the sample audio to obtain the phonemes of the sample text;
[0020] Extract pitch features, spectrograms, and mel-spectrograms of sample audio.
[0021] The audio style vector training method provided in the embodiment of the present invention helps to better understand the pronunciation details of the audio by converting the sample text into a corresponding phoneme sequence, and extracts the audio features of the sample audio to obtain features related to the audio style, which is helpful for subsequent speech synthesis.
[0022] In an optional implementation, using a first speech synthesis model to perform speech synthesis based on a sample style vector and a sample audio feature of a sample audio to obtain a first synthesized audio includes:
[0023] Using the pitch posterior coding layer of the first speech synthesis model, feature extraction is performed based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a first sample latent vector;
[0024] Using the Fourier coding layer of the first speech synthesis model, extracting features based on the sample style vector of the sample audio and the orthogram in the sample audio feature to obtain a second sample latent vector;
[0025] A vocoder of the first speech synthesis model is used to synthesize a first synthesized audio based on the sample style vector, the first sample latent vector and the second sample latent vector of the sample audio.
[0026] The audio style vector training method provided in the embodiment of the present invention better captures emotional information through the pitch posterior coding layer in the first speech synthesis model to make the synthesized audio more expressive, and then more finely captures the frequency domain characteristics through the Fourier coding layer, which helps the model learn subtle changes in the audio. Finally, the final audio is synthesized through a vocoder. By fusing pitch features and frequency domain features and introducing a style vector representing the audio style, it can generate more natural and coherent synthesized audio, ensuring that the style of the synthesized audio is consistent with the original audio, thereby enhancing the accuracy of speech synthesis.
[0027] In an optional implementation, using a second speech synthesis model to perform speech synthesis based on a sample style vector and a sample audio feature of the sample audio to obtain a second synthesized audio includes:
[0028] Using the pitch posterior coding layer of the second speech synthesis model, feature extraction is performed based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a third sample latent vector;
[0029] Using the Fourier coding layer of the second speech synthesis model, extracting features based on the sample style vector of the sample audio and the orthogram in the sample audio feature to obtain a fourth sample latent vector;
[0030] Using the stream coding layer of the second speech synthesis model, extracting features based on the third sample latent vector and the fourth sample latent vector to obtain a fifth sample latent vector;
[0031] Generate a sample text encoding based on the phonemes in the sample audio features of the sample audio using the sample text encoding layer and the protection layer of the second speech synthesis model;
[0032] Using a rhythmic one-way search algorithm, based on the fifth sample latent vector and sample text encoding, the sample rhythmic information is extracted;
[0033] The prediction layer of the second speech synthesis model is used to synthesize the second synthesized audio based on the sample style vector, the sample text encoding and the sample prosody information of the sample audio.
[0034] The audio style vector training method provided by the embodiment of the present invention adopts the pitch posterior coding layer of the second speech synthesis model to better capture emotional information and make the synthesized speech more expressive, and then uses the Fourier coding layer to more finely capture the frequency domain characteristics, which helps the model learn subtle changes in the audio, and then fuses the pitch features and frequency domain features through the stream coding layer, and then adopts the sample text coding layer and the protection layer, and uses phonemes for model supervision, which helps the model better understand the pronunciation details of the audio, and generates sample rhythm information by adopting the rhythm one-way search algorithm, which can ensure that the generated rhythm information is consistent with the audio features and text content, thereby enhancing the naturalness of the synthesized audio, and finally adopts the prediction layer to ensure that the style of the synthesized audio is consistent with the original audio by integrating the pitch features, frequency domain features, sample text coding and rhythm information and the style vector representing the speaker's style, thereby enhancing the accuracy of the synthesized audio.
[0035] In a second aspect, the present invention provides a method for generating an audio style vector, the method comprising:
[0036] Get speaker audio;
[0037] Determine whether the speaker's audio carries corresponding text information, where the text information is used to describe the style of the speaker's audio;
[0038] When the speaker's audio carries corresponding text information, a second target speech synthesis model is used to generate a style vector corresponding to the speaker's audio. The second target speech synthesis model is trained based on the audio style vector training method of the first aspect or any corresponding implementation method thereof.
[0039] The audio style vector generation method provided by the embodiment of the present invention obtains the speaker's audio and determines whether it carries corresponding text information, thereby selecting a corresponding speech synthesis model to extract the audio style vector, so that the audio style vector can be accurately extracted to meet the needs of various application scenarios.
[0040] In an optional implementation, after determining whether the speaker audio carries corresponding text information, the method further includes:
[0041] When the speaker's audio does not carry corresponding text information, a first target speech synthesis model is used to generate a style vector corresponding to the speaker's audio. The first target speech synthesis model is trained based on the audio style vector training method of the first aspect or any corresponding implementation method thereof.
[0042] The audio style vector generation method provided by the embodiment of the present invention can also accurately extract the style vector by adopting an unsupervised first target speech synthesis model when the speaker audio does not carry text information, so as to improve the generalization ability of obtaining the style vector.
[0043] In a third aspect, the present invention provides an audio style vector training device, the device comprising:
[0044] A first acquisition module is used to acquire multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features, wherein any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio;
[0045] A first determination module is used to determine, for any sample audio, a sample style vector of the sample audio based on the sample audio features of the sample audio by using a style coding layer;
[0046] A classification module is used to classify the sample audio using a speaker classifier to obtain a sample classification result;
[0047] A synthesis module, configured to use a first speech synthesis model and a second speech synthesis model to perform speech synthesis based on a sample style vector and a sample audio feature of the sample audio, to obtain a first synthesized audio and a second synthesized audio, respectively, wherein the first speech synthesis model is an unsupervised model of a conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of a conditional variational autoencoding structure;
[0048] A second determination module is used to determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on a sample classification result of the sample audio and a speaker label;
[0049] A training module is used to train the first speech synthesis model, the second speech synthesis model and the style coding layer based on the first synthesis loss, the second synthesis loss and the classification loss, integrate the trained style coding layer into the trained first speech synthesis model and the trained second speech synthesis model, and obtain the first target speech synthesis model and the second target speech synthesis model.
[0050] In a fourth aspect, the present invention provides an audio style vector generating device, the device comprising:
[0051] A second acquisition module is used to acquire speaker audio;
[0052] A judgment module, used to judge whether the speaker's audio carries corresponding text information;
[0053] The first generation module is used to use a second target speech synthesis model to generate a style vector corresponding to the speaker's audio when the speaker's audio carries corresponding text information. The second target speech synthesis model is trained based on the audio style vector training method of the first aspect or any corresponding implementation method thereof.
[0054] In a fifth aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the audio style vector training method of the above-mentioned first aspect or any corresponding embodiment thereof, or executes the audio style vector generation method of the above-mentioned second aspect or any corresponding embodiment thereof by executing the computer instructions.
[0055] In a sixth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the audio style vector training method of the above-mentioned first aspect or any corresponding embodiment thereof, or to execute the audio style vector generation method of the above-mentioned second aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0057] Figure 1 is a flowchart of an audio style vector training method according to an embodiment of the present invention;
[0058] Figure 2 is a schematic diagram of a first speech synthesis model according to an embodiment of the present invention;
[0059] Figure 3 is a schematic diagram of a second speech synthesis model according to an embodiment of the present invention;
[0060] Figure 4 is a flowchart of a method for generating an audio style vector according to an embodiment of the present invention;
[0061] Figure 5 is a structural block diagram of an audio style vector training device according to an embodiment of the present invention;
[0062] Figure 6 is a structural block diagram of an audio style vector generating device according to an embodiment of the present invention;
[0063] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0065] The Tacotron model relies on speaker information when extracting audio style vectors. When the speaker information is unclear or incorrect, the extracted audio style vector may be biased, affecting the accuracy of the audio style vector, and the flexibility and generalization ability of the model are poor. The audio style vector training method provided in an embodiment of the present invention can accurately extract the style vector in the audio without relying on speaker information, and thus perform training based on the style vector, so that the trained speech synthesis model can accurately extract the style vector of the audio and accurately perform speech synthesis.
[0066] According to an embodiment of the present invention, an embodiment of an audio style vector training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0067] In this embodiment, a method for training an audio style vector is provided, which can be used in mobile terminals, such as mobile phones, computers, etc. Figure 1 is a flow chart of an audio style vector training method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0068] Step S101, obtain multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features. Any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio. Specifically, the multiple sample audios include multiple languages, which can be obtained from existing public data sets or other methods. The sample text of any sample audio contains various information, such as age, gender, speaking style, etc., which is used to comprehensively describe the style of the sample audio. The audio features of any sample audio include phonemes, pitch features, spectrograms, and mel-spectrograms. By obtaining multiple sample audios and obtaining corresponding text and audio features, a data basis is provided for subsequent training.
[0069] Step S102, for any sample audio, a style coding layer is used to determine the sample style vector of the sample audio based on the sample audio features of the sample audio. Specifically, the style coding layer includes an audio coding layer and an attention layer. The audio coding layer includes a convolutional layer, a pooling layer, a long short-term memory network, and a fully connected layer. The Mel-spectrogram of the sample audio is input into the audio coding layer, and a feature vector of a fixed length is output through step-by-step processing of each layer. Then, the feature vector is input into the attention layer, and the feature vector is weighted to highlight the more important part of the style expression, so as to generate a vector that more concentrates on expressing the audio style, that is, the sample style vector. Optionally, the style coding layer is an existing structure, and extracting the sample style vector through the style coding layer is a prior art, which will not be repeated here.
[0070] Step S103, using a speaker classifier to classify the sample audio to obtain a sample classification result. Specifically, a pre-trained speaker classifier is used, and the speaker classifier is a neural network structure. The sample audio is classified according to it to determine which speaker the sample audio belongs to. It should be noted that in the related art, the speech synthesis model is trained when the speaker of the audio is clear, but after the sample audio is classified, the present application does not input the sample classification result into the following model for training.
[0071] Step S104, using the first speech synthesis model and the second speech synthesis model, based on the sample style vector and the sample audio features of the sample audio to perform speech synthesis, respectively obtaining the first synthesized audio and the second synthesized audio, the first speech synthesis model is an unsupervised model of the conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of the conditional variational autoencoding structure. Specifically, the first speech synthesis model and the second speech synthesis model are both models of the conditional variational autoencoding structure, but when training the model, the first speech synthesis model does not require supervision, while the second speech synthesis model requires the phonemes in the sample audio features as supervision for more accurate learning. The sample style vector and the sample audio features are used as inputs of the two models to generate synthesized audios respectively.
[0072] Step S105, determining a first synthesis loss based on the first synthesized audio and the sample audio, determining a second synthesis loss based on the second synthesized audio and the sample audio, and determining a classification loss based on the sample classification result of the sample audio and the speaker label. Specifically, the difference between the synthesized audio and the original sample audio is compared, and the synthesis loss is calculated using any loss function. At the same time, the difference between the predicted result of the speaker classifier and the actual speaker label is compared, and the classification loss is calculated, so that the performance of the synthesis task and the classification task can be measured according to the synthesis loss and the classification loss, and the optimization direction of the model is guided.
[0073] Step S106, based on the first synthesis loss, the second synthesis loss and the classification loss, the first speech synthesis model, the second speech synthesis model and the style coding layer are trained, and the trained style coding layer is integrated into the trained first speech synthesis model and the trained second speech synthesis model to obtain the first target speech synthesis model and the second target speech synthesis model. Specifically, all the losses calculated above are combined, and the parameters of the model are updated by back propagation to minimize the loss, so that the model performance is better. After the style coding layer, the first speech synthesis model and the second speech synthesis model are trained, that is, after the model parameters reach the preset threshold or have been trained for a preset number of rounds, the style coding layer is integrated into the two speech synthesis models to form the final first target speech synthesis model and the second target speech synthesis model, so that the final speech synthesis model can directly extract the style vector from the input audio without the participation of speaker information, and generate high-quality synthesized audio based on this vector.
[0074] The audio style vector training method provided by the embodiment of the present invention can cover a variety of different styles and speakers by using multiple sample audios, thereby improving the generalization ability of the model. The sample text provides additional contextual information, which helps the model to better understand the audio content and style. The style coding layer can extract style-related information from audio features without relying on speaker information. The speaker classifier is used for classification, which helps the model learn the characteristics of different speakers. The unsupervised model and the supervised model of the conditional variational automatic encoding structure are used for speech synthesis respectively to obtain the first synthesized audio and the second synthesized audio, providing two different synthesis paths, thereby enhancing the robustness and flexibility of the model. Without relying on speaker information, the style vector in the audio can be accurately extracted, so that training can be performed based on the style vector. By considering the synthesis loss and the classification loss at the same time, the model can be optimized from multiple angles. Finally, the trained style coding layer is respectively integrated into the two trained speech synthesis models, so that the two speech synthesis models can accurately extract the style vector in the audio, and can perform speech synthesis more accurately according to the style vector of the audio.
[0075] In this embodiment, a method for training an audio style vector is provided, which can be used in the above-mentioned mobile terminal, such as a mobile phone, a computer, etc. The method specifically includes the following steps:
[0076] Step S201, obtaining multiple initial audios and initial texts corresponding to the multiple initial audios. Specifically, the initial audios can be obtained from existing public data sets or in other ways. The initial text of any initial audio contains various information, such as age, gender, speaking style, etc., which is used to fully describe the style of the audio.
[0077] Step S202: for any initial audio, adjust the sampling rate of the initial audio to obtain a sample audio corresponding to the initial audio. Specifically, all initial audios are unified to the same sampling rate, such as 22K, to ensure data consistency and compatibility.
[0078] Step S203, regularize the initial text corresponding to the initial audio to obtain a sample text corresponding to the initial text. Specifically, the initial text is regularized to meet the requirements of subsequent model training. The regularization process includes removing irrelevant characters, word segmentation, removing stop words, and normalizing numbers.
[0079] Step S204, obtaining multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features, wherein any sample audio carries a speaker label, the sample text is used to describe the style of the sample audio, and the sample audio features include phonemes, pitch features, orthograms, and mel-spectrograms.
[0080] Specifically, the above step S204 obtains the sample audio features corresponding to the multiple sample audios, including:
[0081] Step S2041, convert the sample text of the sample audio into phonemes to obtain the phonemes of the sample text. Specifically, convert each word and pinyin in the sample text into a corresponding phoneme sequence.
[0082] Step S2042, extracting the pitch features, orthogram and mel-spectrogram of the sample audio. Specifically, a pitch detection algorithm is used to analyze the sample audio to extract pitch features for representing the high and low changes of the sample audio. Using short-time Fourier transform, the sample audio is divided into multiple small segments, and Fourier transform is applied to each segment to obtain its spectrum to extract the orthogram of the sample audio for visualization and analysis of the spectral features of the sample audio. The sample audio is passed through a set of mel filters, which are spaced more closely in the low-frequency part and more sparsely in the high-frequency part to simulate the perceptual characteristics of the human ear to extract the mel-spectrogram of the sample audio.
[0083] Step S205: for any sample audio, use the style coding layer to determine the sample style vector of the sample audio based on the sample audio features of the sample audio. Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.
[0084] Step S206: Use a speaker classifier to classify the sample audio to obtain a sample classification result. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0085] Step S207, using the first speech synthesis model and the second speech synthesis model, perform speech synthesis based on the sample style vector and the sample audio features of the sample audio, and obtain the first synthesized audio and the second synthesized audio respectively, the first speech synthesis model is an unsupervised model of the conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of the conditional variational autoencoding structure.
[0086] Specifically, the above step S207 includes:
[0087] Step S2071, using the pitch posterior coding layer of the first speech synthesis model, based on the sample style vector of the sample audio and the pitch features in the sample audio features, feature extraction is performed to obtain a first sample latent vector. Specifically, the pitch posterior coding layer is used to extract a latent vector representing the audio pitch information from the input pitch features and sample style vectors. First, the input layer of the pitch posterior coding layer receives the pitch features and the sample style vector. Then, a feature fusion layer is used to fuse the pitch features and the sample style vector. Then, a neural network, such as a fully connected layer, a convolutional layer, a recursive neural network, etc., is used to encode the fused features and generate a latent vector. Finally, the first sample latent vector is output through the output layer.
[0088] Step S2072, using the Fourier coding layer of the first speech synthesis model, based on the sample style vector of the sample audio and the spectrogram in the sample audio feature, feature extraction is performed to obtain a second sample latent vector. Specifically, the Fourier coding layer is used to extract a latent vector representing the audio frequency domain information from the input spectrogram and sample style vector. First, the input layer of the Fourier coding layer receives the spectrogram and sample style vector of the sample audio. Then, using the feature fusion layer, the spectrogram and the sample style vector are fused. Then, using the Fourier transform layer, the fused features are Fourier transformed to extract the frequency domain information. Then, a neural network, such as a fully connected layer, a convolutional layer, a recursive neural network, etc., is used to encode the frequency domain information and generate a latent vector. Finally, the final second sample latent vector is output through the output layer.
[0089] Step S2073, using the vocoder of the first speech synthesis model, synthesizes the first synthesized audio based on the sample style vector, the first sample latent vector and the second sample latent vector of the sample audio. Specifically, the vocoder is used to convert multiple latent vectors into audible audio. First, the input layer of the vocoder receives the sample style vector, the first sample latent vector and the second sample latent vector. Then, a feature fusion layer is used to fuse these vectors. Then, a decoding layer is used to decode the fused features using a neural network, such as a fully connected layer, a convolutional layer, a recursive neural network, etc., to generate a mel-spectrogram. Then, an audio synthesis layer is used, and a special waveform generation technology can be used to convert the decoded mel-spectrogram into audio. Finally, the output layer is used to output the final first synthesized audio.
[0090] In some optional embodiments, Figure 2 is a schematic diagram of a first speech synthesis model according to an embodiment of the present invention, such as Figure 2 As shown, after the Mel spectrogram of the sample audio is input into the style coding layer, the sample style vector of the sample audio is obtained, and the sample style vector and the sample audio features are input into the first speech synthesis model for speech synthesis. First, the pitch posterior coding layer is used to extract features based on the sample style vector and the pitch features to obtain the first sample latent vector. Then, the Fourier coding layer is used to extract features based on the sample style vector and the orthogram to obtain the second sample latent vector. Finally, the vocoder is used to decode and synthesize the first sample latent vector, the second sample latent vector and the sample style vector to obtain the first synthesized audio.
[0091] Step S2074, using the pitch posterior coding layer of the second speech synthesis model, based on the sample style vector of the sample audio and the pitch features in the sample audio features, feature extraction is performed to obtain a third sample latent vector. Specifically, using the pitch posterior coding layer of the second speech synthesis model, refer to the above step S2071 to obtain the third sample latent vector.
[0092] Step S2075, using the Fourier coding layer of the second speech synthesis model, extracting features based on the sample style vector of the sample audio and the spectrogram in the sample audio feature, to obtain a fourth sample latent vector. Specifically, using the Fourier coding layer of the second speech synthesis model, refer to the above step S2072 to obtain the fourth sample latent vector.
[0093] Step S2076, using the stream coding layer of the second speech synthesis model, perform feature extraction based on the third sample latent vector and the fourth sample latent vector to obtain the fifth sample latent vector. Specifically, the stream coding layer is used to extract a higher-level, more expressive latent vector from multiple input latent vectors. First, the input layer of the stream coding layer is used to receive the third sample latent vector and the fourth sample latent vector. Then, the feature fusion layer is used to fuse the two latent vectors. Then, the stream transformation layer is used to process the fused features using stream transformation techniques, such as coupling flow and autoregressive flow, to generate a new latent vector. Finally, the output layer is used to output the final fifth sample latent vector.
[0094] Step S2077, using the sample text encoding layer and protection layer of the second speech synthesis model, based on the phonemes in the sample audio features of the sample audio, a sample text encoding is generated. Specifically, the sample text encoding layer is used to convert the input phonemes into a high-level sample text encoding with rich semantic information, and the protection layer is used to ensure that the generated sample text encoding can retain important phoneme information. First, the input layer of the sample text encoding layer is used to receive the phonemes of the sample audio. Then, the embedding layer is used to map the phonemes to a high-dimensional vector space. Then, the encoding layer is used to use a neural network, such as a recurrent neural network, a long short-term memory network, a gated recurrent unit or a transformer, etc., and the self-attention mechanism is used to encode the embedded phoneme vector to generate a sample text encoding. Finally, the protection layer is used to ensure that the generated sample text encoding can retain important phoneme information. Optionally, the protection layer can be implemented in a variety of ways, such as residual connection and attention mechanism, etc. The residual connection ensures that important information is not ignored by adding input to the output of the encoding layer, and the attention mechanism explicitly emphasizes certain key phonemes by adding an attention mechanism during the encoding process.
[0095] Step S2078, using a rhythmic one-way search algorithm, based on the fifth sample latent vector and the sample text encoding, to extract the sample prosody information. Specifically, the rhythmic one-way search algorithm is used to extract prosody information from the sample text encoding and latent vector. Prosody information usually includes features such as intonation, rhythm, and stress, which are crucial for naturalness and expressiveness in speech synthesis. First, the fifth sample latent vector and the sample text encoding are received through the input layer. Then, the fusion layer is used to fuse the fifth sample latent vector and the sample text encoding. Then, the prosody feature extraction layer is used to extract prosody information from the fused features using a one-way search algorithm. The one-way search algorithm is a step-by-step forward search method for extracting prosody features from the fused features. By searching forward step by step, each step updates the prosody features according to the current state and the fused feature vector. Finally, the final sample prosody information is output through the output layer.
[0096] Step S2079, using the prediction layer of the second speech synthesis model, based on the sample style vector, sample text encoding and sample prosody information of the sample audio, synthesize the second synthesized audio. Specifically, the prediction layer is used to convert the input sample style vector, sample text encoding and sample prosody information into the final synthesized audio. First, the input layer of the prediction layer is used to receive the sample style vector, sample text encoding and sample prosody information. Then, the fusion layer is used to fuse these input features to generate a comprehensive feature vector. Then, the acoustic model layer is used to generate a Mel spectrum from the fused feature vector using a neural network such as a recurrent neural network, a long short-term memory network, a gated recurrent unit or a transformer. Then, the vocoder layer is used to convert the Mel spectrum into the final synthesized audio waveform. Finally, the output layer is used to output the final second synthesized audio.
[0097] In some optional embodiments, Figure 3 is a schematic diagram of a second speech synthesis model according to an embodiment of the present invention, such as Figure 3As shown, after the Mel spectrogram of the sample audio is input into the style coding layer, the sample style vector of the sample audio is obtained, and the sample style vector and the sample audio features are input into the second speech synthesis model for speech synthesis. First, the pitch posterior coding layer is used to extract features based on the sample style vector and the pitch feature to obtain the third sample latent vector. Then, the Fourier coding layer is used to extract features based on the sample style vector and the direct spectrogram to obtain the fourth sample latent vector. Then, the stream coding layer is used to extract features in combination with the third sample latent vector and the fourth sample latent vector to obtain the fifth sample latent vector. Then, the sample text coding layer and the protection layer are used to generate the sample text coding based on the phoneme. Then, the rhythmic one-way search algorithm is used to extract based on the fifth sample latent vector and the sample text coding to obtain the sample rhythmic information. Finally, the prediction layer is used to synthesize in combination with the sample style vector, the sample text coding and the sample rhythmic information to obtain the second synthesized audio.
[0098] Step S208: Determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on the sample classification result of the sample audio and the speaker label. Figure 1 Step S105 of the illustrated embodiment will not be described in detail here.
[0099] Step S209: Based on the first synthesis loss, the second synthesis loss and the classification loss, the first speech synthesis model, the second speech synthesis model and the style coding layer are trained, and the trained style coding layer is integrated into the trained first speech synthesis model and the trained second speech synthesis model to obtain the first target speech synthesis model and the second target speech synthesis model. For details, please refer to Figure 1 Step S106 of the illustrated embodiment will not be described in detail here.
[0100] The audio style vector training method provided by the embodiment of the present invention can cover a variety of different styles and speakers by using multiple sample audios, thereby improving the generalization ability of the model. The sample text provides additional contextual information, which helps the model to better understand the audio content and style. The style coding layer can extract style-related information from audio features without relying on speaker information. The speaker classifier is used for classification, which helps the model learn the characteristics of different speakers. The unsupervised model and the supervised model of the conditional variational automatic encoding structure are used for speech synthesis respectively to obtain the first synthesized audio and the second synthesized audio, providing two different synthesis paths, thereby enhancing the robustness and flexibility of the model. Without relying on speaker information, the style vector in the audio can be accurately extracted, so that training can be performed based on the style vector. By considering the synthesis loss and the classification loss at the same time, the model can be optimized from multiple angles. Finally, the trained style coding layer is respectively integrated into the two trained speech synthesis models, so that the two speech synthesis models can accurately extract the style vector in the audio, and can perform speech synthesis more accurately according to the style vector of the audio.
[0101] In this embodiment, a method for generating an audio style vector is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, computers, etc. Figure 4 FIG. 1 is a flow chart of a method for generating an audio style vector according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:
[0102] Step S401, obtaining the speaker's audio. Specifically, in the fields of speech synthesis and voiceprint recognition, there will be large errors when only using audio. If the style of the audio is combined, speech synthesis or voiceprint recognition can be performed more accurately. Therefore, a speaker's audio can be obtained from a recording device, a file, or other source to extract the style vector of the speaker's audio.
[0103] Step S402, determining whether the speaker audio carries corresponding text information, where the text information is used to describe the style of the speaker audio. Specifically, determining whether the speaker audio carries text information describing its style, such as "this is a passionate speech", "this audio is a calm narration", etc.
[0104] Step S403, when the speaker audio carries corresponding text information, a second target speech synthesis model is used to generate a style vector corresponding to the speaker audio, and the second target speech synthesis model is trained based on the audio style vector training method provided in the above embodiment. Specifically, if the speaker audio carries corresponding text information, the text information is used as supervision, and the second target speech synthesis model trained in the above embodiment is used. Since the model includes a trained style encoding layer, the style vector of the speaker audio can be accurately generated.
[0105] Step S404: If the speaker audio does not carry corresponding text information, a first target speech synthesis model is used to generate a style vector corresponding to the speaker audio, where the first target speech synthesis model is trained based on the audio style vector training method provided in the above embodiment. Specifically, if the speaker audio does not carry corresponding text information, the unsupervised first target speech synthesis model trained in the above embodiment can be used to generate a style vector for the speaker audio.
[0106] The audio style vector generation method provided by the embodiment of the present invention can also accurately extract the style vector by adopting an unsupervised first target speech synthesis model when the speaker audio does not carry text information, so as to improve the generalization ability of obtaining the style vector.
[0107] In this embodiment, an audio style vector training device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0108] This embodiment provides an audio style vector training device, such as Figure 5 As shown, including:
[0109] The first acquisition module 501 is used to acquire multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features. Any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio.
[0110] The first determination module 502 is used to determine, for any sample audio, a sample style vector of the sample audio based on the sample audio features of the sample audio by using a style coding layer.
[0111] The classification module 503 is used to classify the sample audio using a speaker classifier to obtain a sample classification result.
[0112] The synthesis module 504 is used to adopt the first speech synthesis model and the second speech synthesis model to perform speech synthesis based on the sample style vector and the sample audio features of the sample audio, and obtain the first synthesized audio and the second synthesized audio respectively. The first speech synthesis model is an unsupervised model of the conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of the conditional variational autoencoding structure.
[0113] The second determination module 505 is used to determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on a sample classification result of the sample audio and a speaker label.
[0114] The training module 506 is used to train the first speech synthesis model, the second speech synthesis model and the style coding layer based on the first synthesis loss, the second synthesis loss and the classification loss, and integrate the trained style coding layer into the trained first speech synthesis model and the trained second speech synthesis model to obtain the first target speech synthesis model and the second target speech synthesis model.
[0115] In some optional implementations, before the first acquisition module 501, the device further includes:
[0116] The initial acquisition module is used to acquire multiple initial audios and initial texts corresponding to the multiple initial audios.
[0117] The first preprocessing module is used to adjust the sampling rate of any initial audio to obtain a sample audio corresponding to the initial audio.
[0118] The second preprocessing module is used to regularize the initial text corresponding to the initial audio to obtain a sample text corresponding to the initial text.
[0119] In some optional implementations, the sample audio features include phonemes, pitch features, orthograms, and mel-spectrograms, and the first acquisition module 501 includes:
[0120] The first determination unit is used to phonemize the sample text of the sample audio to obtain the phonemes of the sample text.
[0121] The extraction unit is used to extract the pitch features, orthogram and mel-spectrogram of the sample audio.
[0122] In some optional implementations, the synthesis module 504 includes:
[0123] The second determination unit is used to use the pitch posterior coding layer of the first speech synthesis model to extract features based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a first sample latent vector.
[0124] The third determination unit is used to use the Fourier coding layer of the first speech synthesis model to extract features based on the sample style vector of the sample audio and the spectrogram in the sample audio feature to obtain a second sample latent vector.
[0125] The first synthesis unit is used to synthesize the first synthesized audio based on the sample style vector, the first sample latent vector and the second sample latent vector of the sample audio by using the vocoder of the first speech synthesis model.
[0126] In some optional implementations, the synthesis module 504 includes:
[0127] The fourth determination unit is used to use the pitch posterior coding layer of the second speech synthesis model to extract features based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a third sample latent vector.
[0128] The fifth determination unit is used to use the Fourier coding layer of the second speech synthesis model to extract features based on the sample style vector of the sample audio and the spectrogram in the sample audio feature to obtain a fourth sample latent vector.
[0129] The sixth determination unit is used to use the stream coding layer of the second speech synthesis model to perform feature extraction based on the third sample latent vector and the fourth sample latent vector to obtain a fifth sample latent vector.
[0130] The seventh determination unit is used to generate a sample text code based on the phonemes in the sample audio features of the sample audio by using the sample text coding layer and the protection layer of the second speech synthesis model.
[0131] The eighth determination unit is used to extract the sample prosody information based on the fifth sample latent vector and the sample text encoding by adopting a prosody one-way search algorithm.
[0132] The second synthesis unit is used to synthesize a second synthesized audio based on the sample style vector, the sample text encoding and the sample prosody information of the sample audio by using the prediction layer of the second speech synthesis model.
[0133] In this embodiment, an audio style vector generation device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0134] This embodiment provides an audio style vector generation device, such as Figure 6 As shown, including:
[0135] The second acquisition module 601 is used to acquire the speaker audio.
[0136] The judgment module 602 is used to judge whether the speaker's audio carries corresponding text information.
[0137] The first generation module 603 is used to use a second target speech synthesis model to generate a style vector corresponding to the speaker audio when the speaker audio carries corresponding text information. The second target speech synthesis model is trained based on the audio style vector training method provided in the above embodiment.
[0138] In some optional implementations, after determining module 602, the device further includes:
[0139] The second generation module is used to use the first target speech synthesis model to generate a style vector corresponding to the speaker's audio when the speaker's audio does not carry corresponding text information. The first target speech synthesis model is trained based on the audio style vector training method provided in the above embodiment.
[0140] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0141] The audio style vector training device and the audio style vector generation device in this embodiment are presented in the form of functional units, where the units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.
[0142] The embodiment of the present invention also provides a computer device having the above Figure 5 The audio style vector training device and Figure 6 The audio style vector generating device shown.
[0143] See also Figure 7 , Figure 7 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 7As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.
[0144] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0145] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0146] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0147] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0148] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0149] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0150] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0151] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for training an audio style vector, characterized in that: The method comprises: Acquire multiple sample audios, sample texts corresponding to the multiple sample audios, and sample audio features, wherein any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio; For any sample audio, using a style coding layer, based on sample audio features of the sample audio, determining a sample style vector of the sample audio; Using a speaker classifier to classify the sample audio to obtain a sample classification result; Using a first speech synthesis model and a second speech synthesis model, speech synthesis is performed based on the sample style vector and the sample audio features of the sample audio, to obtain a first synthesized audio and a second synthesized audio, respectively, wherein the first speech synthesis model is an unsupervised model of a conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of a conditional variational autoencoding structure; Determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on a sample classification result of the sample audio and a speaker label; Based on the first synthesis loss, the second synthesis loss and the classification loss, the first speech synthesis model, the second speech synthesis model and the style coding layer are trained, and the trained style coding layer is integrated into the trained first speech synthesis model and the trained second speech synthesis model to obtain the first target speech synthesis model and the second target speech synthesis model.
2. The method according to claim 1, characterized in that: Before obtaining a plurality of sample audios, sample texts corresponding to the plurality of sample audios, and sample audio features, the method further includes: Acquire multiple initial audios and initial texts corresponding to the multiple initial audios; For any initial audio, adjusting the sampling rate of the initial audio to obtain a sample audio corresponding to the initial audio; Regularize the initial text corresponding to the initial audio to obtain a sample text corresponding to the initial text.
3. The method according to claim 1, characterized in that The sample audio features include phonemes, pitch features, orthograms, and mel-spectrograms. The step of obtaining the sample audio features corresponding to the plurality of sample audios includes: Phonemizing the sample text of the sample audio to obtain phonemes of the sample text; Extract the pitch features, orthogram and mel-spectrogram of the sample audio.
4. The method according to claim 3, characterized in that The method of adopting a first speech synthesis model to perform speech synthesis based on a sample style vector and a sample audio feature of the sample audio to obtain a first synthesized audio includes: Using the pitch posterior coding layer of the first speech synthesis model, feature extraction is performed based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a first sample latent vector; Using the Fourier coding layer of the first speech synthesis model, extracting features based on the sample style vector of the sample audio and the orthogram in the sample audio feature to obtain a second sample latent vector; The first synthesized audio is synthesized by using a vocoder of the first speech synthesis model based on the sample style vector of the sample audio, the first sample latent vector and the second sample latent vector.
5. The method according to claim 3, characterized in that: The adopting the second speech synthesis model to perform speech synthesis based on the sample style vector and the sample audio feature of the sample audio to obtain the second synthesized audio includes: Using the pitch posterior coding layer of the second speech synthesis model, extracting features based on the sample style vector of the sample audio and the pitch features in the sample audio features to obtain a third sample latent vector; Using the Fourier coding layer of the second speech synthesis model, extracting features based on the sample style vector of the sample audio and the orthogram in the sample audio feature to obtain a fourth sample latent vector; Using the stream coding layer of the second speech synthesis model, extracting features based on the third sample latent vector and the fourth sample latent vector to obtain a fifth sample latent vector; Using the sample text encoding layer and the protection layer of the second speech synthesis model, based on the phonemes in the sample audio features of the sample audio, generate a sample text encoding; Using a rhythmic one-way search algorithm, based on the fifth sample latent vector and the sample text encoding, to extract sample rhythmic information; The second synthesized audio is synthesized based on the sample style vector of the sample audio, the sample text encoding and the sample prosody information using the prediction layer of the second speech synthesis model.
6. A method for generating an audio style vector, characterized in that: The method comprises: Get speaker audio; Determining whether the speaker audio carries corresponding text information, where the text information is used to describe the style of the speaker audio; In the case where the speaker audio carries corresponding text information, a second target speech synthesis model is used to generate a style vector corresponding to the speaker audio, and the second target speech synthesis model is trained based on the audio style vector training method according to any one of claims 1 to 5.
7. The method according to claim 6, characterized in that After determining whether the speaker audio carries corresponding text information, the method further includes: When the speaker audio does not carry corresponding text information, a first target speech synthesis model is used to generate a style vector corresponding to the speaker audio, and the first target speech synthesis model is trained based on the audio style vector training method according to any one of claims 1 to 5.
8. An audio style vector training device, characterized in that: The device comprises: A first acquisition module is used to acquire a plurality of sample audios, sample texts corresponding to the plurality of sample audios, and sample audio features, wherein any sample audio carries a speaker label, and the sample text is used to describe the style of the sample audio; A first determination module is used to determine, for any sample audio, a sample style vector of the sample audio based on the sample audio features of the sample audio by using a style coding layer; A classification module, used to classify the sample audio using a speaker classifier to obtain a sample classification result; A synthesis module, configured to use a first speech synthesis model and a second speech synthesis model to perform speech synthesis based on a sample style vector and a sample audio feature of the sample audio, to obtain a first synthesized audio and a second synthesized audio, respectively, wherein the first speech synthesis model is an unsupervised model of a conditional variational autoencoding structure, and the second speech synthesis model is a supervised model of a conditional variational autoencoding structure; a second determination module, configured to determine a first synthesis loss based on the first synthesized audio and the sample audio, determine a second synthesis loss based on the second synthesized audio and the sample audio, and determine a classification loss based on a sample classification result of the sample audio and a speaker label; A training module is used to train the first speech synthesis model, the second speech synthesis model and the style coding layer based on the first synthesis loss, the second synthesis loss and the classification loss, integrate the trained style coding layer into the trained first speech synthesis model and the trained second speech synthesis model, and obtain a first target speech synthesis model and a second target speech synthesis model.
9. An audio style vector generation device, characterized in that: The device comprises: A second acquisition module is used to acquire speaker audio; A judgment module, used to judge whether the speaker audio carries corresponding text information; The first generation module is used to use a second target speech synthesis model to generate a style vector corresponding to the speaker audio when the speaker audio carries corresponding text information, and the second target speech synthesis model is trained based on the audio style vector training method described in any one of claims 1 to 5.
10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the audio style vector training method described in any one of claims 1 to 5, or executes the audio style vector generation method described in any one of claims 6 and 7 by executing the computer instructions.
Citation Information
Patent Citations
Speech synthesis method and device based on style cloning and storage medium
CN118197349A
Speech synthesis system based on cue word, training method and reasoning method
CN118762685A