A self-supervised speech feature enhancement speech synthesis method based on mutual information theory

By introducing a self-supervised speech feature enhancement method based on mutual information theory in the speech synthesis technology, the problem of insufficient integration of acoustic information is solved, and a more natural and high-quality speech synthesis is achieved.

CN119964551BActive Publication Date: 2025-06-24XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510430211.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-24
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing speech synthesis technology has problems such as insufficient integration of acoustic information and limited improvement in speech nature.

Method used

The self-supervised speech feature enhancement method based on mutual information theory is adopted, and acoustic information is fully integrated to improve the naturalness and quality of speech synthesis through self-supervised speech feature extraction, information bottleneck module design and text representation enhancement.

Benefits of technology

Significantly improves the naturalness and quality of speech synthesis, enhances the predictive ability of text representations to key acoustic features, and performs well in multispeakers and across languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964551B_ABST
    Figure CN119964551B_ABST
Patent Text Reader

Abstract

A self-supervised speech feature enhancement speech synthesis method based on mutual information theory, which relates to the field of speech synthesis technology. This method introduces self-supervised speech features as acoustic supplements to the text, designs an information bottleneck module based on maximizing and minimizing mutual information, extracts compact and task-related self-supervised representations from the self-supervised speech features, and enhances the acoustic information of the text representation by maximizing the mutual information between the text representation and the self-supervised representation, thereby improving the naturalness and quality of speech synthesis. It performs excellently in both single-speaker and multi-speaker speech synthesis scenarios, and has good cross-language adaptability, and can effectively improve the speech synthesis quality in different language environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech synthesis, and in particular, to a self-supervised speech feature enhanced speech synthesis method based on the mutual information theory. This method is mainly used to improve the naturalness and quality of speech synthesis, and is also applicable to speech synthesis in multi-speaker scenarios and cross-lingual environments. Background Art

[0002] Speech synthesis technology aims to convert text information into speech output, and is widely used in fields such as intelligent voice assistants, audiobook generation, and voice broadcast systems, which is of great significance for improving the human-computer interaction experience and information dissemination efficiency.

[0003] With the development of deep learning technology, speech synthesis methods based on neural networks have gradually become a research hotspot. For example, the autoregressive speech synthesis model Tacotron2 can generate natural speech, but has problems such as low computational efficiency and difficulty in parallel processing; non-autoregressive models such as FastSpeech2 improve the synthesis efficiency and quality by introducing additional acoustic information (such as pitch, energy, etc.), but are still limited to the modeling of specific acoustic features and lack the modeling of more extensive acoustic information.

[0004] In recent years, self-supervised learning has achieved remarkable results in the field of speech representation learning; self-supervised speech representation models such as wav2vec 2.0, HuBERT, and BYOL-A can learn rich acoustic information from large-scale unsupervised speech data, including duration, energy, pitch, prosody, etc.; however, the application of these self-supervised speech features containing rich acoustic information in speech synthesis is still in the exploratory stage. Based on this, the present invention proposes a method based on the mutual information theory, which can effectively utilize self-supervised speech features to enhance speech synthesis.

[0005] In summary, the existing speech synthesis technology still has problems such as insufficient integration of acoustic information and limited improvement of speech naturalness. Therefore, there is an urgent need for a speech synthesis method that can effectively integrate self-supervised speech features to further improve the quality and applicability of speech synthesis. Summary of the Invention

[0006] The purpose of the present invention is to provide a self-supervised speech feature enhanced speech synthesis method based on the mutual information theory for the problems of insufficient information integration and limited speech naturalness existing in the existing speech synthesis technology. This method introduces self-supervised speech features and enhances text representation based on the mutual information theory to more comprehensively integrate acoustic information, thereby improving the naturalness and quality of speech synthesis.

[0007] To achieve the above-mentioned invention purpose, the present invention provides the following technical solutions.

[0008] A self-supervised speech feature enhanced speech synthesis method based on mutual information theory, comprising the following steps:

[0009] 1) Self-supervised speech feature extraction: Extract self-supervised speech features from a pre-trained self-supervised speech model, and use bilinear interpolation algorithm to upsample the self-supervised speech features along the time axis dimension to align their lengths to the same number of frames as the mel spectrogram, ensuring that the mutual information is calculated frame by frame with the mel spectrogram subsequently. These self-supervised speech features can capture various acoustic information including duration, energy, pitch, timbre, etc., and are used as acoustic information supplements for the text.

[0010] 2) Construct an information bottleneck module based on mutual information maximization and minimization: The self-supervised speech feature S is encoded by a self-supervised encoder to obtain a self-supervised representation Z. By maximizing the mutual information between the self-supervised representation Z and the mel spectrogram M, and at the same time minimizing the mutual information I(Z;S) between the self-supervised representation Z and the self-supervised speech feature S, a compact and task-related self-supervised representation Z is generated.

[0011] 3) Text representation enhancement: By maximizing the mutual information between the text representation and the self-supervised representation obtained in step 2), the text representation contains more information from the self-supervised representation, that is, the acoustic information of the text representation is supplemented to improve the naturalness and quality of speech synthesis.

[0012] 4) Speech synthesis: Construct a speech synthesis model. During training, use step 2) and step 3) to optimize the speech synthesis model. The trained speech synthesis model is used for inference to complete the self-supervised speech feature enhanced speech synthesis based on mutual information theory; step 2) can obtain a compact and task-related self-supervised representation, and step 3) can use this self-supervised representation to enhance the text representation. Therefore, during inference, the text representation contains more acoustic information, which can enhance the performance of the speech synthesis model and improve the naturalness and quality of speech synthesis. Thus, the self-supervised speech feature enhanced speech synthesis based on mutual information theory is completed.

[0013] In step 1), the self-supervised speech model can adopt wav2vec 2.0, HuBERT, BYOL-A, etc., and the selection of its hidden layer is determined according to the model architecture and acoustic feature encoding ability.

[0014] The specific steps for extracting the self-supervised speech features can be: Use a pre-trained self-supervised speech model, input the speech waveform data, and after being encoded by the self-supervised speech model, extract the output of the hidden layer that encodes the most acoustic features of the self-supervised speech model as the self-supervised speech feature S.

[0015] Taking wav2vec 2.0 as an example, the specific steps can be as follows: Use the pre-trained wav2vec 2.0 as a self-supervised speech model. Input the speech waveform data. After being encoded by wav2vec 2.0, select the output of the sixth hidden layer of the wav2vec 2.0 model as the self-supervised speech feature S. Since the sixth hidden layer of this model encodes more acoustic features than other layers, such as key information like duration, pitch, energy, and prosody, which play an important role in improving the naturalness and quality of subsequent speech synthesis, therefore, select the output of the sixth hidden layer of the wav2vec 2.0 model as the self-supervised speech feature.

[0016] In step 2), the information bottleneck module designed based on maximizing and minimizing mutual information consists of a self-supervised encoder and an optimization formula based on mutual information theory; the self-supervised encoder uses the non-causal convolutional network WaveNet to encode the extracted self-supervised speech feature S to obtain the self-supervised representation Z; the non-causal convolutional network WaveNet uses 16 layers of causal convolutions with a fixed dilation factor of 1, with a convolutional kernel size of 5 for each layer and 192 channels; stable training is achieved through a dual-path residual structure (skip summation + residual accumulation) and 10% probability Dropout;

[0017] The optimization formula based on mutual information theory is used to enhance the self-supervised representation Z. By the optimization formula, a compact and task-related self-supervised representation Z is extracted. The optimization formula is:

[0018]

[0019] Among them, represents the mutual information between the self-supervised representation Z and the mel spectrogram M. By maximizing this mutual information, it is ensured that the self-supervised representation Z learns the acoustic information related to the speech synthesis task; represents the mutual information between the self-supervised representation Z and the self-supervised speech feature S. By minimizing this mutual information, the compactness of the self-supervised representation Z is maintained and redundant information is eliminated; γ is a weight hyperparameter used to balance the relationship between the two; specifically, it is necessary to ensure both the compactness of Z and that Z can encode task-related information. The value range of the weight hyperparameter γ can be 0.8 to 1.2; in a preferred embodiment, the same weight is taken for the optimization of the two mutual informations, that is, γ is set to 1.0 to balance the compactness and task relevance of the self-supervised representation Z;

[0020] In addition, , both are estimated using the MINE framework. MINE estimates the mutual information through the following inequality:

[0021]

[0022] Among them, a neural network is used to approximate the value of the expression on the right side of the inequality sign, and then an estimated value of the mutual information is obtained. The neural network consists of a 3-layer MLP, and the dimension of the hidden layer is 256. In the formula, represents the mutual information between the random variables X and Y, which measures the statistical dependence between the two; represents the function under the joint distribution of X and Y; represents the expectation of when X and Y are independently distributed.

[0023] The joint distribution sampling can be realized by maintaining the time alignment between the self-supervised features and the mel spectrogram, and the independent distribution sampling can be realized by randomly shuffling the time alignment relationship of the samples within the batch.

[0024] In step 3), in the text representation enhancement step, by maximizing the mutual information between the text representation T and the self-supervised representation Z, the acoustic information of the text representation is enhanced; the specific steps can be:

[0025] The text encoder encodes the text sequence X into the text representation T, and then upsamples the text representation T by the length predicted by the length predictor to match the length of the mel spectrogram; by maximizing the mutual information between the text representation T and the self-supervised representation Z, the acoustic information of the text representation is enhanced; the specific optimization formula is:

[0026]

[0027] Among them, is the mutual information between the text representation T and the self-supervised representation Z, which is also estimated using the MINE framework. The specific structure of the neural network is the same as that in step 2); by maximizing this mutual information, the text representation T can better capture the acoustic information in the self-supervised representation Z, thereby enhancing the acoustic information of the text representation T.

[0028] In step 4), a variational inference (VAE) network, a flow-based prior enhancement network (VP - Flow), and a post-processing network (post-net) are used as the basic architecture of the speech synthesis model. When training the speech synthesis model, the two optimization formulas proposed in step 2) and step 3) are added to complete the self-supervised speech feature enhanced speech synthesis based on the mutual information theory;

[0029] During training: The text encoder encodes the text sequence into a text representation T, and then upsamples it according to the Mel spectrogram frame length predicted by the length predictor for each text representation to match the length of the Mel spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel spectrogram and inputs it into the VAE decoder for Mel spectrogram restoration; meanwhile, the posterior distribution is processed by VP-Flow and transformed into a standard Gaussian distribution; and the post-net also converts the Mel spectrogram samples into a standard Gaussian distribution; both VP-Flow and post-net are reversible. During training, VP-Flow transforms the posterior distribution into a standard Gaussian distribution conditioned on the text representation, and during inference, it can sample from the standard Gaussian distribution and approximately restore the posterior distribution according to the text representation; while during training, the post-net converts the real Mel spectrogram into a standard Gaussian distribution conditioned on the preliminary estimate of the Mel spectrogram decoded by the VAE decoder, and during inference, it can sample from the standard Gaussian distribution to convert the preliminary estimate of the Mel spectrogram generated by the VAE decoder into a more delicate Mel spectrogram.

[0030] During inference, the text encoder encodes the text sequence into a text representation T, and then upsamples it according to the Mel spectrogram frame length predicted by the length predictor for each text representation to match the length of the Mel spectrogram; the enhanced text representation T and the latent variable sampled from the standard Gaussian distribution are input into VP-Flow for inverse transformation to obtain the latent variable of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel spectrogram; finally, the Post–net samples the latent variable from the standard Gaussian distribution and performs inverse transformation conditioned on the preliminary estimate of the Mel spectrogram to obtain the final Mel spectrogram; and then the Mel spectrogram is converted into a speech waveform output by means of a vocoder, thus completing the entire self-supervised speech feature enhanced speech synthesis process based on the mutual information theory.

[0031] Compared with the prior art, the advantages and prominent technical effects of the present invention are as follows:

[0032] 1. Introduction of self-supervised features: By introducing self-supervised speech features, the present invention can capture various acoustic information including duration, energy, pitch, etc. and the correlations between the information. Compared with the prior art methods that only focus on specific acoustic features, the self-supervised features introduced by the present invention integrate acoustic information more comprehensively, which helps to improve the naturalness and quality of the synthesized speech.

[0033] 2. Use of self-supervised features: In order to efficiently integrate self-supervised speech features into the TTS model, the present invention proposes an optimization method based on the mutual information theory, extracts compact and task-related self-supervised representations from the self-supervised speech features, and uses these self-supervised representations to enhance the text representation.

[0034] 3. The present invention conducts experiments on Chinese single-speaker and multi-speaker datasets. Compared with multiple speech synthesis methods, the Mean Opinion Score (MOS) is increased by 0.09 and 0.21 respectively compared to the best comparison model.

[0035] 4. Good cross-language adaptability: The method of the present invention has good cross-language adaptability. Experiments show that even in a low-resource language (Minnan dialect) environment not covered by the pre-trained model, it can still effectively improve the quality of speech synthesis and further expand the application scope.

[0036] 5. The text representation generated by the present invention significantly enhances the prediction ability for key acoustic features (including energy, duration, and pitch), indicating that the text representation encodes richer and more comprehensive acoustic information. Description of the Drawings

[0037] Figure 1 It is a schematic diagram of the overall framework process.

[0038] Figure 2 It is an example of the model training process of the present invention.

[0039] Figure 3 It is the optimization process of the information bottleneck module designed by the present invention for self-supervised representation, showing the process of information change contained in the self-supervised representation.

[0040] Figure 4 It is the text representation enhancement process, showing the process of enhancing the text representation by maximizing mutual information and supplementing the acoustic information missing in the text representation.

[0041] Figure 5 It is the model inference process of the present invention. Detailed Embodiments

[0042] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following embodiments will further illustrate the present invention in conjunction with the accompanying drawings.

[0043] Embodiments of the present invention introduce self-supervised speech features and enhance the text representation based on the mutual information theory to more comprehensively integrate acoustic information, thereby improving the naturalness and quality of speech synthesis.

[0044] As Figure 1 shown, the specific steps of the embodiments of the present invention are as follows.

[0045] 1) Self-supervised speech feature extraction: Select a pre-trained self-supervised speech model. In this embodiment, wav2vec2.0 is adopted. Load the wav2vec 2.0 model that has been pre-trained on large-scale unsupervised speech data from the pre-trained model library, and preprocess the original speech waveform data to ensure that the data format meets the input requirements of the wav2vec 2.0 model; in this embodiment, the original speech data is uniformly sampled to 22kHz, and the audio signal amplitude is normalized to the range of [-1, 1]; after the preprocessing is completed, input the speech data into the loaded wav2vec 2.0 model in an appropriate batch size.

[0046] Extract the output of the hidden layer as the self-supervised speech feature. Taking wav2vec 2.0 as an example in this embodiment, the wav2vec2.0 model encodes the input speech waveform data, and extracts the output of the hidden layer (the sixth hidden layer of wav2vec2.0) of the model as the self-supervised speech feature during the encoding process. Use the bilinear interpolation algorithm to upsample the self-supervised speech feature along the time axis dimension to align its length with the number of frames of the mel spectrogram; use the non-causal convolutional network WaveNet as the self-supervised encoder to encode the extracted self-supervised speech feature S to obtain the self-supervised representation Z; WaveNet uses 16 layers of causal convolutions with a fixed dilation factor of 1, the convolution kernel size of each layer is 5, and the number of channels is 192; stable training is achieved through a two-path residual structure (skip summation + residual accumulation) and 10% probability Dropout.

[0047] Specifically, based on the analysis of the outputs of different hidden layers of wav2vec 2.0 in "What all do audio transformer models hear? probing acoustic representations for language delivery and its structure" written by Jui Shah et al. and published in "arXiv preprint (arXiv:2101.00387, 2021)", since the sixth hidden layer of this model encodes more acoustic features than other layers, such as key information like duration, pitch, energy, and prosody, which play an important role in improving the naturalness and quality of subsequent speech synthesis, therefore, select the output of the sixth hidden layer as the self-supervised speech feature.

[0048] Select a pre-trained self-supervised speech model (such as wav2vec 2.0, HuBERT, BYOL-A, etc.) and load the pre-trained model from the pre-trained model library. The selection of the hidden layer of different models needs to be determined according to their architectures and acoustic feature encoding capabilities: for the HuBERT model, the output of its intermediate layer can be selected, specifically depending on the encoding efficiency of the model for acoustic information; for the BYOL-A model, the optimal hidden layer position can be determined through ablation experiments.

[0049] Upsample the self-supervised speech features through the bilinear interpolation algorithm to align their lengths with the number of frames of the mel spectrogram, so as to ensure frame-by-frame matching in subsequent mutual information calculations.

[0050] 2) Design of the information bottleneck module:

[0051] The information bottleneck module is one of the key components of the present invention. It is designed based on the principles of maximizing and minimizing mutual information, aiming to extract a compact self-supervised representation from the self-supervised speech features that is closely related to the speech synthesis task.

[0052] The overall architecture of the information bottleneck module consists of a self-supervised encoder and two optimization branches. The self-supervised encoder can adopt the non-causal convolutional network WaveNet to encode the extracted self-supervised speech features S to obtain the self-supervised representation Z. The two optimization branches are respectively used to process different mutual information, so as to realize the optimization of the self-supervised representation.

[0053] Furthermore, the non-causal convolutional network WaveNet adopts 16 layers of causal convolutions with a fixed dilation factor of 1. Each layer has a convolutional kernel size of 5 and 192 channels, and stable training is achieved through a dual-path residual structure (skip summation + residual accumulation) and 10% probability Dropout.

[0054] Furthermore, the two optimization branches are respectively used to process the mutual information between the self-supervised representation Z and the mel spectrogram M , and the mutual information between the self-supervised representation Z and the self-supervised feature S ; Each branch constructs a neural network structure with a 3-layer MLP and a hidden layer dimension of 256 to effectively estimate and optimize the mutual information.

[0055] The self-supervised representation Z is a compact feature processed by the information bottleneck module, which is the output of the self-supervised speech feature S encoded by the information bottleneck module, and is generated by maximizing and minimizing . Design the information bottleneck module based on maximizing and minimizing mutual information, and its optimization formula is:

[0056]

[0057] Among them, represents the mutual information between the self-supervised representation Z and the mel spectrogram M. By maximizing this mutual information, it is ensured that the self-supervised representation Z learns the acoustic information related to the speech synthesis task; represents the mutual information between the self-supervised representation Z and the self-supervised feature S. By minimizing this mutual information, the compactness of the self-supervised representation Z is maintained and redundant information is eliminated; γ is a weight hyperparameter used to balance the relationship between the two.

[0058] Specifically, it is necessary to ensure both the compactness of the self-supervised representation Z and that the self-supervised representation Z can encode task-related information. Therefore, in this embodiment, γ can be set to 1.0 to balance the compactness and task relevance of the self-supervised representation Z; when the value of γ is too large, the model will overly focus on the mutual information between the self-supervised representation Z and the mel spectrogram M, resulting in the self-supervised representation Z possibly containing too much redundant information and being unable to effectively maintain compactness; when the value of γ is too small, the model will overly emphasize the mutual information between the self-supervised representation Z and the self-supervised feature S, causing the self-supervised representation Z to lose some key acoustic information closely related to the speech synthesis task. Therefore, the same weight is taken for the two mutual informations, that is, γ is set to 1.0.

[0059] In addition, , both are estimated using the MINE framework, and MINE estimates the mutual information through the following inequality;

[0060]

[0061] where a neural network is used to approximate the value of the expression on the right side of the inequality sign, and then the estimated value of the mutual information is obtained. represents the mutual information between the random variables X and Y, measuring the statistical dependence between the two; represents the function under the joint distribution of X and Y; represents when X and Y are independently distributed, the expectation of; a neural network is used to approximate the value of the expression on the right side of the inequality sign, and then the estimated value of the mutual information is obtained. The neural network is composed of a 3-layer MLP, the hidden layer dimension is 256, the RELU activation function is used, and the optimization objective is:

[0062]

[0063] The Adam optimizer is used as the optimizer, where takes 0.9, takes 0.98, Take 10 to the power of negative 9. The learning rate decay refers to that in "Attention is all you need" published in "Advances in Neural Information Processing Systems" (30, 2017) by Vaswani et al. The number of training iterations and batch size are the same as those of the speech synthesis model. Taking the Chinese dataset CSMSC as an example: the number of iterations is 300K, and the batch size is 64.

[0064] 3) Enhancement of text representation:

[0065] To enable the text representation to capture acoustic information more effectively, the text representation obtained by encoding the text encoder is enhanced. The text encoder uses an encoder based on the Transformer architecture, and through the self-attention mechanism, it can effectively capture long-range dependencies in the text sequence and extract rich semantic features.

[0066] First, initialize the text encoder, set model parameters according to task requirements, such as the dimension of the hidden layer, the number of attention heads, etc.; preprocess the text sequence X in the format required by the encoder, including operations such as word segmentation and adding position encoding, and then input it into the text encoder; after multiple layers of processing by the encoder, the text encoder encodes the text sequence X into a text representation T. At this time, this text representation T contains the semantic and structural information of the text sequence X, but does not fully integrate acoustic information.

[0067] Since the length of the text representation T is usually inconsistent with the length of the mel spectrogram, a length predictor is introduced to predict the length of the mel spectrogram frames corresponding to the text representation to perform upsampling processing on the text representation T to achieve the matching of their lengths; by maximizing the mutual information between the text representation T and the self-supervised representation Z, the acoustic information of the text representation is enhanced; the network parameters are continuously optimized through training to ensure that the upsampled text representation can not only retain the original text features, but also be adapted to the mel spectrogram in length and contain more acoustic information. The specific optimization formula is:

[0068]

[0069] where It is the mutual information between the text representation T and the self-supervised representation Z. The estimation of the mutual information also uses the MINE framework for estimation. The specific structure of the neural network in this MINE framework is the same as the neural network structure when estimating the mutual information in step 2). By maximizing this mutual information, the text representation T can better capture the acoustic information in the self-supervised representation Z, thereby enhancing the acoustic information of the text representation. When the model converges, the obtained text representation T contains rich acoustic information, laying a solid foundation for high-quality speech synthesis.

[0070] 4) Speech synthesis:

[0071] Adopt a variational inference (VAE) network: a variational encoder and a variational decoder, a flow-based prior enhancement network (VP-Flow), and a post-processing network (Post-net), etc. as the basic architecture of the speech synthesis model. When training the speech synthesis model, add the two optimization formulas proposed in step 2) and step 3) to complete the self-supervised speech feature-enhanced speech synthesis based on the mutual information theory.

[0072] The following gives specific application examples.

[0073] Example 1: Single-speaker synthesis

[0074] 1. Data preparation

[0075] Taking the single-speaker Chinese dataset CSMSC as an example, use the single-speaker Chinese dataset CSMSC, which contains 10,000 speech samples with a total duration of 12 h. Divide the dataset into a training set, a validation set, and a test set, where the training set is used for model training, the validation set is used for model tuning, and the test set is used for evaluating model performance.

[0076] 2. Self-supervised feature extraction

[0077] In this embodiment, taking wav2vec 2.0 as an example, the original speech data is uniformly sampled to a sampling rate of 22 kHz, and the amplitude of the audio signal is normalized to the range of [-1, 1]; after preprocessing, the speech data is input into the loaded wav2vec 2.0 model with an appropriate batch size. The wav2vec 2.0 model encodes the input speech waveform data, and during the encoding process, the output of the hidden layer of the model (the sixth hidden layer of wav2vec 2.0) is extracted as the self-supervised speech feature. The bilinear interpolation algorithm is used to upsample the self-supervised speech feature along the time axis dimension to align its length with the number of frames of the Mel spectrogram; the non-causal convolutional network WaveNet is used as the self-supervised encoder to encode the extracted self-supervised speech feature S to obtain the self-supervised representation Z; WaveNet uses 16 layers of causal convolutions with a fixed dilation factor of 1, the convolutional kernel size of each layer is 5, and the number of channels is 192; stable training is achieved through a dual-path residual structure (skip summation + residual accumulation) and 10% probability Dropout.

[0078] 3. Self-Supervised Representation Enhancement

[0079] Extract the compact and task-related self-supervised representation S through the information bottleneck module; the overall architecture of the information bottleneck module consists of a self-supervised encoder and two optimization branches. The non-causal convolutional network WaveNet is used as the self-supervised encoder to encode the extracted self-supervised speech feature S to obtain the self-supervised representation Z; the non-causal convolutional network WaveNet uses 16 layers of causal convolutions with a fixed dilation factor of 1, the convolutional kernel size of each layer is 5, and the number of channels is 192. Stable training is achieved through a dual-path residual structure (skip summation + residual accumulation) and 10% probability Dropout; the two optimization branches are respectively used to process the mutual information between the self-supervised representation Z and the Mel spectrogram M, and the mutual information between the self-supervised representation Z and the self-supervised feature S; each branch constructs a neural network structure of a 3-layer MLP with a hidden layer dimension of 256 to effectively estimate and optimize the mutual information; the specific optimization process is as follows: maximize the mutual information between the self-supervised representation Z and the Mel spectrogram M , ensuring that the extracted representation contains information related to the speech synthesis task; calculate the mutual information between the self-supervised representation Z and the self-supervised feature S , and maintain the compactness of the extracted self-supervised representation by minimizing this mutual information; the specific optimization formula is as follows:

[0080] Obtain the final self-supervised representation Z; Figure 3 Illustrate the information conversion process of the self-supervised representation Z: through the optimization of the mutual information loss function, the self-supervised representation Z encodes the necessary information required to reconstruct the Mel spectrogram M.

[0081] 4. Enhancement of Text Representation

[0082] To enable the text representation to capture acoustic information more effectively, the text representation obtained by encoding with the text encoder is enhanced. The text encoder uses an encoder based on the Transformer architecture, which can effectively capture long-distance dependencies in the text sequence through the self-attention mechanism and extract rich semantic features.

[0083] First, the text encoder is initialized, which consists of 4 encoders of the Transformer architecture, with a hidden layer dimension of 192 and 2 attention heads. After preprocessing the text sequence X in the format required by the encoder, including word segmentation, adding position encoding, etc., it is input into the text encoder. After multiple layers of processing by the encoder, the text encoder encodes the text sequence X into a text representation T. At this time, this text representation T contains the semantic and structural information of the text sequence X, but does not fully integrate acoustic information.

[0084] Since the length of the text representation T is usually inconsistent with the length of the mel spectrogram, a length predictor is introduced to predict the length of the mel spectrogram frames corresponding to the text representation to perform upsampling processing on the text representation T to achieve the matching of their lengths; by maximizing the mutual information between the text representation T and the self-supervised representation Z, the acoustic information of the text representation is enhanced.

[0085] By maximizing the mutual information between the text representation T and the self-supervised representation Z,

[0086]

[0087] where is the mutual information between the text representation T and the self-supervised representation Z. By maximizing this mutual information, the text representation T can better capture the acoustic information in the self-supervised representation Z, thereby enhancing the acoustic information of the text representation. When the model converges, the obtained text representation T contains rich acoustic information, laying a solid foundation for high-quality speech synthesis. Figure 4 Describe the change process of the information contained in the text representation T, that is, more acoustic information from the self-supervised representation Z is supplemented. Using this enhanced text representation T in the subsequent speech synthesis process can improve the naturalness and quality of speech synthesis.

[0088] 5. Speech Synthesis

[0089] The variational inference (VAE) network, the flow-based prior enhancement network (VP-Flow), and the post-processing network (Post-net) described by Ren Yi et al. in "Portaspeech: Portable and high-quality generative text-to-speech" published in "Advances in Neural Information Processing Systems (34, 2021)" are used as the basic architecture of the speech synthesis model. As Figure 2 shown, the variational encoder receives the Mel spectrogram input, performs temporal downsampling through five layers of causal convolution (convolution kernel 5, stride 4), gradually expanding the number of channels from 80 to 192. At the end, a bidirectional LSTM is connected to extract context features, and finally mapped to the mean and variance parameters of a 16-dimensional latent variable. After reparameterized sampling of the latent variable, it is input into the flow-based prior enhancement network (VP-Flow), which contains a four-step reversible transformation chain. Each step sequentially performs channel permutation and affine coupling operations. The coupling layer generates a dynamic scaling factor through a three-layer fully connected network (hidden unit 64) to convert the output of the variational encoder into a standard Gaussian distribution. The latent variable is input into the variational decoder, and the coarse-grained Mel spectrogram is reconstructed through four layers of non-causal WaveNet (a residual convolution stack with increasing dilation coefficients), and the original temporal resolution is restored using transposed convolution (convolution kernel 5, stride 4). The flow network refines the output of the variational decoder. The first layer multiplies the 80-channel Mel spectrogram by 2 to 160 channels through a time-axis compression operation, and then performs twelve flow transformations. Every three flow transformations share a parameter group, each group containing activation normalization, four-group reversible neighborhood convolution (weights initialized with an orthogonal matrix), and a coupling block. The coupling block uses three layers of WaveNet (convolution kernel 3, hidden unit 192) to generate a spectral correction factor, and finally restores the time dimension through a decompression operation and outputs a high-fidelity Mel spectrogram. In the key parameter configuration of the model, the encoder WaveNet contains eight layers of dilated convolution, each coupling block of the post-processing network uses three layers of convolution, the latent variable optimization process realizes distribution enhancement through four-step reversible transformation, and all modules adopt dynamic conditional injection technology to map text features into convolution weight modulation signals to improve the content-rhythm alignment ability. This architecture, through a non-autoregressive parallel generation design (temporal modeling is completed by a single forward pass of the dilated convolution stack) and an explicit latent variable optimization strategy (the VP-Flow reversible transformation chain), significantly improves the inference efficiency while ensuring the synthesis quality.

[0090] During the training of the speech synthesis model, the two optimization formulas proposed in steps 3 and 4 are added to complete self-supervised speech feature enhancement for speech synthesis based on the mutual information theory.

[0091] As Figure 2 shown, during training: the text encoder encodes the text sequence into a text representation T, and then upsamples it according to the Mel spectrogram frame length predicted by the length predictor for each text representation to match the length of the Mel spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel spectrogram and inputs it into the VAE decoder for Mel spectrogram restoration; meanwhile, the posterior distribution is processed by VP-Flow and then pulled closer to the standard Gaussian distribution through the KL divergence, converting it into the standard Gaussian distribution; then the post-net also converts the Mel spectrogram samples into the standard Gaussian distribution; both VP-Flow and post-net are reversible. During training, VP-Flow processes the posterior distribution conditioned on the text representation and pulls it closer to the standard Gaussian distribution through the KL divergence, converting it into the standard Gaussian distribution. During inference, samples can be drawn from the standard Gaussian distribution and the posterior distribution can be approximately restored according to the text representation; while during training, the post-net converts the real Mel spectrogram into the standard Gaussian distribution conditioned on the preliminary estimate of the Mel spectrogram decoded by the VAE decoder, and during inference, samples can be drawn from the standard Gaussian distribution to convert the preliminary estimate of the Mel spectrogram generated by the VAE decoder into a more delicate Mel spectrogram. As shown, during inference, the text encoder encodes the text sequence into a text representation T, and then upsamples it according to the Mel spectrogram frame length predicted by the length predictor for each text representation to match the length of the Mel spectrogram; since the text representation T has been enhanced during training, the text representation T at this time has encoded rich acoustic information. The enhanced text representation T and the latent variable sampled from the prior distribution - the standard Gaussian distribution

[0092] As Figure 5 shown, during inference, the text encoder encodes the text sequence into a text representation T, and then upsamples it according to the Mel spectrogram frame length predicted by the length predictor for each text representation to match the length of the Mel spectrogram; since the text representation T has been enhanced during training, the text representation T at this time has encoded rich acoustic information. The enhanced text representation T and the latent variable sampled from the prior distribution - the standard Gaussian distribution are input into VP-Flow for inverse transformation to obtain the latent variable of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel spectrogram; finally, the Post - net samples the latent variable from the prior distribution - the standard Gaussian distribution and performs inverse transformation conditioned on the preliminary estimate of the Mel spectrogram to obtain the final Mel spectrogram; then the Mel spectrogram is converted into a speech waveform output with the help of a vocoder, thus completing the entire self-supervised speech feature enhancement speech synthesis process based on the mutual information theory.

[0093] Embodiment 2: Multi-speaker synthesis

[0094] Similar to Example 1, the difference lies in Step 1, data preparation: Taking the multi-speaker Chinese dataset AISHELL-3 as an example, the multi-speaker Chinese dataset AISHELL-3 is used, which contains 88,035 speech samples from 218 speakers, with a total duration of 85 hours; the dataset is divided into a training set, a validation set, and a test set, where the training set is used for model training, the validation set is used for model tuning, and the test set is used to evaluate model performance.

[0095] Steps 2-5 are the same as those in Example 1 for single speakers.

[0096] Experiments are conducted on two datasets: the single-speaker dataset CSMSC and the multi-speaker dataset AISHELL-3; the CSMSC dataset contains 10,000 utterances from female speakers, with a total audio duration of 12 hours; in contrast, the AISHELL-3 dataset is a multi-speaker dataset, containing 88,035 utterances from 218 speakers, with a total duration of 85 hours; for each dataset, 100 utterances are designated as the validation set, 500 utterances are designated as the test set, and the remaining data is used to train the model.

[0097] An average opinion score (MOS) subjective test is conducted, and the results are reported with a 95% confidence interval.

[0098] The synthesized speech by the method proposed in the present invention is subjectively evaluated and scored (MOS) compared with other systems, including 1) GT, the real audio; 2) GT (Mel + HiFi-GAN), where the real Mel spectrogram is extracted and converted into audio using the HiFi-GAN vocoder; 3) Tacotron 2, an autoregressive TTS model, and in the training process of the multi-speaker TTS system of the present invention, the speaker embedding is connected to the output of the text encoder; 4) FastSpeech 2, a non-autoregressive TTS model, and the speaker embedding setting for multi-speaker TTS training is the same as that of Tacotron2; 5) HierSpeech, an end-to-end TTS model, and the multi-speaker TTS training setting is the same as that of Tacotron2; 6) ParrotTTS, which replaces the Mel spectrogram in traditional speech synthesis with self-supervised speech features, and the multi-speaker TTS training setting is the same as that of Tacotron2.

[0099] Table 1

[0100]

[0101] As can be seen from Table 1, the model of the present invention achieves higher MOS scores than other models on both datasets; it is worth noting that on the multi-speaker dataset, the method of the present invention is significantly better than the existing TTS models and only slightly lags behind the theoretical performance upper limit (GT + HiFi-GAN); it is speculated that this is because in the context of large datasets, the MINE mutual information estimation framework used in the present invention becomes more accurate and stable, which is beneficial to the optimization process of the text representation of the present invention.

[0102] Example 3: Low-resource Language Extension

[0103] To further verify the generalization ability of the method of the present invention, additional experiments are conducted in a low-resource language environment, and Minnan dialect is selected as the dataset. It is worth noting that the pre-trained self-supervised speech model is not trained on the Minnan dialect; this setting can evaluate the performance of the method proposed in the present invention when dealing with languages not included in the pre-trained corpus, thereby further testing its cross-language generalization ability; in the experiment, SuiSiann is selected as the Minnan dialect dataset, where the dataset size is 3,467 speeches and the total duration is close to 5 hours. 100 utterances are designated as the validation set, 500 utterances are designated as the test set, and the remaining data is used to train the model, and the MOS scores are compared with the comparison methods in the experimental part of Table 1.

[0104] Table 2

[0105]

[0106] The experimental results in Table 2 show that although the pre-trained model is not trained on the Minnan dialect, the method of the present invention still significantly improves the quality of speech synthesis; specifically, the MOS scores of the system combined with enhanced text representation are higher, indicating that the rich acoustic information captured by enhancing text representation helps to improve the naturalness and fluency of the synthesized speech; these results show that the method proposed in the present invention has strong cross-language adaptability and can effectively improve TTS performance even for languages not involved in the pre-training process; by enhancing self-supervised features, the method of the present invention can effectively capture rich acoustic information without being affected by the language, thereby further improving the speech quality of the TTS system and enhancing the robustness and scalability of the method.

[0107] To verify whether the present invention effectively enhances the acoustic information contained in the text representation, an analysis experiment was conducted by evaluating the ability of the text representation to predict acoustic features; specifically, three typical acoustic features were predicted: pitch, energy, and duration; the architecture of the prediction network follows the predictor in FastSpeech2, where the output of the text encoder is used as the text representation and used to predict the corresponding acoustic features; the prediction model was trained and optimized for 150k steps until convergence, and the mean square error of the predicted value relative to the true value was taken as the evaluation index.

[0108] Table 3

[0109]

[0110] In Table 3, the "baseline model" refers to the TTS model that does not adopt the method proposed in the present invention, and "our method" refers to the method of self-supervised feature-enhanced speech synthesis based on the mutual information theory proposed in the present invention; the results in Table 3 show that the text representation generated by the method of the present invention has significantly enhanced the prediction ability for key acoustic features, including energy (pitch), duration, and pitch (energy); this indicates that the method of the present invention enables the text representation to encode richer and more comprehensive acoustic information; it is worth noting that the improvement in pitch prediction accuracy highlights a stronger ability to capture prosodic changes, while more accurate energy and duration estimates reflect a more accurate representation of speech rhythm and intensity; these findings further verify the effectiveness of the method of the present invention in compensating for the acoustic information in the text representation, ultimately contributing to more natural and expressive speech synthesis.

[0111] The existing methods involved in the table are described as follows:

[0112] Tacotron2 [ICASSP, 2018]: A non-autoregressive speech synthesis model that uses information from previous frames to predict the next frame for generation, but has a high time cost.

[0113] Wav2vec 2.0 [Arxiv, 2020]: A pre-trained self-supervised speech model.

[0114] Hifi-gan [NIPS, 2020]: A vocoder for speech synthesis, used to restore the mel spectrogram to a speech waveform diagram.

[0115] FastSpeech2 [ICLR 2021]: Proposes the modeling of pitch and energy to alleviate the information gap between text and speech.

[0116] HierSpeech [NIPS 2022]: End-to-end speech synthesis.

[0117] ParrotTTS [EACL 2024]: Replace the mel spectrogram in traditional speech synthesis with self-supervised speech features to verify the potential of self-supervised features in speech synthesis, but no further in-depth optimization is performed on self-supervised speech features.

[0118] The above embodiments are only preferred embodiments of the present invention and should not be considered as limiting the scope of implementation of the present invention. Any equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory, characterized in that The following steps are involved: 1) Self-supervised speech feature extraction: Extract self-supervised speech features from the pre-trained self-supervised speech model, use a bilinear interpolation algorithm to upsample the self-supervised speech features along the time axis, and align their length to the same number of frames as the Mel-spectrogram; 2) Construct an information bottleneck module based on mutual information maximization and minimization: The self-supervised speech feature S is encoded by the self-supervised encoder to obtain the self-supervised representation Z. The mutual information I(Z;M) between the self-supervised representation Z and the mel-spectrogram M is maximized through the optimization formula maxI(Z;M)-γI(Z;S), while minimizing the mutual information I(Z;S) between the self-supervised representation Z and the self-supervised speech feature S. This generates a compact and task-related self-supervised representation Z. γ is a weight hyperparameter used to balance the relationship between the two. The information bottleneck module based on mutual information maximization and minimization is composed of a self-supervised encoder and an optimization formula based on mutual information theory; the self-supervised encoder uses a non-causal convolutional network WaveNet to encode the extracted self-supervised speech feature S to obtain a compact and task-related self-supervised representation Z; The optimization formula based on mutual information theory is: maxI(Z;M)-γI(Z;S) Where I(Z;M) represents the mutual information between the self-supervised representation Z and the mel-spectrogram M. By maximizing the mutual information, the self-supervised representation Z is ensured to learn the acoustic information related to the speech synthesis task. I(Z;S) represents the mutual information between the self-supervised representation Z and the self-supervised speech feature S. By minimizing the mutual information, the compactness of the self-supervised representation Z is maintained and redundant information is eliminated. γ is a weight hyperparameter used to balance the relationship between the two. I(Z;M) and I(Z;S) are estimated using the MINE framework, which estimates mutual information through the following inequality: I(X;Y)≥E(x,y)[f(x,y)]-logE p (x)[E p (and)e f(x,y) ] Where I(X; Y) represents the mutual information between random variables X and Y, which measures the statistical dependence between the two; E(x, y)[f(x, y)] represents the expected value of the function f(x; y) under the joint distribution P(x; y) of X and Y; E p (x)[E p (y)e f(x,y) ] means that when X and Y are independently distributed, e f(x,y) ; use a neural network to approximate the value of the expression on the right side of the inequality sign, and then obtain an estimated value of the mutual information, wherein the neural network is composed of 3 layers of MLP; 3) Text representation enhancement: By optimizing the formula maxI(T; Z), the mutual information I(T; Z) between the text representation T and the self-supervised representation Z obtained in step 2) is maximized, so that the text representation T contains more information from the self-supervised representation and enhances the acoustic information of the text representation; where I(T; Z) is the mutual information between the text representation T and the self-supervised representation Z, estimated using the MINE framework; 4) Speech synthesis: A speech synthesis model is constructed based on the variational inference VAE network, the flow-based prior enhancement network VP-Flow and the post-processing network Post-net, and the optimization formulas of steps 2) and 3) are added to train the model; the trained speech synthesis model is used for inference to complete the self-supervised speech feature enhanced speech synthesis based on mutual information theory.

2. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 1), the specific steps of extracting self-supervised speech features are: using a pre-trained self-supervised speech model, inputting speech waveform data, and after encoding by the self-supervised speech model, extracting the output of the hidden layer of the self-supervised speech model that encodes the most acoustic features as the self-supervised speech feature S.

3. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 2, characterized in that The self-supervised speech model adopts wav2vec 2.0, HuBERT or BYOL-A, and the hidden layer selection is determined according to the model architecture and acoustic feature encoding capability.

4. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 3), the text representation enhancement step is as follows: the text encoder encodes the text sequence X into a text representation T, and then upsamples the text representation T by the length predicted by the length predictor to match the length of the mel-spectrogram; by maximizing the mutual information I(T; Z) between the text representation T and the self-supervised representation Z, the acoustic information of the text representation is enhanced; the specific optimization formula is: maxI(T;Z) Among them, I(T;Z) is the mutual information between the text representation T and the self-supervised representation Z, estimated using the MINE framework.

5. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 4), during model training: the text encoder encodes the text sequence into a text representation T, and upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel-spectrogram, and inputs it into the VAE decoder for Mel-spectrogram restoration; at the same time, the posterior distribution is processed by the flow-based prior enhancement network VP-Flow and converted into a standard Gaussian distribution; the post-processing network post-net converts the Mel-spectrogram samples into a standard Gaussian distribution; the flow-based prior enhancement network VP-Flow and the post-processing network post-net are both reversible. During training, the flow-based prior enhancement network VP-Flow converts the posterior distribution into a standard Gaussian distribution based on the text representation, and samples from the standard Gaussian distribution during inference and approximately restores the posterior distribution based on the text representation; and during training, the post-processing network post-net converts the real Mel-spectrogram into a standard Gaussian distribution based on the preliminary estimate of the Mel-spectrogram decoded by the VAE decoder, and samples from the standard Gaussian distribution during inference to convert the preliminary estimate of the Mel-spectrogram generated by the VAE decoder into a more delicate Mel-spectrogram.

6. A method for self-supervised speech feature enhancement speech synthesis based on mutual information theory as claimed in claim 5, characterized in that During inference, the text encoder encodes the text sequence into a text representation T, and upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the enhanced text representation T and the latent variables sampled from the standard Gaussian distribution are input into the flow-based prior enhancement network VP-Flow for inverse transformation to obtain the latent variables of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel-spectrogram; finally, the post-processing network post-net samples the latent variables from the standard Gaussian distribution and performs inverse transformation based on the preliminary estimate of the Mel-spectrogram to obtain the final Mel-spectrogram; the Mel-spectrogram is then converted into a speech waveform output with the help of a vocoder, thus completing the entire self-supervised speech feature enhanced speech synthesis process based on mutual information theory.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the self-supervised speech feature enhanced speech synthesis method based on mutual information theory as described in any one of claims 1 to 6 is implemented.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the self-supervised speech feature enhanced speech synthesis method based on mutual information theory as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Speech conversion method and device for semi-optimized cycle consistent adversarial networks (CycleGAN) model

    CN110246488A

  • Model training method and speech synthesis method based on semi-supervised knowledge distillation

    CN116092469A