Self-supervised speech feature enhanced speech synthesis method based on mutual information theory

By introducing a self-supervised speech feature enhancement method based on mutual information theory in the speech synthesis technology, the problems of insufficient integration of acoustic information and limited improvement in speech nature are solved, and more natural and high-quality speech synthesis is achieved.

CN119964551AActive Publication Date: 2025-05-09XIAMEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510430211.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing speech synthesis technology has problems such as insufficient integration of acoustic information and limited improvement in speech nature.

Method used

The self-supervised speech feature enhancement method based on mutual information theory is adopted to more comprehensively integrate acoustic information through self-supervised speech feature extraction, information bottleneck module design and text representation enhancement, thereby improving the naturalness and quality of speech synthesis.

Benefits of technology

It significantly improves the naturalness and quality of speech synthesis, enhances the predictive ability of text representation to key acoustic features, and has good cross-language adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964551A_ABST
    Figure CN119964551A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised speech feature enhanced speech synthesis method based on a mutual information theory, and relates to the technical field of speech synthesis. According to the method, self-supervised speech features are introduced to serve as acoustic supplementation of texts, an information bottleneck module based on mutual information maximization and minimization is designed, compact self-supervised representation related to tasks is extracted from the self-supervised speech features, and through maximizing mutual information between text representation and self-supervised representation, the self-supervised speech features are extracted. And acoustic information of text representation is enhanced, so that the naturalness and the quality of speech synthesis are improved. The method is excellent in performance in single-speaker and multi-speaker speech synthesis scenes, has good cross-language adaptability, and can effectively improve the speech synthesis quality in different language environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech synthesis technology, and in particular to a self-supervised speech feature enhanced speech synthesis method based on mutual information theory. The method is mainly used to improve the naturalness and quality of speech synthesis, and is also applicable to speech synthesis in multi-speaker scenarios and cross-language environments. Background Art

[0002] Speech synthesis technology aims to convert text information into speech output. It is widely used in intelligent voice assistants, audio book generation, voice broadcast systems and other fields. It is of great significance to improving human-computer interaction experience and information dissemination efficiency.

[0003] With the development of deep learning technology, speech synthesis methods based on neural networks have gradually become a research hotspot. For example, the autoregressive speech synthesis model Tacotron2 can generate natural speech, but it has problems such as low computational efficiency and difficulty in parallel processing; non-autoregressive models such as FastSpeech2 improve synthesis efficiency and quality by introducing additional acoustic information (such as pitch, energy, etc.), but are still limited to modeling specific acoustic features and lack modeling of broader acoustic information.

[0004] In recent years, self-supervised learning has achieved remarkable results in the field of speech representation learning. Self-supervised speech representation models such as wav2vec 2.0, HuBERT, and BYOL-A can learn rich acoustic information from large-scale unsupervised speech data, including duration, energy, pitch, rhythm, etc. However, the application of these self-supervised speech features containing rich acoustic information in speech synthesis is still in the exploratory stage. Based on this, the present invention proposes a method based on mutual information theory, which can effectively utilize self-supervised speech features to enhance speech synthesis.

[0005] In summary, existing speech synthesis technologies still have problems such as insufficient integration of acoustic information and limited improvement of speech naturalness. Therefore, there is an urgent need for a speech synthesis method that can effectively integrate self-supervised speech features to further improve the quality and applicability of speech synthesis. Summary of the invention

[0006] The purpose of the present invention is to provide a method for self-supervised speech feature enhanced speech synthesis based on mutual information theory to address the problems of insufficient information integration and limited speech naturalness in existing speech synthesis technologies. The method introduces self-supervised speech features and enhances text representation based on mutual information theory to more comprehensively integrate acoustic information, thereby improving the naturalness and quality of speech synthesis.

[0007] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions.

[0008] A self-supervised speech feature enhanced speech synthesis method based on mutual information theory comprises the following steps:

[0009] 1) Self-supervised speech feature extraction: Extract self-supervised speech features from the pre-trained self-supervised speech model, use a bilinear interpolation algorithm to upsample the self-supervised speech features along the time axis, and align their length to the same number of frames as the Mel spectrogram to ensure that the mutual information is calculated frame by frame with the Mel spectrogram. These self-supervised speech features can capture a variety of acoustic information including duration, energy, pitch, timbre, etc., as a supplement to the acoustic information of the text;

[0010] 2) Construct an information bottleneck module based on mutual information maximization and minimization: The self-supervised speech feature S is encoded by the self-supervised encoder to obtain the self-supervised representation Z, and the mutual information between the self-supervised representation Z and the mel-spectrogram M is maximized. , while minimizing the mutual information I(Z;S) between the self-supervised representation Z and the self-supervised speech feature S, generating a compact and task-related self-supervised representation Z;

[0011] 3) Text representation enhancement: By maximizing the mutual information between the text representation and the self-supervisory representation obtained in step 2), the text representation contains more information from the self-supervisory representation, that is, the acoustic information of the text representation is supplemented to improve the naturalness and quality of speech synthesis;

[0012] 4) Speech synthesis: Build a speech synthesis model. Use step 2) and step 3) to optimize the speech synthesis model during training. The trained speech synthesis model is used for reasoning to complete self-supervised speech feature enhanced speech synthesis based on mutual information theory. Step 2) can obtain a compact and task-related self-supervised representation. Step 3) can use the self-supervised representation to enhance the text representation. Therefore, during reasoning, the text representation contains more acoustic information, which can enhance the performance of the speech synthesis model and improve the naturalness and quality of speech synthesis. This completes self-supervised speech feature enhanced speech synthesis based on mutual information theory.

[0013] In step 1), the self-supervised speech model may adopt wav2vec 2.0, HuBERT, BYOL-A, etc., and the hidden layer selection is determined according to the model architecture and acoustic feature encoding capability.

[0014] The specific steps of extracting the self-supervised speech features may be: using a pre-trained self-supervised speech model, inputting speech waveform data, and after encoding by the self-supervised speech model, extracting the output of the hidden layer of the self-supervised speech model that encodes the most acoustic features as the self-supervised speech feature S.

[0015] Taking wav2vec 2.0 as an example, the specific steps are as follows: use the pre-trained wav2vec 2.0 as the self-supervised speech model, input the speech waveform data, and after encoding by wav2vec 2.0, select the output of the sixth hidden layer of the wav2vec 2.0 model as the self-supervised speech feature S. Since the sixth hidden layer of the model encodes more acoustic features than other layers, such as duration, pitch, energy, rhythm and other key information, which is important for the naturalness and quality improvement of subsequent speech synthesis, the output of the sixth hidden layer of the wav2vec 2.0 model is selected as the self-supervised speech feature.

[0016] In step 2), the information bottleneck module designed based on mutual information maximization and minimization is composed of a self-supervised encoder and an optimization formula based on mutual information theory; the self-supervised encoder adopts a non-causal convolutional network WaveNet to encode the extracted self-supervised speech feature S to obtain a self-supervised representation Z; the non-causal convolutional network WaveNet adopts 16 layers of causal convolution with a fixed expansion factor of 1, the convolution kernel size of each layer is 5, and the number of channels is 192; stable training is achieved through a dual-path residual structure (jump sum + residual accumulation) and a 10% probability Dropout;

[0017] The optimization formula based on mutual information theory is used to enhance the self-supervisory representation Z. The compact and task-related self-supervisory representation Z is extracted through the optimization formula. The optimization formula is:

[0018]

[0019] in, represents the mutual information between the self-supervised representation Z and the Mel-spectrogram M, and by maximizing the mutual information, it is ensured that the self-supervised representation Z learns the acoustic information related to the speech synthesis task; represents the mutual information between the self-supervised representation Z and the self-supervised speech feature S, and the compactness of the self-supervised representation Z is maintained by minimizing the mutual information, and redundant information is eliminated; γ is a weight hyperparameter, which is used to balance the relationship between the two; specifically, it is necessary to ensure the compactness of Z while ensuring that Z can encode task-related information, and the value range of the weight hyperparameter γ can be 0.8 to 1.2; in a preferred embodiment, the same weight is taken for the two mutual information optimizations, that is, γ is set to 1.0, so as to take into account both the compactness and task relevance of the self-supervised representation Z;

[0020] also, , Both are estimated using the MINE framework, which estimates mutual information through the following inequality:

[0021]

[0022] A neural network is used to approximate the value of the expression on the right side of the inequality sign, and then the estimated value of the mutual information is obtained. The neural network consists of 3 layers of MLP, and the hidden layer dimension is 256. In the formula, Represents the mutual information between random variables X and Y, measuring the statistical dependence between the two; Representation function The joint distribution of X and Y Expected value under It means that when X and Y are independently distributed, expectations.

[0023] Joint distribution sampling can be achieved by maintaining the temporal alignment of self-supervised features and mel-spectrograms, and independent distribution sampling can be achieved by randomly disrupting the temporal alignment of samples within a batch.

[0024] In step 3), in the text representation enhancement step, the mutual information between the text representation T and the self-supervised representation Z is maximized. , enhance the acoustic information of text representation; the specific steps can be:

[0025] The text encoder encodes the text sequence X into a text representation T, and then upsamples the text representation T to match the length of the Mel spectrogram by the length predicted by the length predictor; by maximizing the mutual information between the text representation T and the self-supervised representation Z , enhance the acoustic information of text representation; the specific optimization formula is:

[0026]

[0027] in, is the mutual information between the text representation T and the self-supervised representation Z, which is also estimated using the MINE framework. The specific structure of the neural network is consistent with step 2). By maximizing the mutual information, the text representation T can better capture the acoustic information in the self-supervised representation Z, thereby enhancing the acoustic information of the text representation T.

[0028] In step 4), a variational inference (VAE) network, a flow-based prior enhancement network (VP-Flow), and a post-processing network (post-net) are used as the basic architecture of the speech synthesis model. When training the speech synthesis model, the two optimization formulas proposed in steps 2) and 3) are added to complete the self-supervised speech feature enhanced speech synthesis based on mutual information theory;

[0029] During training: the text encoder encodes the text sequence into a text representation T, and then upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel-spectrogram, and inputs it into the VAE decoder for Mel-spectrogram restoration; at the same time, the posterior distribution is processed by VP-Flow and converted into a standard Gaussian distribution; and post-net will also convert the Mel-spectrogram samples into a standard Gaussian distribution; VP-Flow and post-net are both reversible. During training, VP-Flow converts the posterior distribution into a standard Gaussian distribution based on the text representation. During inference, it can sample from the standard Gaussian distribution and approximately restore the posterior distribution based on the text representation; while post-net training converts the real Mel-spectrogram into a standard Gaussian distribution based on the preliminary estimate of the Mel-spectrogram decoded by the VAE decoder. During inference, it can sample from the standard Gaussian distribution to convert the preliminary estimate of the Mel-spectrogram generated by the VAE decoder into a more delicate Mel-spectrogram;

[0030] During inference, the text encoder encodes the text sequence into a text representation T, and then upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the enhanced text representation T and the latent variables sampled from the standard Gaussian distribution are input into VP-Flow for inverse transformation to obtain the latent variables of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel-spectrogram; finally, Post-net samples the latent variables from the standard Gaussian distribution and performs an inverse transformation based on the preliminary estimate of the Mel-spectrogram to obtain the final Mel-spectrogram; the Mel-spectrogram is then converted into a speech waveform output with the help of a vocoder, thus completing the entire self-supervised speech feature enhanced speech synthesis process based on mutual information theory.

[0031] Compared with the prior art, the advantages and outstanding technical effects of the present invention are:

[0032] 1. Introduction of self-supervisory features: By introducing self-supervisory speech features, the present invention can capture a variety of acoustic information including duration, energy, pitch, etc. and the relationship between the information. Compared with the prior art method that only focuses on specific acoustic features, the self-supervisory features introduced in the present invention more comprehensively integrate acoustic information, which helps to improve the naturalness and quality of synthesized speech.

[0033] 2. Use of self-supervised features: In order to efficiently integrate self-supervised speech features into the TTS model, the present invention proposes an optimization method based on mutual information theory to extract compact, task-related self-supervised representations from self-supervised speech features, and use these self-supervised representations to enhance text representation.

[0034] 3. The present invention conducts experiments on Chinese single-speaker and multi-speaker datasets and compares with multiple speech synthesis methods. The mean opinion score (MOS) is improved by 0.09 and 0.21 respectively compared with the best comparison model.

[0035] 4. Good cross-language adaptability: The method of the present invention has good cross-language adaptability. Experiments have shown that even in a low-resource language (Hokkien) environment that is not covered by the pre-training model, it can still effectively improve the quality of speech synthesis and further expand the scope of application.

[0036] 5. The text representation generated by the present invention has significantly enhanced predictive power for key acoustic features, including energy, duration, and pitch, which indicates that the text representation encodes richer and more comprehensive acoustic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of the overall framework flow.

[0038] Figure 2 This is an example of the model training process of the present invention.

[0039] Figure 3 This is the optimization process of the information bottleneck module designed by the present invention for the self-supervisory representation. It shows the information change process contained in the self-supervisory representation.

[0040] Figure 4 It is the text representation enhancement process. It shows the process of enhancing text representation by maximizing mutual information and supplementing the lack of acoustic information in text representation.

[0041] Figure 5 It is the model reasoning process of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in the following embodiments in conjunction with the accompanying drawings.

[0043] The embodiments of the present invention introduce self-supervised speech features and enhance text representation based on mutual information theory to more comprehensively integrate acoustic information, thereby improving the naturalness and quality of speech synthesis.

[0044] like Figure 1 As shown, the specific steps of the embodiment of the present invention are as follows:

[0045] 1) Self-supervised speech feature extraction: A pre-trained self-supervised speech model is selected. This embodiment uses wav2vec2.0, and a wav2vec 2.0 model that has been pre-trained on large-scale unsupervised speech data is loaded from the pre-trained model library. The original speech waveform data is pre-processed to ensure that the data format meets the input requirements of the wav2vec 2.0 model; in this embodiment, the original speech data is sampled at a unified rate of 22kHz, and the audio signal amplitude is normalized to the range of [-1,1]; after preprocessing, the speech data is input into the loaded wav2vec 2.0 model with an appropriate batch size.

[0046] Extract the hidden layer output as the self-supervised speech feature. This embodiment takes wav2vec 2.0 as an example. The wav2vec2.0 model encodes the input speech waveform data. During the encoding process, the output of the hidden layer of the model (the sixth hidden layer of wav2vec2.0) is extracted as the self-supervised speech feature. The self-supervised speech feature is upsampled along the time axis using a bilinear interpolation algorithm, and its length is aligned to the same number of frames as the mel spectrogram. The non-causal convolutional network WaveNet is used as a self-supervised encoder to encode the extracted self-supervised speech feature S to obtain the self-supervised representation Z. WaveNet uses 16 layers of causal convolution with a fixed expansion factor of 1, each layer of convolution kernel size is 5, and the number of channels is 192. Stable training is achieved through a dual-path residual structure (jump sum + residual accumulation) and a 10% probability Dropout.

[0047] In particular, based on the analysis of the outputs of different hidden layers of wav2vec 2.0 in “What all do audio transformer models hear? probingacoustic representations for language delivery and its structure” by Jui Shah et al. and published in “arXiv preprint” (arXiv:2101.00387, 2021), the output of the sixth hidden layer was selected as the self-supervised speech feature because the sixth hidden layer of the model encodes more acoustic features than other layers, such as duration, pitch, energy, and rhythm, which are key information that plays an important role in improving the naturalness and quality of subsequent speech synthesis.

[0048] Select a pre-trained self-supervised speech model (such as wav2vec 2.0, HuBERT, BYOL-A, etc.) and load the pre-trained model from the pre-trained model library. The hidden layer selection of different models should be determined according to their architecture and acoustic feature encoding capabilities: for the HuBERT model, the output of its middle layer can be selected, depending on the model's encoding efficiency of acoustic information; for the BYOL-A model, the optimal hidden layer position can be determined through ablation experiments.

[0049] The self-supervised speech features are upsampled by a bilinear interpolation algorithm to align their length with the number of Mel-spectrogram frames to ensure frame-by-frame matching in subsequent mutual information calculations.

[0050] 2) Information bottleneck module design:

[0051] The information bottleneck module is one of the key components of the present invention, which is designed based on the principle of maximizing and minimizing mutual information, and aims to extract a compact self-supervised representation that is closely related to the speech synthesis task from the self-supervised speech features.

[0052] The overall architecture of the information bottleneck module consists of a self-supervised encoder and two optimization branches. The self-supervised encoder can use a non-causal convolutional network WaveNet to encode the extracted self-supervised speech feature S to obtain a self-supervised representation Z; the two optimization branches are used to process different mutual information respectively, so as to optimize the self-supervised representation.

[0053] Furthermore, the non-causal convolutional network WaveNet adopts 16 layers of causal convolution with a fixed dilation factor of 1, a convolution kernel size of 5 in each layer, and 192 channels, and achieves stable training through a dual-path residual structure (jump summation + residual accumulation) and 10% probability Dropout.

[0054] Furthermore, the two optimization branches are used to process the mutual information between the self-supervised representation Z and the Mel-spectrogram M respectively. , and the mutual information between the self-supervised representation Z and the self-supervised feature S ; Each branch constructs a 3-layer MLP neural network structure with a hidden layer dimension of 256 to achieve effective estimation and optimization of mutual information.

[0055] The self-supervised representation Z is a compact feature processed by the information bottleneck module, which is the output of the self-supervised speech feature S after being encoded by the information bottleneck module. and minimize Generate. The information bottleneck module is designed based on mutual information maximization and minimization, and its optimization formula is:

[0056]

[0057] in, represents the mutual information between the self-supervisory representation Z and the Mel-spectrogram M, and by maximizing the mutual information, it is ensured that the self-supervisory representation Z learns the acoustic information related to the speech synthesis task; It represents the mutual information between the self-supervised representation Z and the self-supervised feature S. By minimizing the mutual information, the compactness of the self-supervised representation Z is maintained and redundant information is eliminated. γ is a weight hyperparameter used to balance the relationship between the two.

[0058] Specifically, it is necessary to ensure the compactness of the self-supervised representation Z and to ensure that the self-supervised representation Z can encode task-related information. Therefore, in this embodiment, γ can be set to 1.0 to take into account the compactness and task relevance of the self-supervised representation Z. When the value of γ is too large, the model will pay too much attention to the mutual information between the self-supervised representation Z and the mel-spectrogram M, resulting in the self-supervised representation Z may contain too much redundant information and cannot effectively maintain compactness. When the value of γ is too small, the model will over-emphasize the mutual information between the self-supervised representation Z and the self-supervised feature S, causing the self-supervised representation Z to lose some key acoustic information closely related to the speech synthesis task. Therefore, the two mutual information are given the same weight, that is, γ is set to 1.0.

[0059] also, , Both are estimated using the MINE framework, and MINE estimates the mutual information through the following inequality;

[0060]

[0061] A neural network is used to approximate the value of the expression on the right side of the inequality sign, thereby obtaining an estimate of the mutual information. Represents the mutual information between random variables X and Y, measuring the statistical dependence between the two; Representation function The joint distribution of X and Y Expected value under It means that when X and Y are independently distributed, ; use a neural network to approximate the value of the expression on the right side of the inequality sign, and then get the estimated value of the mutual information. The neural network consists of 3 layers of MLP, the hidden layer dimension is 256, and the RELU activation function is used. The optimization goal is:

[0062]

[0063] The optimizer uses the Adam optimizer, where Take 0.9, Take 0.98, Take 10 to the negative 9th power, and the learning rate decay refers to "Attention is all you need" by Vaswani et al. published in "Advances in Neural Information Processing Systems" (30, 2017). The number of training iterations and batch size are consistent with the speech synthesis model. Taking the Chinese dataset CSMSC as an example: the number of iterations is 300K times and the batch size is 64.

[0064] 3) Text representation enhancement:

[0065] In order to enable the text representation to capture acoustic information more effectively, the text representation encoded by the text encoder is enhanced. The text encoder adopts an encoder based on the Transformer architecture, which can effectively capture the long-distance dependencies in the text sequence through the self-attention mechanism and extract rich semantic features.

[0066] First, the text encoder is initialized and model parameters are set according to task requirements, such as hidden layer dimension, number of attention heads, etc. The text sequence X is preprocessed according to the format required by the encoder, including word segmentation, adding position encoding, etc., and then input into the text encoder. After multiple layers of processing by the encoder, the text encoder encodes the text sequence X into a text representation T. At this time, the text representation T contains the semantic and structural information of the text sequence X, but does not fully integrate the acoustic information.

[0067] Since the length of the text representation T is usually inconsistent with the length of the Mel-spectrogram, a length predictor is needed to predict the Mel-spectrogram frame length corresponding to the text representation to upsample the text representation T to achieve the matching of the two lengths; by maximizing the mutual information between the text representation T and the self-supervised representation Z , enhance the acoustic information of text representation; the network parameters are continuously optimized through training to ensure that the upsampled text representation can not only retain the original text features, but also adapt to the Mel spectrogram in length, and can contain more acoustic information. The specific optimization formula is:

[0068]

[0069] in is the mutual information between the text representation T and the self-supervised representation Z. The mutual information is also estimated using the MINE framework. The specific structure of the neural network in the MINE framework is consistent with the neural network structure when estimating the mutual information in step 2). By maximizing the mutual information, the text representation T can better capture the acoustic information in the self-supervised representation Z, thereby enhancing the acoustic information of the text representation. When the model converges, the obtained text representation T contains rich acoustic information, laying a solid foundation for high-quality speech synthesis.

[0070] 4) Speech Synthesis:

[0071] The variational inference (VAE) network: variational encoder and variational decoder, flow-based prior enhancement network (VP-Flow) and post-processing network (Post-net) are used as the basic architecture of the speech synthesis model. When training the speech synthesis model, the two optimization formulas proposed in steps 2) and 3) are added to complete the self-supervised speech feature enhanced speech synthesis based on mutual information theory.

[0072] Specific application examples are given below.

[0073] Example 1: Single speaker synthesis

[0074] 1. Data preparation

[0075] Taking the single-speaker Chinese dataset CSMSC as an example, the single-speaker Chinese dataset CSMSC is used, which contains 10,000 speech samples with a total length of 12 hours; the dataset is divided into training set, validation set and test set, where the training set is used for model training, the validation set is used for model tuning, and the test set is used to evaluate model performance.

[0076] 2. Self-supervised feature extraction

[0077] This embodiment takes wav2vec 2.0 as an example, unifies the sampling rate of the original speech data to 22kHz, and normalizes the amplitude of the audio signal to the range of [-1, 1]; after preprocessing, the speech data is input into the loaded wav2vec 2.0 model in an appropriate batch size, and the wav2vec 2.0 model encodes the input speech waveform data. During the encoding process, the output of the hidden layer of the model (the sixth hidden layer of wav2vec 2.0) is extracted as the self-supervised speech feature, and the self-supervised speech feature is upsampled along the time axis dimension using a bilinear interpolation algorithm, and its length is aligned to the same number of frames as the mel spectrogram; a non-causal convolutional network WaveNet is used as a self-supervised encoder to extract the self-supervised speech feature S Encoding is performed to obtain the self-supervised representation Z; WaveNet uses 16 layers of causal convolution with a fixed dilation factor of 1, the convolution kernel size of each layer is 5, and the number of channels is 192; stable training is achieved through a dual-path residual structure (jump summation + residual accumulation) and 10% probability Dropout.

[0078] 3. Self-supervised representation enhancement

[0079] A compact and task-related self-supervised representation S is extracted through the information bottleneck module; the overall architecture of the information bottleneck module consists of a self-supervised encoder and two optimization branches, and a non-causal convolutional network WaveNet is used as a self-supervised encoder to encode the extracted self-supervised speech feature S to obtain a self-supervised representation Z; the non-causal convolutional network WaveNet adopts 16 layers of causal convolution with a fixed expansion factor of 1, a convolution kernel size of 5 in each layer, and 192 channels, and stable training is achieved through a dual-path residual structure (jump sum + residual accumulation) and 10% probability Dropout; the two optimization branches are used to process the mutual information between the self-supervised representation Z and the Mel-spectrogram M, and the mutual information between the self-supervised representation Z and the self-supervised feature S, respectively; each branch constructs a 3-layer MLP, a neural network structure with a hidden layer dimension of 256 to achieve effective estimation and optimization of mutual information; the specific optimization process is as follows: maximize the mutual information between the self-supervised representation Z and the Mel-spectrogram M , ensuring that the extracted representation contains information relevant to the speech synthesis task; calculating the mutual information between the self-supervised representation Z and the self-supervised feature S , by minimizing the mutual information, the compactness of the extracted self-supervisory representation is maintained; the specific optimization formula is as follows:

[0080] Get the final self-supervised representation Z; Figure 3 Show the information conversion process of the self-supervised representation Z: Through the optimization of the mutual information loss function, the self-supervised representation Z encodes the necessary information required to reconstruct the mel-spectrogram M.

[0081] 4. Text Representation Enhancement

[0082] In order to enable the text representation to capture acoustic information more effectively, the text representation encoded by the text encoder is enhanced. The text encoder adopts an encoder based on the Transformer architecture, which can effectively capture the long-distance dependencies in the text sequence through the self-attention mechanism and extract rich semantic features.

[0083] First, the text encoder is initialized, which consists of 4 Transforme architecture encoders, with a hidden layer dimension of 192 and 2 attention heads. The text sequence X is preprocessed according to the format required by the encoder, including word segmentation, adding position encoding and other operations, and then input into the text encoder. After multiple layers of processing by the encoder, the text encoder encodes the text sequence X into a text representation T. At this time, the text representation T contains the semantic and structural information of the text sequence X, but does not fully integrate the acoustic information.

[0084] Since the length of the text representation T is usually inconsistent with the length of the Mel-spectrogram, it is necessary to introduce a length predictor to predict the Mel-spectrogram frame length corresponding to the text representation to upsample the text representation T to achieve the matching of the two lengths; by maximizing the mutual information between the text representation T and the self-supervised representation Z , enhancing the acoustic information of text representation.

[0085] By maximizing the mutual information between the text representation T and the self-supervised representation Z , enhance the acoustic information of text representation; the specific optimization formula is:

[0086]

[0087] in, is the mutual information between the text representation T and the self-supervisory representation Z. By maximizing the mutual information, the text representation T can better capture the acoustic information in the self-supervisory representation Z, thereby enhancing the acoustic information of the text representation. When the model converges, the obtained text representation T contains rich acoustic information, laying a solid foundation for high-quality speech synthesis. Figure 4 The change process of the information contained in the text representation T is described, that is, more acoustic information from the self-supervisory representation Z is supplemented; the enhanced text representation T is used in the subsequent speech synthesis process, which can improve the naturalness and quality of speech synthesis.

[0088] 5. Speech Synthesis

[0089] Speech synthesis uses the variational inference (VAE) network, flow-based prior enhancement network (VP-Flow), and post-processing network (Post-net) in "Portaspeech: Portable and high-quality generative text-to-speech" by Ren Yi et al., published in "Advances in Neural Information Processing Systems" (34, 2021) as the basic architecture of the speech synthesis model. Figure 2 As shown in the figure, the variational encoder receives the Mel spectrum input, and performs temporal downsampling through five layers of causal convolution (convolution kernel 5, step size 4), gradually expanding the number of channels from 80 to 192, and connecting to the bidirectional LSTM at the end to extract contextual features, and finally mapping them to the mean and variance parameters of 16-dimensional latent variables. After the latent variables are reparameterized and sampled, they are input into the flow-based prior enhancement network (VP-Flow), which contains a four-step reversible transformation chain, each step performs channel permutation and affine coupling operations in turn, where the coupling layer generates a dynamic scaling factor through a three-layer fully connected network (hidden unit 64) to convert the variational encoder output into a standard Gaussian distribution. The latent variables are input into the variational decoder, and the coarse-grained Mel spectrum is reconstructed through a four-layer non-causal WaveNet (residual convolution stack with increasing dilation coefficient), and the original temporal resolution is restored using transposed convolution (convolution kernel 5, step size 4). The flow network refines the output of the variational decoder. The first layer multiplies the 80-channel Mel spectrum to 160 channels through the time axis compression operation, followed by a twelve-step flow transformation. Each three-step flow transformation shares a parameter group, each group contains activation normalization, four-group reversible neighbor convolution (orthogonal matrix initialization weights) and coupling blocks, where the coupling block uses a three-layer WaveNet (convolution kernel 3, hidden unit 192) to generate spectral correction factors, and finally restores the time dimension through decompression operation and outputs high-fidelity Mel spectrum. In the key parameter configuration of the model, the encoder WaveNet contains eight layers of dilated convolution, and the post-processing network uses three layers of convolution for each coupling block. The latent variable optimization process achieves distribution enhancement through a four-step reversible transformation. All modules use dynamic conditional injection technology to map text features into convolution weight modulation signals to improve content-rhythm alignment capabilities. This architecture significantly improves inference efficiency while ensuring synthesis quality through non-autoregressive parallel generation design (dilated convolution stack single feedforward to complete timing modeling) and latent variable explicit optimization strategy (VP-Flow reversible transformation chain).

[0090] When training the speech synthesis model, the two optimization formulas proposed in step 3 and step 4 are added to complete the self-supervised speech feature enhanced speech synthesis based on mutual information theory.

[0091] like Figure 2 As shown in the figure, during training: the text encoder encodes the text sequence into a text representation T, and then upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel-spectrogram and inputs it into the VAE decoder for Mel-spectrogram restoration; at the same time, the posterior distribution is processed by VP-Flow and then compared with the standard Gaussian distribution through KL divergence. The posterior distribution is processed by zooming in and converting it to a standard Gaussian distribution. The post-processing network (post-net) also converts the Mel spectrum samples to a standard Gaussian distribution. Both VP-Flow and post-net are reversible. During training, VP-Flow processes the posterior distribution based on the text representation and compares it with the standard Gaussian distribution through KL divergence. During inference, we can sample from the standard Gaussian distribution and approximately restore the posterior distribution based on the text representation. During post-net training, we convert the real Mel-spectrogram into a standard Gaussian distribution based on the preliminary estimate of the Mel-spectrogram decoded by the VAE decoder. During inference, we can sample from the standard Gaussian distribution and convert the preliminary estimate of the Mel-spectrogram generated by the VAE decoder into a more delicate Mel-spectrogram.

[0092] like Figure 5 As shown in Figure 1, during inference, the text encoder encodes the text sequence into a text representation T, and then upsamples it to match the length of the Mel spectrum frame predicted by the length predictor for each text representation; since the text representation T has been enhanced during training, the text representation T at this time has encoded rich acoustic information, and the enhanced text representation T and the latent variable sampled from the prior distribution - the standard Gaussian distribution are combined. The input is inversely transformed into VP-Flow to obtain the hidden variables of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel spectrogram; finally, Post-net samples the hidden variables from the prior distribution - the standard Gaussian distribution The initial estimate of the Mel-spectrogram is used as a condition for inverse transformation to obtain the final Mel-spectrogram; the Mel-spectrogram is then converted into a speech waveform output with the help of a vocoder, thus completing the entire self-supervised speech feature enhanced speech synthesis process based on mutual information theory.

[0093] Example 2: Multi-speaker synthesis

[0094] Similar to Example 1, the difference lies in step 1, data preparation: taking the multi-speaker Chinese dataset AISHELL-3 as an example, the multi-speaker Chinese dataset AISHELL-3 is used, which contains 88,035 speech samples from 218 speakers with a total length of 85 hours; the dataset is divided into a training set, a validation set and a test set, where the training set is used for model training, the validation set is used for model tuning, and the test set is used to evaluate model performance.

[0095] Steps 2-5 are the same as those in Example 1 for a single speaker.

[0096] Experiments are conducted on two datasets: the single-speaker dataset CSMSC and the multi-speaker dataset AISHELL-3; the CSMSC dataset contains 10,000 utterances from a female speaker with a total audio duration of 12 hours; in contrast, the AISHELL-3 dataset is a multi-speaker dataset containing 88,035 utterances from 218 speakers with a total audio duration of 85 hours; for each dataset, 100 utterances are designated as the validation set, 500 utterances as the test set, and the remaining data is used to train the model.

[0097] Mean opinion score (MOS) subjective testing was performed and the results were reported with 95% confidence intervals.

[0098] The speech synthesized by the method proposed in the present invention is compared with other systems in terms of subjective evaluation score (MOS), including 1) GT, real audio; 2) GT (Mel + HiFi-GAN), in which real Mel spectrograms are extracted and converted into audio using HiFi-GAN vocoder; 3) Tacotron 2, an autoregressive TTS model, in which the speaker embedding is connected to the output of the text encoder during the training process of the multi-speaker TTS system; 4) FastSpeech 2, a non-autoregressive TTS model, and the speaker embedding setting of the multi-speaker TTS training is consistent with that of Tacotron2; 5) HierSpeech, an end-to-end TTS model, and the multi-speaker TTS training setting is consistent with that of Tacotron2; 6) ParrotTTS, which replaces the Mel spectrogram in traditional speech synthesis with self-supervised speech features, and the multi-speaker TTS training setting is consistent with that of Tacotron2.

[0099] Table 1

[0100]

[0101] As can be seen from Table 1, the model of the present invention achieves higher MOS scores than other models on both datasets; it is worth noting that on the multi-speaker dataset, the method of the present invention is significantly better than the existing TTS model, and is only slightly behind the theoretical performance ceiling (GT + HiFi-GAN); it is speculated that this is because in the context of large datasets, the MINE mutual information estimation framework used in the present invention becomes more accurate and stable, which is beneficial to the optimization process of the present invention for text representation.

[0102] Example 3: Low-resource language expansion

[0103] In order to further verify the generalization ability of the method of the present invention, additional experiments were conducted in a low-resource language environment, and the Minnan dialect was selected as the data set. It is worth noting that the pre-trained self-supervised speech model was not trained on the Minnan dialect; this setting can evaluate the performance of the method proposed in the present invention when processing languages ​​not included in the pre-training corpus, thereby further testing its cross-language generalization ability; in the experiment, SuiSiann was selected as the Minnan dialect data set, where the data set size was 3467 utterances with a total duration of nearly 5 hours. 100 utterances were designated as the validation set and 500 utterances as the test set. The remaining data was used to train the model, and the MOS score was compared with the comparison method in the experimental part of Table 1.

[0104] Table 2

[0105]

[0106] The experimental results in Table 2 show that although the pre-trained model is not trained on the Minnan dialect, the method of the present invention still significantly improves the quality of speech synthesis; specifically, the MOS score of the system combined with the enhanced text representation is higher, indicating that the rich acoustic information captured by the enhanced text representation helps to improve the naturalness and fluency of the synthesized speech; These results show that the method proposed in the present invention has strong cross-language adaptability and can effectively improve the TTS performance even for languages ​​that are not involved in the pre-training process; By enhancing the self-supervised features, the method of the present invention can effectively capture rich acoustic information without being affected by the language, thereby further improving the speech quality of the TTS system and enhancing the robustness and scalability of the method.

[0107] In order to verify whether the present invention effectively enhances the acoustic information contained in the text representation, an analytical experiment was conducted by evaluating the ability of the text representation to predict acoustic features. Specifically, three typical acoustic features were predicted: pitch, energy, and duration. The architecture of the prediction network followed the predictor in FastSpeech2, where the output of the text encoder was used as the text representation and used to predict the corresponding acoustic features. The prediction model was trained and optimized for 150k steps until convergence, and the mean square error of the predicted value relative to the true value was taken as the evaluation indicator.

[0108] Table 3

[0109]

[0110] In Table 3, "baseline model" refers to the TTS model that does not adopt the method proposed in the present invention, and "our method" refers to the method of self-supervised feature-enhanced speech synthesis based on mutual information theory proposed in the present invention; the results in Table 3 show that the text representation generated by the method of the present invention has significantly enhanced prediction capabilities for key acoustic features, including energy (pitch), duration (duration) and pitch (energy); this shows that the method of the present invention enables text representation to encode richer and more comprehensive acoustic information; it is worth noting that the improvement in pitch prediction accuracy highlights a stronger ability to capture prosodic changes, while more accurate energy and duration estimates reflect more accurate speech rhythm and intensity representations; these findings further verify the effectiveness of the method of the present invention in compensating for acoustic information in text representation, and ultimately contribute to more natural and expressive speech synthesis.

[0111] The existing methods involved in the table are described as follows:

[0112] Tacotron2 [ICASSP, 2018]: A non-autoregressive speech synthesis model that uses information from previous frames to predict the next frame, but it is time-consuming and expensive.

[0113] Wav2vec 2.0 [Arxiv, 2020]: A pre-trained self-supervised speech model.

[0114] Hifi-gan [NIPS, 2020]: A vocoder for speech synthesis that converts mel-spectrograms into speech waveforms.

[0115] FastSpeech2 [ICLR 2021]: Proposes modeling of pitch and energy to alleviate the information gap between text and speech.

[0116] HierSpeech [NIPS 2022]: End-to-end speech synthesis.

[0117] ParrotTTS [EACL 2024]: Replaces the mel-spectrogram in traditional speech synthesis with self-supervised speech features, verifying the potential of self-supervised features in speech synthesis, but does not make further optimizations for self-supervised speech features.

[0118] The above embodiments are only preferred embodiments of the present invention and cannot be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.

Claims

1. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory, characterized in that The following steps are involved: 1) Self-supervised speech feature extraction: Extract self-supervised speech features from the pre-trained self-supervised speech model, use a bilinear interpolation algorithm to upsample the self-supervised speech features along the time axis, and align their length to the same number of frames as the Mel-spectrogram; 2) Construct an information bottleneck module based on mutual information maximization and minimization: The self-supervised speech feature S is encoded by the self-supervised encoder to obtain the self-supervised representation Z. Maximize the mutual information between the self-supervised representation Z and the mel-spectrogram M , while minimizing the mutual information between the self-supervised representation Z and the self-supervised speech feature S , generates a compact and task-related self-supervised representation Z, γ is a weight hyperparameter used to balance the relationship between the two; 3) Text representation enhancement: by optimizing the formula Maximize the mutual information between the text representation T and the self-supervised representation Z obtained in step 2) , so that the text representation T contains more information from the self-supervised representation and enhances the acoustic information of the text representation; 4) Speech synthesis: A speech synthesis model is constructed based on the variational inference VAE network, the flow-based prior enhancement network VP-Flow, and the post-processing network Post-net. The optimization formulas of steps 2) and 3) are added to train the model. The trained speech synthesis model is used for inference to complete self-supervised speech feature enhanced speech synthesis based on mutual information theory.

2. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 1), the specific steps of extracting the self-supervised speech feature are: using a pre-trained self-supervised speech model, inputting speech waveform data, and after encoding by the self-supervised speech model, extracting the output of the hidden layer of the self-supervised speech model that encodes the most acoustic features as the self-supervised speech feature S.

3. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 2, characterized in that The self-supervised speech model adopts wav2vec 2.0, HuBERT or BYOL-A, and the hidden layer selection is determined according to the model architecture and acoustic feature encoding capability.

4. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 2), the information bottleneck module based on mutual information maximization and minimization consists of a self-supervised encoder and an optimization formula based on mutual information theory; the self-supervised encoder uses a non-causal convolutional network WaveNet to encode the extracted self-supervised speech feature S to obtain a compact and task-related self-supervised representation Z.

5. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 4, characterized in that The optimization formula based on mutual information theory is: in, represents the mutual information between the self-supervised representation Z and the Mel-spectrogram M, and by maximizing the mutual information, it is ensured that the self-supervised representation Z learns the acoustic information related to the speech synthesis task; represents the mutual information between the self-supervised representation Z and the self-supervised speech feature S. By minimizing the mutual information, the compactness of the self-supervised representation Z is maintained and redundant information is eliminated. γ is a weight hyperparameter used to balance the relationship between the two. , Both are estimated using the MINE framework, which estimates mutual information through the following inequality: in, Represents the mutual information between random variables X and Y, measuring the statistical dependence between the two; Representation function The joint distribution of X and Y Expected value under It means that when X and Y are independently distributed, ; use a neural network to approximate the value of the expression on the right side of the inequality sign, and then obtain an estimated value of the mutual information, wherein the neural network is composed of 3 layers of MLP.

6. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 3), the text representation enhancement step is as follows: the text encoder encodes the text sequence X into a text representation T, and then upsamples the text representation T by the length predicted by the length predictor to match the length of the mel-spectrogram; by maximizing the mutual information between the text representation T and the self-supervised representation Z , enhance the acoustic information of text representation; the specific optimization formula is: in, is the mutual information between the text representation T and the self-supervised representation Z, estimated using the MINE framework.

7. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 1, characterized in that In step 4), when the model is trained: the text encoder encodes the text sequence into a text representation T, and upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the VAE encoder generates a latent variable of the posterior distribution based on the Mel-spectrogram, and inputs it into the VAE decoder for Mel-spectrogram restoration; at the same time, the posterior distribution is processed by the flow-based prior enhancement network VP-Flow and converted into a standard Gaussian distribution; the post-processing network post-net converts the Mel-spectrogram samples into a standard Gaussian distribution; the flow-based prior enhancement network VP-Flow and the post-processing network post-net are both reversible, and the flow-based prior enhancement network VP-Flow is used during training. Flow converts the posterior distribution into a standard Gaussian distribution conditioned on the text representation, samples from the standard Gaussian distribution during inference and approximately restores the posterior distribution based on the text representation; while during post-net training, the post-processing network converts the real Mel-spectrogram into a standard Gaussian distribution conditioned on the preliminary estimate of the Mel-spectrogram decoded by the VAE decoder, and samples from the standard Gaussian distribution during inference to convert the preliminary estimate of the Mel-spectrogram generated by the VAE decoder into a more delicate Mel-spectrogram.

8. A self-supervised speech feature enhanced speech synthesis method based on mutual information theory as claimed in claim 7, characterized in that During inference, the text encoder encodes the text sequence into a text representation T, and upsamples the Mel-spectrogram frame length of each text representation predicted by the length predictor to match the length of the Mel-spectrogram; the enhanced text representation T and the latent variables sampled from the standard Gaussian distribution are input into the flow-based prior enhancement network VP-Flow for inverse transformation to obtain the latent variables of the posterior distribution, and then the VAE decoder generates a preliminary estimate of the Mel-spectrogram; finally, the post-processing network post-net samples the latent variables from the standard Gaussian distribution and performs an inverse transformation based on the preliminary estimate of the Mel-spectrogram to obtain the final Mel-spectrogram; the Mel-spectrogram is then converted into a speech waveform output with the help of a vocoder, thus completing the entire self-supervised speech feature enhanced speech synthesis process based on mutual information theory.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the self-supervised speech feature enhanced speech synthesis method based on mutual information theory as described in any one of claims 1 to 8 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the self-supervised speech feature enhanced speech synthesis method based on mutual information theory as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Speech conversion method and device for semi-optimized cycle consistent adversarial networks (CycleGAN) model

    CN110246488A

  • Speech synthesis method and device

    CN113539236A

  • Model training method and speech synthesis method based on semi-supervised knowledge distillation

    CN116092469A

  • Method and apparatus for synthesizing unified voice wave based on self-supervised learning

    US20240347037A1