Speech synthesis method and device based on decoupled VQ-VAE, equipment and storage medium
By introducing a global reference encoder and a decoupled VQ-VAE single-codebook speech codec into the speech codec, the efficiency and robustness issues caused by the multi-codebook structure are resolved, and more efficient speech reconstruction results are achieved.
Patent Information
- Application Number
- CN202411531770.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-10-29
AI Technical Summary
In existing technologies, speech codecs use a multi-codebook structure, which requires the language model to predict multiple discrete sequences, affecting the efficiency and robustness of the codec.
A speech synthesis method based on decoupled VQ-VAE is adopted. The time-invariant features are decoupled through a global reference encoder, and a single codebook speech codec is used to decouple the speech signal into a discrete sequence with time-invariant features and rich speech information, thus avoiding multi-sequence prediction.
It improves the efficiency and robustness of speech codecs, and achieves better speech reconstruction quality, especially under low bandwidth conditions.
Smart Images

Figure CN119252225B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and in particular to a speech synthesis method, apparatus, device, and storage medium based on decoupled VQ-VAE. Background Technology
[0002] Speech codecs aim to convert speech signals into compact discrete representations and reconstruct the original audio when needed. Speech codecs have wide applications in speech communication, speech storage, and speech synthesis. In Large Language Model-Based Speech Synthesis (LLM-TTS) systems, the speech codec is responsible for converting speech signals into discrete representations (tokens), enabling the large language model to process speech like text and reconstructing the discrete representations generated by the large language model into high-quality speech waveforms.
[0003] Currently, most mainstream speech codecs in the industry adopt a multi-codebook structure, requiring the language model to predict multiple discrete sequences, which seriously affects the efficiency and robustness of the codec. Summary of the Invention
[0004] This invention provides a speech synthesis method, apparatus, device, and storage medium based on decoupled VQ-VAE, to solve the technical problem in the prior art where the speech codec adopts a multi-codebook structure and the language model needs to predict multiple discrete sequences, which affects the working efficiency and robustness of the codec.
[0005] Firstly, a speech synthesis method based on decoupled VQ-VAE is provided, including:
[0006] The first frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment. The reference frame segment is input into the global reference encoder. The time-invariant features are decoupled from the reference frame segment by the global reference encoder to obtain a global representation of speech information separated from the time-varying features.
[0007] A second frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized. The second frame segment and the global representation of the speech information are input together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE.
[0008] The discrete sequence is decoded by a decoder to obtain speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of speech information to generate a reconstructed Mel spectrogram;
[0009] The reconstructed Mel spectrogram is converted into a speech waveform using a vocoder to obtain the speech synthesis result.
[0010] Secondly, a speech synthesis device based on decoupled VQ-VAE is provided, comprising:
[0011] Reference coding module: used to randomly select the first frame segment from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment, input the reference frame segment into the global reference encoder, and decouple the time-invariant features from the reference frame segment through the global reference encoder to obtain a global representation of speech information separated from the time-varying features;
[0012] Audio encoding module: used to randomly select a second frame segment from the Mel spectrogram of the speech signal to be synthesized, and input the second frame segment and the global representation of speech information together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through decoupled VQ-VAE.
[0013] Audio decoding module: used to decode the discrete sequence through a decoder to obtain speech detail information of the speech signal to be synthesized, and add the speech detail information to the global representation of speech information to generate a reconstructed Mel spectrogram;
[0014] Speech synthesis module: used to convert the reconstructed Mel spectrogram into a speech waveform using a vocoder to obtain the speech synthesis result.
[0015] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described speech synthesis method based on decoupled VQ-VAE.
[0016] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described speech synthesis method based on decoupled VQ-VAE.
[0017] The aforementioned speech synthesis method, apparatus, computer equipment, and storage medium based on decoupled VQ-VAE decouple time-invariant features from the speech signal by introducing a global reference encoder before encoding and decoding. This separates the global representation of the speech signal from time-varying content information, allowing the speech codec to embed more speech content information during encoding and decoding. Furthermore, a single-codebook speech codec based on decoupled VQ-VAE decouples the speech signal into discrete sequences rich in time-invariant features and speech information. Quantization of these discrete sequences using only a single codebook avoids the problem of multi-sequence prediction, improving the efficiency and robustness of the speech codec. It achieves better speech reconstruction quality than multi-codebook codecs with lower bandwidth. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of an application environment for a speech synthesis method based on decoupled VQ-VAE in one embodiment of the present invention;
[0020] Figure 2 This is a flowchart illustrating a speech synthesis method based on decoupled VQ-VAE in the first embodiment of the present invention;
[0021] Figure 3 This is a flowchart illustrating a speech synthesis method based on decoupled VQ-VAE in the second embodiment of the present invention;
[0022] Figure 4 This is a schematic diagram of a speech synthesis device based on decoupled VQ-VAE in one embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0024] Figure 6 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] The speech synthesis method based on decoupled VQ-VAE provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can randomly select a first frame segment from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment through the client. The reference frame segment is input into a global reference encoder, which decouples the time-invariant features from the reference frame segment to obtain a global representation of the speech information separated from the time-varying features. The server can then randomly select a second frame segment from the Mel spectrogram of the speech signal to be synthesized, and input the second frame segment and the global representation of the speech information together into a single-codebook speech codec based on decoupled VQ-VAE. The single-codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE. The decoder decodes the discrete sequence to obtain the speech detail information of the speech signal to be synthesized, and adds the speech detail information to the global representation of the speech information to generate a reconstructed Mel spectrogram. The vocoder converts the reconstructed Mel spectrogram into a speech waveform to obtain the speech synthesis result, which is then fed back to the client. This invention is applicable to speech synthesis for customer service in the financial sector. Specifically, for speech synthesis in financial customer service, a global reference encoder can be used to decouple time-invariant features from the speech signal, and a single-codebook speech codec based on decoupled VQ-VAE is employed to decouple the speech signal into a discrete sequence rich in time-invariant features and speech information, greatly improving the efficiency and robustness of the speech codec. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0027] Please see Figure 2 As shown, Figure 2 The flowchart of the speech synthesis method based on decoupled VQ-VAE provided in the first embodiment of the present invention is shown, including the following steps:
[0028] S100: Randomly select the first frame segment from the Mel spectrogram of the speech signal to be synthesized as the reference frame segment, input the reference frame segment into the global reference encoder, and decouple the time-invariant features from the reference frame segment through the global reference encoder to obtain the global representation of speech information separated from the time-varying features;
[0029] In this step, time-invariant and time-variant features are used to describe the changes in the speech codec system's response to the input speech signal. Time-invariant features indicate that the characteristics and behavior of the speech codec system do not change over time. That is, given an input speech signal, regardless of when the input begins, its output will depend only on the speech signal itself, and is independent of the start time of the input speech signal. Time-variant features, on the other hand, indicate that the characteristics and behavior of the speech codec system will vary depending on the start time of the input speech signal. This embodiment of the invention introduces a global reference encoder before speech encoding and decoding. The global reference encoder consists of convolutional layers and GRU (Gate Recurrent Unit) layers. After the reference frame fragment is input into the global reference encoder, it first undergoes convolution processing through the convolutional layers to obtain convolutional features. Then, the GRU uses a gating mechanism to decouple the time-invariant features from the convolutional features, and after linear layer transformation, a global representation of speech information including timbre and acoustic environment is obtained. This allows the single-codebook speech codec to embed more speech content information during speech encoding and decoding.
[0030] S110: Randomly select the second frame segment from the Mel spectrogram of the speech signal to be synthesized, and input the second frame segment and the global representation of the speech information together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE.
[0031] In this step, VQ-VAE (Vector Quantized Variational Autoencoder) is a deep learning model that combines variational autoencoder (VAE) and quantization techniques. This embodiment of the invention uses decoupled VQ-VAE to decouple the speech signal into time-invariant embeddings and discrete sequences rich in speech information. Only a single codebook is used to quantize the discrete sequences, avoiding the multi-sequence prediction problem of speech codecs and improving the efficiency and robustness of the speech codec. Specifically, the single-codebook speech codec based on decoupled VQ-VAE includes an encoder, a decoder, and a discrete codebook. After inputting the second frame segment and the global representation of speech information into the single-codebook speech codec based on decoupled VQ-VAE, the encoder first extracts the speech-related phoneme information from the second frame segment and subtracts it from the global representation of speech information to obtain the potential speech content representation. Then, the vector quantizer maps the potential speech content representation to the discrete codebook space to obtain the quantized speech content representation.
[0032] Furthermore, the encoder includes a hybrid sampling module, a Conformer module, a resampling module, and a BLSTM (Bidirectional Long Short-Term Memory) module. The encoder's encoding algorithm specifically includes:
[0033] S111: The second frame segment is downsampled by the hybrid sampling module, and the downsampled features are fed into multiple convolutional blocks for convolutional feature extraction;
[0034] S112: The convolutional features are fed into the BLSTM module for context modeling to obtain context information between adjacent frames, and the phoneme information related to speech in the convolutional features is extracted through the resampling module.
[0035] It is understood that the embodiments of the present invention improve the clustering efficiency of speech content by introducing a BLSTM module into the encoder to obtain contextual information and using this contextual information to help discover the correlation between adjacent frames. Introducing a resampling module encourages the encoder to extract more phoneme information with lower short-term variance and related to speech from the acoustic sequence.
[0036] S113: After the phoneme information is transformed by a linear layer, it is subtracted from the global representation of the speech information, and then upsampled through the hybrid sampling module to obtain the upsampled features;
[0037] It should be noted that in S111 and S113, the hybrid sampling module combines convolution and pooling operations to downsample the second frame segment, and combines transposed convolution and copying to upsample the subtracted phoneme information, thereby reducing the distortion of upsampling and downsampling.
[0038] S114: The upsampled features are fed into the BLSTM module for context modeling, and the output of the BLSTM module is transformed by a linear layer to obtain the potential speech content representation;
[0039] S115: The potential speech content representation is fed into the vector quantizer, which maps the potential speech content representation to the discrete codebook space to obtain the quantized speech content representation.
[0040] S120: The discrete sequence is decoded by the decoder to obtain the speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of the speech information to generate the reconstructed Mel spectrogram;
[0041] In this step, after the encoder obtains the quantized speech content representation, the speech content representation is input into the decoder. The decoder decodes the speech content representation to obtain the reconstructed Mel spectrogram. Specifically, the decoding algorithm for the speech content representation includes:
[0042] S121: The speech content representation is downsampled using the hybrid sampling module to obtain downsampled features;
[0043] S122: The downsampled features are fed into multiple transposed convolutional blocks for feature decoding. The decoded features are fed into a residual block combined with a BLSTM module. Context modeling is performed through the residual block, and the speech detail information of the speech signal is further recovered through the resampling module.
[0044] S123: The output of the residual block is transformed by a linear layer and added to the global representation of the speech information. The added features are then fed into the BLSTM module for context modeling.
[0045] S124: The output of the BLSTM module is upsampled by the hybrid sampling module and then transformed by a linear layer to obtain the reconstructed Mel spectrum.
[0046] S130: The reconstructed Mel spectrogram is converted into a speech waveform using a vocoder to obtain the speech synthesis result;
[0047] In this step, the reconstructed Mel spectrogram is fed into a pre-trained BigVGAN vocoder, which then converts the reconstructed Mel spectrogram into a speech waveform.
[0048] As can be seen, in the above scheme, the speech synthesis method based on decoupled VQ-VAE provided in the first embodiment of the present invention decouples the time-invariant features from the speech signal by introducing a global reference encoder before encoding and decoding, separating the global representation of the speech signal from the time-varying content information, so that the speech codec can embed more speech content information during encoding and decoding. Furthermore, it employs a single-codebook speech codec based on decoupled VQ-VAE to decouple the speech signal into discrete sequences with time-invariant features and rich speech information. Only a single codebook is used to quantize the discrete sequences, thereby avoiding the problem of multi-sequence prediction, improving the working efficiency and robustness of the speech codec, and achieving better speech reconstruction quality than multi-codebook codecs at lower bandwidth.
[0049] Please see Figure 3 As shown, Figure 3 The flowchart of the speech synthesis method based on decoupled VQ-VAE provided in the second embodiment of the present invention includes the following steps:
[0050] S200: The Fourier transform method is used to convert the speech signal to be synthesized into a Mel spectrogram;
[0051] In this step, the Mel spectrogram is a method for converting the spectral representation of a speech signal to a Mel frequency scale. This embodiment of the invention uses the Fourier transform method to transform the speech signal to be synthesized from the time domain to the frequency domain, which can better simulate the human ear's perception of frequency and facilitate the analysis of the intensity and distribution of each frequency component in the speech signal. It can be understood that the Mel spectrogram obtained by transforming the speech signal to be synthesized using the Fourier transform method is the original Mel spectrogram.
[0052] S210: Randomly select the first frame segment from the Mel spectrogram as the reference frame segment, input the reference frame segment into the global reference encoder, and decouple the time-invariant features from the reference frame segment through the global reference encoder to obtain a global representation of speech information separated from the time-varying features;
[0053] It is understood that the processing procedure of the global reference encoder in this step is the same as S100 in the first embodiment. To avoid redundancy, it will not be described again here.
[0054] S220: Randomly select the second frame segment from the Mel spectrogram, and input the second frame segment and the global representation of the speech information together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE.
[0055] It is understood that the encoding process of the single codebook speech codec in this step is the same as S110 in the first embodiment. To avoid redundancy, it will not be described again here.
[0056] S230: The discrete sequence is decoded by the decoder to obtain the speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of the speech information to generate the reconstructed Mel spectrogram;
[0057] It is understood that the encoding process of the single codebook speech codec in this step is the same as S120 in the first embodiment. To avoid redundancy, it will not be described again here.
[0058] S240: The reconstructed Mel spectrogram is fed into the pre-trained BigVGAN vocoder, which converts the reconstructed Mel spectrogram into a speech waveform.
[0059] S250: The reconstructed Mel spectrogram is fed into the discriminator for real vs. fake discrimination, and the discriminator is trained adversarially based on the discrimination results to make the reconstructed Mel spectrogram closer to the original Mel spectrogram.
[0060] In this step, the total loss function for adversarial training of the discriminator is:
[0061] L = L_recon + L_commit + L_adv (1)
[0062] Here, L_recon represents the reconstruction loss, which measures the difference between the reconstructed Mel spectrogram and the original Mel spectrogram. L_commit represents the commitment loss, which measures the difference between the speech content representation before and after quantization, encouraging the encoder output to be closer to the cluster centers in the codebook. L_adv represents the adversarial loss, which improves the reconstruction quality of the Mel spectrogram, making the reconstructed Mel spectrogram closer to the true distribution.
[0063] As can be seen, in the above scheme, the speech synthesis method based on decoupled VQ-VAE provided in the second embodiment of the present invention uses a discriminator to distinguish between real and fake Mel spectrograms and performs adversarial training on the discriminator, so that the reconstructed Mel spectrogram is closer to the original Mel spectrogram, thereby improving the speech reconstruction quality of the speech codec system.
[0064] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0065] In one embodiment, a speech synthesis device based on decoupled VQ-VAE is provided, which corresponds one-to-one with the speech synthesis method based on decoupled VQ-VAE in the above embodiments. For example... Figure 4 As shown, the speech synthesis device based on decoupled VQ-VAE includes an audio conversion module 101, a reference encoding module 102, an audio encoding module 103, an audio decoding module 104, a speech synthesis module 105, and an adversarial training module 106. Detailed descriptions of each functional module are as follows:
[0066] Audio conversion module 101: This module uses Fourier transform to convert the speech signal to be synthesized into a Mel spectrogram. The Mel spectrogram is a method for converting the spectral representation of a speech signal to a Mel frequency scale. In this embodiment, the Fourier transform converts the speech signal to be synthesized from the time domain to the frequency domain, which better simulates the human ear's perception of frequency and facilitates the analysis of the intensity and distribution of each frequency component in the speech signal. It can be understood that the Mel spectrogram obtained by converting the speech signal to be synthesized using the Fourier transform is the original Mel spectrogram.
[0067] Reference coding module 102: used to randomly select the first frame segment from the Mel spectrogram as the reference frame segment, input the reference frame segment into the global reference encoder, and decouple the time-invariant features from the reference frame segment through the global reference encoder to obtain a global representation of speech information separated from the time-varying features;
[0068] Audio encoding module 103: used to randomly select the second frame segment from the Mel spectrogram, input the second frame segment and the global representation of speech information together into a single codebook speech codec based on decoupled VQ-VAE, the single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE;
[0069] Audio decoding module 104: used to decode discrete sequences through a decoder, obtain speech detail information of the speech signal to be synthesized, and add the speech detail information to the global representation of speech information to generate a reconstructed Mel spectrogram;
[0070] Speech synthesis module 105: used to convert the reconstructed Mel spectrogram into a speech waveform through a vocoder to obtain the speech synthesis result; wherein, the reconstructed Mel spectrogram is fed into a pre-trained BigVGAN vocoder, and the reconstructed Mel spectrogram is converted into a speech waveform through the BigVGAN vocoder.
[0071] Adversarial training module 106: This module is used to feed the reconstructed Mel spectrogram into the discriminator for true / false discrimination and to perform adversarial training on the discriminator based on the discrimination results, so that the reconstructed Mel spectrogram is closer to the original Mel spectrogram.
[0072] In one embodiment, the reference encoding module 102 is specifically used for:
[0073] The second frame segment is downsampled using a hybrid sampling module, and the downsampled features are then fed into multiple convolutional blocks for convolutional feature extraction.
[0074] The convolutional features are fed into the BLSTM module for context modeling to obtain context information between adjacent frames. The context information is then input into the resampling module, which extracts phoneme information related to speech from the speech signal.
[0075] After the phoneme information is transformed by a linear layer and subtracted from the global representation of the speech information, it is upsampled through a hybrid sampling module to obtain upsampled features.
[0076] The upsampled features are fed into the BLSTM module for context modeling, and the output of the BLSTM module is transformed by a linear layer to obtain the latent content representation.
[0077] The latent content representation is fed into a vector quantizer, which maps the latent content representation to a discrete codebook space to obtain the quantized content representation.
[0078] In one embodiment, the audio encoding module 103 is specifically used for:
[0079] The quantized content representation is downsampled using a hybrid sampling module to obtain downsampled features;
[0080] The downsampled features are fed into multiple transposed convolutional blocks for feature decoding. The decoded features are then fed into a residual block that incorporates a BLSTM module. Context modeling is performed through the residual block, and the speech details of the speech signal are further recovered through a resampling module.
[0081] The output of the residual block is transformed by a linear layer and added to the global representation of the speech information. The added features are then fed into the BLSTM module for context modeling.
[0082] The output of the BLSTM module is upsampled by the hybrid sampling module and then transformed by a linear layer to obtain the reconstructed Mel spectrum.
[0083] The speech synthesis device based on decoupled VQ-VAE provided by this invention decouples time-invariant features from the speech signal by introducing a global reference encoder before encoding and decoding, separating the global representation of the speech signal from time-varying content information. This allows the speech codec to embed more speech content information during encoding and decoding. Furthermore, a single-codebook speech codec based on decoupled VQ-VAE decouples the speech signal into discrete sequences with time-invariant features and rich speech information. Quantization of these discrete sequences is performed using only a single codebook, thus avoiding the problem of multi-sequence prediction, improving the efficiency and robustness of the speech codec, and achieving better speech reconstruction quality than multi-codebook codecs with lower bandwidth.
[0084] Specific limitations regarding the speech synthesis device based on decoupled VQ-VAE can be found in the limitations of the speech synthesis method based on decoupled VQ-VAE mentioned above, and will not be repeated here. Each module in the aforementioned speech synthesis device based on decoupled VQ-VAE can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0085] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side speech synthesis method based on decoupled VQ-VAE.
[0086] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a decoupled VQ-VAE-based speech synthesis method.
[0087] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0088] The first frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized as the reference frame segment. The reference frame segment is input into the global reference encoder. The time-invariant features are decoupled from the reference frame segment by the global reference encoder to obtain a global representation of the speech information separated from the time-varying features.
[0089] The second frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized. The second frame segment and the global representation of the speech information are input together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE.
[0090] The discrete sequence is decoded by a decoder to obtain the speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of the speech information to generate the reconstructed Mel spectrogram;
[0091] The reconstructed Mel spectrogram is converted into a speech waveform using a vocoder to obtain the speech synthesis result.
[0092] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0093] The first frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized as the reference frame segment. The reference frame segment is input into the global reference encoder. The time-invariant features are decoupled from the reference frame segment by the global reference encoder to obtain a global representation of the speech information separated from the time-varying features.
[0094] The second frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized. The second frame segment and the global representation of the speech information are input together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and rich speech information through the decoupled VQ-VAE.
[0095] The discrete sequence is decoded by a decoder to obtain the speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of the speech information to generate the reconstructed Mel spectrogram;
[0096] The reconstructed Mel spectrogram is converted into a speech waveform using a vocoder to obtain the speech synthesis result.
[0097] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0098] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other storage media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0099] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0100] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech synthesis method based on decoupled VQ-VAE, characterized in that, include: The first frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment. The reference frame segment is input into the global reference encoder, which includes a convolutional layer and a GRU layer. After the reference frame segment is input into the global reference encoder, the convolutional layer performs convolution processing on the reference frame segment to obtain convolutional features. The GRU layer uses a gating mechanism to decouple the time-invariant features from the convolutional features, and after linear layer transformation, a global representation of the speech information is obtained. A second frame segment is randomly selected from the Mel spectrogram of the speech signal to be synthesized. The second frame segment and the global representation of the speech information are input together into a single codebook speech codec based on decoupled VQ-VAE. The second frame segment is downsampled by the hybrid sampling module, and the downsampled features are sent into multiple convolutional blocks for convolutional feature extraction. The convolutional features are fed into the BLSTM module for context modeling to obtain context information between adjacent frames. The phoneme information related to speech in the convolutional features is extracted through the resampling module. The phoneme information is transformed by a linear layer and then subtracted from the global representation of the speech information. The result is then upsampled using a hybrid sampling module to obtain upsampled features. The upsampled features are fed into the BLSTM module for context modeling, and the output of the BLSTM module is transformed by a linear layer to obtain the potential speech content representation. The speech content representation is fed into a vector quantizer, which maps the speech content representation to a discrete codebook space to obtain a quantized speech content representation, which constitutes a discrete sequence with time-invariant features and speech information. The discrete sequence is decoded by a decoder to obtain speech detail information of the speech signal to be synthesized, and the speech detail information is added to the global representation of speech information to generate a reconstructed Mel spectrogram; The reconstructed Mel spectrogram is converted into a speech waveform using a vocoder to obtain the speech synthesis result.
2. The speech synthesis method based on decoupled VQ-VAE as described in claim 1, characterized in that, The step of decoding the discrete sequence using a decoder to obtain speech detail information of the speech signal to be synthesized, and adding the speech detail information to the global representation of the speech information to generate a reconstructed Mel spectrogram, includes: The speech content representation is downsampled using the hybrid sampling module to obtain downsampled features. The downsampled features are fed into a transposed convolutional block for feature decoding, the decoded features are fed into the residual block of the BLSTM module for context modeling, and the speech detail information is recovered through a resampling module. The output of the residual block is transformed by a linear layer and added to the global representation of the speech information. The added features are then fed into the BLSTM module for context modeling. The output of the BLSTM module is upsampled by the hybrid sampling module and then reconstructed by linear layer transformation to obtain the Mel spectrum.
3. The speech synthesis method based on decoupled VQ-VAE as described in claim 1, characterized in that, Before randomly selecting a first frame segment from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment, the method further includes: The speech signal to be synthesized is converted into a Mel spectrogram using the Fourier transform method.
4. The speech synthesis method based on decoupled VQ-VAE as described in any one of claims 1-3, characterized in that, After decoding the discrete sequence using a decoder to obtain speech detail information of the speech signal to be synthesized, and adding the speech detail information to the global representation of the speech information to generate the reconstructed Mel spectrogram, the method further includes: The reconstructed Mel spectrogram is fed into a discriminator for real / false detection. Based on the detection results, the discriminator is subjected to adversarial training to make the reconstructed Mel spectrogram closer to the true distribution.
5. The speech synthesis method based on decoupled VQ-VAE as described in claim 4, characterized in that, The total loss function for adversarial training of the discriminator is: L = L_recon + L_commit + L_adv Where L_recon represents the reconstruction loss, which measures the difference between the reconstructed Mel spectrogram and the original Mel spectrogram; L_commit represents the commitment loss, which measures the difference between the speech content representation before and after quantization; and L_adv represents the adversarial loss, which is used to improve the reconstruction quality of the Mel spectrogram.
6. A speech synthesis apparatus based on decoupled VQ-VAE, used to implement the speech synthesis method based on decoupled VQ-VAE as described in any one of claims 1-5, characterized in that, include: Reference coding module: used to randomly select the first frame segment from the Mel spectrogram of the speech signal to be synthesized as a reference frame segment, input the reference frame segment into the global reference encoder, and decouple the time-invariant features from the reference frame segment through the global reference encoder to obtain a global representation of speech information separated from the time-varying features; Audio encoding module: used to randomly select a second frame segment from the Mel spectrogram of the speech signal to be synthesized, and input the second frame segment and the global representation of speech information together into a single codebook speech codec based on decoupled VQ-VAE. The single codebook speech codec decouples the second frame segment into a discrete sequence with time-invariant features and speech information through decoupled VQ-VAE. Audio decoding module: used to decode the discrete sequence through a decoder to obtain speech detail information of the speech signal to be synthesized, and add the speech detail information to the global representation of speech information to generate a reconstructed Mel spectrogram; Speech synthesis module: used to convert the reconstructed Mel spectrogram into a speech waveform using a vocoder to obtain the speech synthesis result.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the speech synthesis method based on decoupled VQ-VAE as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the speech synthesis method based on decoupled VQ-VAE as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Synthesis by generation and concatenation of multi-form segments
CN101828218A
Parallel speech synthesis method and device based on variational auto-encoder
CN113450761A