Synthetic speech generation method and device, computer equipment and storage medium
Through the combination of potential diffusion model and pre-trained language model, the problem of speech synthesis method dependence on labeled data is solved, effective speech generation and efficient long audio sequence processing under limited data are realized, and speech alignment accuracy and naturalness of generating speech are improved.
Patent Information
- Application Number
- CN202510726460.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-26
AI Technical Summary
Existing speech synthesis methods need to rely on a large amount of labeled data, especially in areas with scarce resources, it is difficult to effectively learn generative models.
The latent diffusion model and pre-trained automatic encoder are used to map high-dimensional speech data to a compact latent space, combined with the representation of the pre-trained language model as conditional information, the audio Transformer model is used for audio prediction, and the model performance is improved through the position-aware cross-attention mechanism.
Reduces dependence on labeled data, can effectively generalize to different text inputs under limited data, efficiently process long audio sequences, and improves the alignment accuracy between speech and text and the naturalness of generating speech.
Smart Images

Figure CN120544538A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and is applicable to the financial or medical fields, and in particular to a method, apparatus, computer equipment, and storage medium for generating synthetic speech. Background Art
[0002] Neural methods have revolutionized generative speech modeling, with autoregressive and diffusion-based models in particular driving recent progress. However, these improvements come at a cost. Generative models are extremely data-intensive, with state-of-the-art systems requiring increasingly large amounts of annotated data. This poses a challenge for applying such methods to resource-limited domains and languages. Learning effective generative models with limited data remains an unresolved challenge.
[0003] Currently, speech synthesis models are typically diffuse TTS models such as NaturalSpeech2 (NS2) and VoiceBox. They rely on phonemers and aligners to generate frame-level phoneme transcriptions, which can introduce errors. Furthermore, both models require phoneme duration annotations for generation, which means an external model is needed for phoneme duration prediction.
[0004] It can be seen from this that traditional speech synthesis methods have the problem of relying on a large amount of labeled data. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to propose a method, apparatus, computer device and storage medium for generating synthetic speech to solve the problem that traditional speech synthesis methods need to rely on a large amount of labeled data.
[0006] In order to solve the above technical problems, the present invention provides a method for generating synthetic speech, which adopts the following technical solutions:
[0007] Obtaining a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data;
[0008] Inputting the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedded data;
[0009] Performing a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data;
[0010] Inputting the text mark position data into a position encoder for position encoding operation to obtain mark position embedded data;
[0011] The XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data are input into an audio Transformer model for performing an audio prediction operation to obtain audio prediction data.
[0012] Furthermore, the step of performing a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data specifically includes the following steps:
[0013] Inputting the speech data to be synthesized into a SoundStream encoder for speech encoding operation to obtain latent speech data;
[0014] The potential speech data is input into a U-Net encoder for high-resolution encoding to obtain the downsampled registration mark data.
[0015] Furthermore, after the step of inputting the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data into the audio Transformer model for performing an audio prediction operation to obtain audio prediction data, the following step is also included:
[0016] The audio prediction data is input into the SoundStream decoder to perform an audio sampling point restoration operation to obtain target audio prediction data.
[0017] Furthermore, the step of inputting the audio prediction data into a SoundStream decoder to perform an audio sampling point restoration operation to obtain target audio prediction data specifically includes the following steps:
[0018] Inputting the audio prediction data into a U-Net decoder for audio decoding operation to obtain denoised audio prediction data;
[0019] The denoised audio prediction data is input into a SoundStream decoder to perform an audio sampling point restoration operation to obtain the target audio prediction data.
[0020] Furthermore, the audio Transformer model is trained according to an asymmetric diffusion loss weighting strategy, wherein the asymmetric diffusion loss weighting strategy is expressed as:
[0021]
[0022] Among them, λ t represents the noise level, γ and x0 are the parameters of the Cauchy distribution, and μ and σ are the parameters of the normal distribution.
[0023] Furthermore, the audio Transformer model uses the XLM-RoBERTa model as a generator and trains it to generate speech duration based on text transcription. Specifically:
[0024] The input of the XLM-RoBERTa model is represented as follows: given a text transcription T, the model generates the corresponding speech duration D;
[0025] The generation process of the XLM-RoBERTa model is expressed as: using kernel sampling (p=0.95) to perform autoregressive generation, wherein the process of the autoregressive generation is expressed as:
[0026] D = XLM-RoBERTa(T;θ)
[0027] Among them, θ represents the parameters of the model;
[0028] The loss function of the XLM-RoBERTa model is expressed as:
[0029]
[0030] Where N represents the number of samples, D i Indicates the actual duration. In order to solve the above technical problems, the embodiment of the present application also provides a synthetic speech generation device, which adopts the following technical solution:
[0031] a request acquisition module, configured to acquire a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data;
[0032] A text encoding module is used to input the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedded data;
[0033] A downsampling module, configured to perform a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data;
[0034] A position encoding module, configured to input the text mark position data into a position encoder for performing a position encoding operation to obtain mark position embedded data;
[0035] The audio prediction module is used to input the XLM tag embedding data, the downsampled registration tag data and the tag position embedding data into the audio Transformer model to perform an audio prediction operation to obtain audio prediction data.
[0036] Furthermore, the downsampling module includes:
[0037] A speech encoding submodule, configured to input the speech data to be synthesized into a SoundStream encoder for speech encoding to obtain latent speech data;
[0038] The high-resolution encoding submodule is used to input the potential speech data into a U-Net encoder for high-resolution encoding operation to obtain the downsampled registration mark data.
[0039] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0040] The invention comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the synthetic speech generation method as described above when executing the computer-readable instructions.
[0041] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0042] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the synthetic speech generation method described above.
[0043] The present application provides a method for generating synthesized speech, comprising: obtaining a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text tag position data; inputting the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedding data; performing a downsampling operation on the speech data to be synthesized to obtain downsampled registration tag data; inputting the text tag position data into a position encoder for position encoding operation to obtain tag position embedding data; inputting the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data into an audio Transformer model for audio prediction operation to obtain audio prediction data. Compared with the prior art, the present application uses a latent diffusion model to map high-dimensional speech data to a compact latent space using a pre-trained autoencoder, thereby reducing dependence on labeled data. At the same time, the representation of the pre-trained language model is used as conditional information, so that it can effectively generalize to different text inputs under limited data. The method adopts a new diffusion architecture that can efficiently process long audio sequences and improve model performance through a position-aware cross-attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0046] Figure 2 This is a flowchart of the implementation of the synthetic speech generation method provided in the embodiment of the present application;
[0047] Figure 3 This is a schematic diagram of the structure of an embodiment of the audio Transformer model provided in the embodiment of the present application;
[0048] Figure 4 Schematic diagram of the structure of the synthetic speech generation device provided in an embodiment of the present application;
[0049] Figure 5 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0051] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0052] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0053] like Figure 1As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0054] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0055] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0056] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0057] It should be noted that the synthetic speech generation method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the synthetic speech generation device is generally set in the server / terminal device.
[0058] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0059] Continue to refer Figure 2 , shows a flow chart of an embodiment of a synthetic speech generation method according to the present application. The synthetic speech generation method comprises: step S201, step S202, step S203, step S204 and step S205.
[0060] In step S201 , a synthesized speech generation request sent by a user terminal is obtained, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data.
[0061] In the embodiments of the present application, the user terminal refers to a terminal device used to execute the image processing method for preventing document abuse provided by the present application. The user terminal can be a mobile terminal such as a mobile phone, a smart phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), a navigation device, etc., as well as a fixed terminal such as a digital TV, a desktop computer, etc. It should be understood that the examples of user terminals here are only for convenience of understanding and are not used to limit the present application.
[0062] In an embodiment of the present application, users input text data to be synthesized, voice data to be synthesized, and text mark position data through their terminal devices (such as mobile phones, computers, etc.), and this input is received by the system. The text data to be synthesized is the specific content that the user wants the system to synthesize that is highly similar to the synthesized voice data. Specifically, the text data to be synthesized can be "transaction data or payment data or business data or purchase data" related to financial institutions (such as banks, etc.). The text data to be synthesized can also be medical data related to medical scenarios, as an example, such as personal health records, prescriptions, examination reports and other data. It should be understood that the examples of text data to be synthesized here are only for convenience of understanding and are not used to limit this application.
[0063] In step S202, the text data to be synthesized is input into an XLM encoder for text encoding operation to obtain XLM tag embedded data.
[0064] In the embodiment of the present application, the text encoding operation is mainly used to semantically encode the text using XLM (cross-language pre-trained model) to generate XLM tag embedded data, wherein XLM can capture the semantic and grammatical structure of the text through multilingual pre-training and support multilingual scenarios.
[0065] In an embodiment of the present application, the output of the text encoding operation is a vector representation of the text (eg, each word is converted into a 512-dimensional vector), retaining semantic information for subsequent fusion.
[0066] In step S203, a down-sampling operation is performed on the speech data to be synthesized to obtain down-sampled registration mark data.
[0067] In step S204, the text mark position data is input into a position encoder for position encoding operation to obtain mark position embedded data.
[0068] In the embodiment of the present application, the position encoding operation mainly converts the text tag position data (such as ["today"@0.5s, "weather"@0.8s]) into tag position embedding data, and then uses a sine function or a learnable position encoding to mark the text rhythm information.
[0069] In the embodiment of the present application, the output of the position encoding operation is a position-sensitive vector (eg, a vector of “today” with a time offset added).
[0070] In step S205, the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data are input into the audio Transformer model for audio prediction operation to obtain audio prediction data.
[0071] In the embodiment of the present application, the audio prediction operation mainly integrates XLM tag embedding (semantics), downsampled registration tags (acoustics), and tag position embedding (rhythm) as input. Transformer aligns text and speech features across modalities through the self-attention mechanism to predict the intermediate representation of the target audio.
[0072] In the embodiment of the present application, the output of the audio prediction operation is an audio prediction sequence (such as a predicted Mel spectrum) that mixes semantic and acoustic features.
[0073] In the embodiment of the present application, the audio Transformer model proposed in the present application is a hybrid architecture that combines the advantages of U-Net and Transformer, such as Figure 3 As shown in the figure, U-Net excels at processing high-resolution data, while Transformer excels at capturing long-range dependencies and integrating conditional information. This application first uses a 1D U-Net to downsample audio features from 1504 frames to 188 frames. This step enables the application of a deep Transformer on the compressed sequence and integrates information from the text transcription. Directly using a Transformer to process data at its original resolution would be computationally very expensive.
[0074] In an embodiment of the present application, in order to enhance the Transformer's ability to model global information, the present application draws on the latest technology of the visual Transformer and adds 8 learnable register tags before the downsampled features. These tags act as global memory slots, enabling the Transformer to process global information more efficiently. After applying the Transformer, the register tags are removed, and the U-Net decoder only upsamples the corresponding audio features and restores them to the original sequence length for final prediction. The hybrid U-Net / Transformer architecture has shown good results in processing high-resolution image diffusion, which prompted the present application to apply this architecture to the audio field.
[0075] In an embodiment of the present application, in text-to-speech synthesis, accurately aligning the generated speech with the input text is a key challenge. In order to improve the alignment problem, the present application introduces a position-aware cross-attention layer in the Transformer model, which focuses on the text representation from the frozen XLM-RoBERTa encoder. In order to explicitly integrate the relative position information of the tokens in the text, the present application introduces a neural position encoder that maps the relative positions of the text tokens to key vectors. The present application adds these position key vectors to the corresponding key vectors in the XLM-RoBERTa embedding and applies them to the cross-attention mechanism. This enables the model to directly search and focus on relevant positions in the text when generating each audio frame. Specifically:
[0076] The logic for calculating cross attention in this application is as follows:
[0077] Let Q be the query vector, K be the key vector, and V be the value vector. K and V come from the output of the XLM-RoBERTa encoder, while Q comes from the output of the decoder layer. This application first calculates the position key vector K pos , which represents the relative position information of the text mark. Then, this application will K pos The key vector K in the XLM-RoBERTa embedding XLM-RoBERTa Add together to get the final key vector K final :
[0078] K final =K XLM-RoBERTa +K pos
[0079] Next, the application calculates the cross-attention score:
[0080]
[0081] Among them, d kis the dimension of the key vector. In this way, the model can more accurately locate the corresponding position in the text when generating each audio frame, thereby improving the alignment accuracy of speech and text.
[0082] To further enhance the position-aware cross-attention, this application also introduces a position mask mechanism. Specifically, when calculating the attention score, this application uses a mask matrix M to suppress the attention weights of irrelevant positions:
[0083]
[0084] The design of the mask matrix M ensures that the model focuses only on the current position and its surroundings in the text when generating each audio frame, preventing the model from straying from the current text position during the generation process. This position-aware cross-attention mechanism not only improves the alignment accuracy of speech synthesis but also enhances the model's ability to understand and generate text content.
[0085] In this way, the model can more accurately capture the positional information in the text when generating speech, ensuring that the generated speech is consistent with the input text on the timeline. This approach is particularly important in text-to-speech synthesis tasks because it directly affects the naturalness and comprehensibility of the generated speech.
[0086] In an embodiment of the present application, unlike the autoregressive method that relies on discrete tags, the diffusion model is good at generating continuous representations, which not only avoids the information loss caused by quantization, but also processes long sequences more efficiently. In order to take full advantage of these advantages, the present application trains a diffusion model on the continuous potential embedding of a pre-trained audio autoencoder. Specifically, the present application adopts the public Soundstream autoencoder to convert a 24kHz audio waveform into 75 potential vectors per second. Soundstream uses residual vector quantization to map each continuous potential vector to multiple discrete tags to capture finer details. However, the present application does not model these discrete tags, but instead trains a diffusion model to generate a 75Hz, 128-dimensional continuous embedding output by the Soundstream encoder before quantization. This continuous potential diffusion method significantly reduces the sequence length compared to tag modeling - 10 seconds of audio contains only 750 potential vectors, while the number of tags after quantization is 24,000 (a 32-fold reduction). The continuous potential vectors generated during the inference process can be further quantized and decoded by Soundstream to generate the final audio waveform.
[0087] In the embodiment of the present application, in the inference stage, the generation process is as follows:
[0088] 1. Initialization: Sample the initial latent vector z0 from a standard normal distribution.
[0089] 2. Iterative generation: For each time step t, the latent vector is updated according to the following formula:
[0090]
[0091] Among them, α t and β t are predefined parameters, is α t The cumulative product of noise is noise sampled from a standard normal distribution.
[0092] 3. Bootstrap parameters: At each time step, the bootstrap parameters w are used to adjust the outputs of the conditional and unconditional models:
[0093]
[0094] In this way, the present application can flexibly control the application intensity of conditional information during the generation process, thereby achieving more flexible and diverse generation results.
[0095] -Classifier-free bootstrapping: A technique that jointly trains conditional and unconditional models and dynamically adjusts the strength of the conditional information at inference time.
[0096] - Bootstrap weight W: A parameter that controls the degree to which the model relies on conditional information. w>1 indicates greater reliance on conditional information, while w<1 indicates greater reliance on unconditional information.
[0097] -Loss function: Contains the losses of the conditional and unconditional models, ensuring that the model learns both generative methods during training.
[0098] - Inference process: The final audio is generated by iteratively updating the latent vector and applying the guided parameters.
[0099] In the embodiments of the present application, this classifier-free guidance method not only improves the generation quality of the model, but also enhances the adaptability and diversity of the model under complex conditions.
[0100] In an embodiment of the present application, a method for generating a synthesized speech is provided, comprising: obtaining a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text tag position data; inputting the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedding data; performing a downsampling operation on the speech data to be synthesized to obtain downsampled registration tag data; inputting the text tag position data into a position encoder for position encoding operation to obtain tag position embedding data; inputting the XLM tag embedding data, downsampled registration tag data, and tag position embedding data into an audio Transformer model for audio prediction operation to obtain audio prediction data. Compared with the prior art, the present application uses a potential diffusion model to map high-dimensional speech data to a compact latent space using a pre-trained autoencoder, thereby reducing dependence on labeled data. At the same time, the representation of the pre-trained language model is used as conditional information, so that it can be effectively generalized to different text inputs under limited data. The method adopts a new diffusion architecture that can efficiently process long audio sequences and improve model performance through a position-aware cross-attention mechanism.
[0101] In some optional implementations of the embodiments of the present application, the step of performing a downsampling operation on the speech data to be synthesized to obtain the downsampled registration mark data specifically includes the following steps:
[0102] Input the speech data to be synthesized into the SoundStream encoder for speech encoding operation to obtain potential speech data;
[0103] The latent speech data is input into the U-Net encoder for high-resolution encoding to obtain downsampled registration marker data.
[0104] In the embodiment of the present application, the SoundStream encoder is a neural audio codec that compresses the original speech waveform into latent speech data (low-dimensional dense vectors) and extracts acoustic features such as Mel spectrum and fundamental frequency through convolution and residual connection.
[0105] In an embodiment of the present application, the output of the SoundStream encoder is an acoustic feature encoding of the reference speech (eg, 16kHz speech compressed into a 200-dimensional latent vector).
[0106] In the embodiment of the present application, the U-Net encoder mainly downsamples the potential speech data to generate downsampled registration mark data. It also retains local details (such as breathing sounds and labiodental sounds) through skip connections to enhance the realism of the speech.
[0107] In an embodiment of the present application, the output of the U-Net encoder is multi-scale acoustic features (such as feature maps of different time resolutions).
[0108] In some optional implementations of the embodiments of the present application, after the step of inputting the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data into the audio Transformer model for audio prediction operation to obtain audio prediction data, the following steps are also included:
[0109] The audio prediction data is input into the SoundStream decoder for audio sampling point restoration operation to obtain the target audio prediction data.
[0110] In the embodiment of the present application, the audio sampling point restoration operation can be to convert the denoised spectrum into a time domain waveform, restore the sampling points through a vector quantized variational autoencoder (VQ-VAE), and dynamically adjust the generation speed (e.g., 16kHz→48kHz super-resolution).
[0111] In the embodiment of the present application, the audio sampling point restoration operation outputs the final target audio (such as a voice file in 16-bit PCM format).
[0112] In some optional implementations of the embodiments of the present application, the step of inputting the audio prediction data into the SoundStream decoder to perform an audio sampling point restoration operation to obtain the target audio prediction data specifically includes the following steps:
[0113] The audio prediction data is input into the U-Net decoder for audio decoding operation to obtain the denoised audio prediction data;
[0114] The denoised audio prediction data is input into the SoundStream decoder for audio sampling point restoration to obtain the target audio prediction data.
[0115] In this embodiment of the present application, the audio prediction data is upsampled and detailed reconstruction is performed to generate denoised audio prediction data. A skip connection is used to fuse encoder features and restore high-frequency details (such as the spectrum of the fricative sound "s").
[0116] In the embodiment of the present application, the output of the audio decoding operation is an intermediate spectrum close to the real speech (such as an 80-dimensional Mel spectrum).
[0117] In some optional implementations of the embodiments of the present application, the audio Transformer model is trained according to an asymmetric diffusion loss weighting strategy, where the asymmetric diffusion loss weighting strategy is expressed as:
[0118]
[0119] Among them, λ trepresents the noise level, γ and x0 are the parameters of the Cauchy distribution, and μ and σ are the parameters of the normal distribution.
[0120] In the embodiments of the present application, accurately highlighting the diffuse noise level that is critical to the perceived quality is the key to improving the synthesis effect. Past studies have mostly used symmetrical weighting methods, such as V-weighting, or unimodal distributions centered around a moderate noise level. However, in text-to-speech synthesis tasks, even at high noise levels, the input text and voice prompts can still provide important signals, while the signals in the latent space are relatively limited. These high noise levels are precisely where conditional information can play the most role, helping to determine the global structure of the speech and align it with the text.
[0121] To capitalize on this, this application adjusts the weighting strategy of the loss function to give greater weight to conditions with high noise levels. This allows the model to more effectively utilize the conditional information in the text and voice prompts during training, ensuring that the overall structure and content of the generated speech are consistent with the input text. This approach not only improves the quality of speech synthesis but also enhances the model's generalization ability with complex text inputs.
[0122] Therefore, this application proposes an asymmetric diffusion loss weighting strategy that assigns greater weight at high noise levels because at these levels, the model relies primarily on text and voice cues to estimate the original speech. Compared with symmetric weighting, the weighting method proposed in this application devotes more model capacity to handling details such as word position and localization. Specifically, this application uses a heavy-tailed Cauchy distribution to parameterize the weights at high noise levels (λt<-1), while using a unimodal normal distribution at low noise levels:
[0123]
[0124] Among them, λ t represents the noise level, γ and x0 are the parameters of the Cauchy distribution, and μ and σ are the parameters of the normal distribution. Through this weighting approach, the present application can more effectively utilize the conditional information in text and voice prompts under high noise levels, thereby improving the quality of speech synthesis and alignment accuracy.
[0125] The core of this asymmetric weighting strategy is:
[0126] 1. Cauchy distribution weighting at high noise level: At high noise level (λ t <-1), the model relies primarily on text and voice cues to infer the global structure of the original speech. In this case, the application uses the Cauchy distribution to parameterize the weights because the Cauchy distribution has a heavier tail and can provide greater weight under extreme noise levels, thus ensuring that the model can fully utilize the conditional information in these critical stages.
[0127] 2. Normal distribution weighting at low noise level: At low noise level (λ t ≥-1), the model is already able to reconstruct the speech signal well. At this time, this application uses normal distribution to parameterize the weights to ensure that the model remains stable in detail reconstruction and fine-tuning.
[0128] This design allows the proposed model to more effectively leverage conditional information from text and voice prompts in high-noise environments, resulting in better alignment and more natural speech generation in speech synthesis. Furthermore, the model maintains fine-grained control over details in low-noise environments, ensuring high-quality generated speech.
[0129] This asymmetric weighting strategy not only improves the model's generalization ability under complex text input, but also significantly improves the overall quality and naturalness of speech synthesis.
[0130] In some optional implementations of the embodiments of the present application, the audio Transformer model uses the XLM-RoBERTa model as a generator and trains it to generate speech duration based on text transcription. Specifically:
[0131] The input of the XLM-RoBERTa model is represented as follows: given a text transcription T, the model generates the corresponding speech duration D;
[0132] The generation process of the XLM-RoBERTa model is expressed as: using kernel sampling (p = 0.95) for autoregressive generation, where the process of autoregressive generation is expressed as:
[0133] D = XLM-RoBERTa(T;θ)
[0134] Among them, θ represents the parameters of the model;
[0135] The loss function of the XLM-RoBERTa model is expressed as:
[0136]
[0137] Where N represents the number of samples, D i Indicates the actual duration. Indicates the forecast duration.
[0138] In the embodiments of the present application, the method of the present application does not require phoneme duration prediction as an intermediate step, which may introduce errors. During the training process, the present application provides the diffusion network with noise potential vectors with the correct sequence length corresponding to the complete speech duration. During inference, the present application only specifies the overall duration without providing the duration of individual phonemes. In contrast, models such as NaturalSpeech2 and VoiceBox require an external phoneme duration prediction model.
[0139] The model of the present application learns to parse phoneme duration from text transcription in an end-to-end manner through diffusion training. In the duration prediction of the inference stage, the present application only needs to fine-tune the XLM-RoBERTa model to a random duration predictor and train it in a sequence-to-sequence manner. Given a text transcription, the model uses kernel sampling (p=0.95) to autoregressively generate speech duration (for example, "4.51 seconds") and achieves a root mean square error (RMSE) of 1.4 seconds. Importantly, the method of the present application is insensitive to the duration selection method, avoiding the cascade error caused by explicit phoneme duration modeling.
[0140] To achieve duration prediction, this application uses the XLM-RoBERTa model as a generator and trains it to generate speech duration D based on text transcription T. The specific steps are as follows:
[0141] 1. Model input: Given a text transcription T, the model generates the corresponding speech duration D.
[0142] 2. Generation process: Use kernel sampling (p = 0.95) for autoregressive generation. The generation process can be expressed as:
[0143] D = XLM-RoBERTa(T;θ)
[0144] Among them, θ represents the parameters of the model.
[0145] 3. Loss function: This application uses root mean square error (RMSE) as the loss function to optimize model parameters:
[0146]
[0147] Where N is the number of samples, D i Is the real duration, is the predicted duration.
[0148] In this way, the model of this application can directly generate speech duration during inference without relying on an external duration prediction model. This end-to-end approach not only simplifies the process but also avoids the error accumulation caused by explicit phoneme duration modeling.
[0149] In some optional implementations of the embodiments of the present application, in order to achieve the above-mentioned non-classifier guidance, the present application randomly discards text with a probability p = 0.1 during the training process, and jointly trains a conditional model and an unconditional model. In the inference phase, the present application introduces a sampling parameter w and calculates it in the following way:
[0150]
[0151] Among them, x cond is the output of the conditional model, x uncond is the output of the unconditional model, and w is the guidance weight. By adjusting the value of w, this application can control the degree to which the model relies on the condition during the generation process. When w > 1, the model will tend to generate results that are more consistent with the conditional information (such as text), while when w < 1, the model will tend to generate more general results.
[0152] During training, the loss function Contains the losses of conditional and unconditional models:
[0153]
[0154] Where x is the input data, t is the time step, ∈ is the noise, and ∈ θ is the noise predicted by the model, c is the conditional information (such as text), and λ t is the weight at time step t.
[0155] In some optional implementations of the present invention, the ability to generate speaker prompts is a valuable feature in text-to-speech (TTS) systems. This capability uses a reference audio clip to guide the generation of speech that matches the target speaker's voice characteristics. Diffusion models can implement this speaker-prompted TTS through audio restoration. This application employs a multi-task learning approach, simultaneously training a denoising network to support TTS synthesis of both text-only and speaker prompts.
[0156] During the training process, the application trains the network with probability p = 0.5 for audio restoration. The specific method is to concatenate the clean audio latent vector with the noise latent vector. The application samples a time length d and concatenates the beginning of the clean audio latent representation x[:d] with the noise latent vector z t The final part of [d:] is concatenated to form the input. Furthermore, we introduce a binary embedding to identify corrupted frames, which is added to the input after the initial projection. When computing the loss, we mask the frames corresponding to the clean audio.
[0157] For the duration of the prompt, we sample a proportion d∈[0,1] of the input and use it as a clean prompt. For example, if we sample d=0.1 for a 10-second audio segment, we use the frame corresponding to the first 1 second of audio as the clean prompt. To sample the duration, we use a Beta distribution with a mode of 0.01 and a density of 5 to highlight the challenging case of very short prompts. During the inference phase, we pre-add the speaker's reference audio and the corresponding text to perform TTS synthesis of the speaker prompt.
[0158] During training, the loss function of audio restoration can be expressed as:
[0159]
[0160] Among them, N is the number of samples, Mask(x i ) is a mask, ⊙ represents element-wise multiplication, Indicates splicing, x i It's clean audio, z t is the noise latent vector, and d is the sampling duration.
[0161] During the inference phase, this application performs speaker-prompted TTS synthesis through the following steps:
[0162] 1. Input preparation: concatenate the speaker’s reference audio and the corresponding text.
[0163] 2. Audio restoration: Use the trained denoising network to process the spliced input and generate audio that matches the voice characteristics of the target speaker.
[0164] 3. Output generation: Output the generated audio to complete the TTS synthesis of the speaker prompt.
[0165] In this way, the model of this application can support TTS synthesis of text only and speaker prompts simultaneously under the multi-task learning framework, and flexibly adapt to different application scenarios.
[0166] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0167] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0168] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0169] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0170] Further references Figure 4 , as a response to the above Figure 2 The present application provides an embodiment of a synthetic speech generating device, which is similar to the embodiment of the present invention. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0171] like Figure 4 As shown, the synthetic speech generation device 200 of the embodiment of the present application includes:
[0172] The request acquisition module 210 is configured to acquire a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data;
[0173] The text encoding module 220 is used to input the text data to be synthesized into the XLM encoder for text encoding operation to obtain XLM tag embedded data;
[0174] A downsampling module 230 is configured to perform a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data;
[0175] The position encoding module 240 is used to input the text mark position data into the position encoder for position encoding operation to obtain the mark position embedded data;
[0176] The audio prediction module 250 is used to input the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data into the audio Transformer model to perform an audio prediction operation to obtain audio prediction data.
[0177] In an embodiment of the present application, a synthetic speech generation device 200 is provided, comprising: a request acquisition module 210 for acquiring a synthetic speech generation request sent by a user terminal, wherein the synthetic speech generation request includes text data to be synthesized, speech data to be synthesized, and text tag position data; a text encoding module 220 for inputting the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedding data; a downsampling module 230 for downsampling the speech data to be synthesized to obtain downsampled registration tag data; a position encoding module 240 for inputting the text tag position data into a position encoder for position encoding operation to obtain tag position embedding data; an audio prediction module 250 for inputting the XLM tag embedding data, downsampled registration tag data, and tag position embedding data into an audio Transformer model for audio prediction operation to obtain audio prediction data. Compared with the prior art, the present application uses a latent diffusion model and a pre-trained autoencoder to map high-dimensional speech data into a compact latent space, thereby reducing dependence on labeled data. At the same time, the representation of the pre-trained language model is used as conditional information, so that it can effectively generalize to different text inputs under limited data. This method adopts a novel diffusion architecture that can efficiently process long audio sequences and improves model performance through a position-aware cross-attention mechanism.
[0178] In some optional implementations of the embodiments of the present application, the downsampling module 230 includes:
[0179] The speech encoding submodule is used to input the speech data to be synthesized into the SoundStream encoder for speech encoding operation to obtain potential speech data;
[0180] The high-resolution encoding submodule is used to input the potential speech data into the U-Net encoder for high-resolution encoding operation to obtain downsampled registration mark data.
[0181] To solve the above technical problems, the present application also provides a computer device. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device according to an embodiment of the present application.
[0182] The computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected through a system bus. It should be noted that the figure only shows the computer device 300 having components 310-330, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0183] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0184] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as a hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk equipped on the computer device 300, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 310 may also include both the internal storage unit of the computer device 300 and its external storage device. In the embodiment of the present application, the memory 310 is generally used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for the synthetic speech generation method. In addition, the memory 310 can also be used to temporarily store various data that has been output or is about to be output.
[0185] In some embodiments, the processor 320 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 320 is generally used to control the overall operation of the computer device 300. In the embodiment of the present application, the processor 320 is used to execute computer-readable instructions or process data stored in the memory 310, such as computer-readable instructions for executing the method for generating synthesized speech.
[0186] The network interface 330 may include a wireless network interface or a wired network interface. The network interface 330 is generally used to establish a communication connection between the computer device 300 and other electronic devices.
[0187] The computer device provided in this application uses a latent diffusion model and a pre-trained autoencoder to map high-dimensional speech data into a compact latent space, thereby reducing reliance on labeled data. Furthermore, the representation of the pre-trained language model is used as conditional information, enabling effective generalization to diverse text inputs with limited data. This method utilizes a novel diffusion architecture that efficiently processes long audio sequences and improves model performance through a position-aware cross-attention mechanism.
[0188] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the synthetic speech generation method as described above.
[0189] The computer-readable storage medium provided in this application uses a latent diffusion model and a pre-trained autoencoder to map high-dimensional speech data into a compact latent space, thereby reducing the reliance on annotated data. Furthermore, the representation of the pre-trained language model is used as conditioning information, enabling effective generalization to diverse text inputs with limited data. This method utilizes a novel diffusion architecture that efficiently processes long audio sequences and improves model performance through a position-aware cross-attention mechanism.
[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0191] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for generating synthetic speech, characterized in that: The steps include: Obtaining a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data; Inputting the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedded data; Inputting the speech data to be synthesized into a SoundStream encoder for speech encoding operation to obtain latent speech data; Inputting the potential speech data into a U-Net encoder for high-resolution encoding to obtain the downsampled registration mark data; Inputting the text mark position data into a position encoder for position encoding operation to obtain mark position embedded data; The XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data are input into an audio Transformer model for performing an audio prediction operation to obtain audio prediction data.
2. The method for generating synthetic speech according to claim 1, wherein: The step of performing a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data specifically includes the following steps: Inputting the speech data to be synthesized into a SoundStream encoder for speech encoding operation to obtain latent speech data; The potential speech data is input into a U-Net encoder for high-resolution encoding to obtain the downsampled registration mark data.
3. The method for generating synthetic speech according to claim 1, wherein: After the step of inputting the XLM tag embedding data, the downsampled registration tag data, and the tag position embedding data into the audio Transformer model for performing an audio prediction operation to obtain audio prediction data, the following step is also included: The audio prediction data is input into the SoundStream decoder to perform an audio sampling point restoration operation to obtain target audio prediction data.
4. The method for generating synthetic speech according to claim 3, wherein: The step of inputting the audio prediction data into a SoundStream decoder to perform an audio sampling point restoration operation to obtain target audio prediction data specifically includes the following steps: Inputting the audio prediction data into a U-Net decoder for audio decoding operation to obtain denoised audio prediction data; The denoised audio prediction data is input into a SoundStream decoder to perform an audio sampling point restoration operation to obtain the target audio prediction data.
5. The method for generating synthetic speech according to claim 1, wherein: The audio Transformer model is trained according to an asymmetric diffusion loss weighting strategy, wherein the asymmetric diffusion loss weighting strategy is expressed as: Among them, λ t represents the noise level, γ and x0 are the parameters of the Cauchy distribution, and μ and σ are the parameters of the normal distribution.
6. The method for generating synthetic speech according to claim 1, wherein: The audio Transformer model uses the XLM-RoBERTa model as a generator and trains it to generate speech duration based on text transcription. Specifically: The input of the XLM-RoBERTa model is represented as follows: given a text transcription T, the model generates the corresponding speech duration D; The generation process of the XLM-RoBERTa model is expressed as: using kernel sampling (p=0.95) to perform autoregressive generation, wherein the process of the autoregressive generation is expressed as: D = XLM-RoBERTa(T;θ) Among them, θ represents the parameters of the model; The loss function of the XLM-RoBERTa model is expressed as: Where N represents the number of samples, D i Indicates the actual duration. Indicates the forecast duration.
7. A synthetic speech generating device, characterized in that: include: a request acquisition module, configured to acquire a synthesized speech generation request sent by a user terminal, wherein the synthesized speech generation request includes text data to be synthesized, speech data to be synthesized, and text mark position data; A text encoding module is used to input the text data to be synthesized into an XLM encoder for text encoding operation to obtain XLM tag embedded data; A downsampling module, configured to perform a downsampling operation on the speech data to be synthesized to obtain downsampled registration mark data; A position encoding module, configured to input the text mark position data into a position encoder for performing a position encoding operation to obtain mark position embedded data; The audio prediction module is used to input the XLM tag embedding data, the downsampled registration tag data and the tag position embedding data into the audio Transformer model to perform an audio prediction operation to obtain audio prediction data.
8. The synthetic speech generating device according to claim 7, wherein: The downsampling module includes: A speech encoding submodule, configured to input the speech data to be synthesized into a SoundStream encoder for speech encoding to obtain latent speech data; The high-resolution encoding submodule is used to input the potential speech data into a U-Net encoder for high-resolution encoding operation to obtain the downsampled registration mark data.
9. A computer device comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the synthetic speech generation method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the synthetic speech generation method according to any one of claims 1 to 6.