Speech synthesis method, diffusion model training method, device and equipment

By using a diffusion model for multi-scale acoustic feature extraction and residual learning, the problem of high-frequency speech detail loss in text-to-speech systems has been solved, improving speech synthesis performance and promoting the development of financial and healthcare businesses.

CN120998174APending Publication Date: 2025-11-21PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511066167.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing text-to-speech systems are prone to losing high-frequency speech details during speech synthesis, resulting in poor speech synthesis quality and hindering the progress of financial and healthcare businesses.

Method used

A diffusion model is used for text encoding, acoustic feature extraction, text-speech alignment, and high-frequency compensation. The speech synthesis effect is improved by multi-scale acoustic feature extraction and residual learning of the diffusion model.

Benefits of technology

It enhances the prosodic balance and speech quality of speech synthesis, improves the speech synthesis effect, and promotes the development of financial and healthcare businesses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998174A_ABST
    Figure CN120998174A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech synthesis, is used for artificial intelligence scenes, financial business scenes and medical health business scenes, and provides a speech synthesis method, a diffusion model training method, a device and equipment, and the method comprises the steps: obtaining text information; carrying out coding processing on the text information based on a text coding sub-model of a diffusion model to obtain a vector sequence; based on an acoustic feature extraction sub-model, performing acoustic feature extraction processing on the vector sequence to obtain a first Mel spectrum; based on a context sensing sub-model, according to the vector sequence, performing text-voice alignment processing on the first Mel spectrum to obtain a second Mel spectrum; performing residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum based on a high-frequency compensation diffusion sub-model to obtain a target Mel spectrum; and based on the diffusion sub-model, according to the target Mel spectrum, determining the target voice so as to improve the voice synthesis effect and further promote the development of financial services and medical health services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and in particular to a speech synthesis method, a training method for a diffusion model, an apparatus, and a device. Background Technology

[0002] Currently, most text-to-speech (TTS) systems employ a phased architecture consisting of text encoding, acoustic models, and vocoders. During the waveform reconstruction stage, TTS relies on traditional vocoders, which can easily lead to significant loss of high-frequency speech details (such as fricatives and sibilants), resulting in synthesized speech that often sounds mechanical or metallic, ultimately leading to poor speech synthesis performance in TTS systems.

[0003] Since TTS systems can be deployed in different business scenarios, such as financial business scenarios and healthcare business scenarios, poor speech synthesis performance of TTS systems can easily have an adverse impact on the business progress of financial business scenarios and healthcare business scenarios. Summary of the Invention

[0004] The main purpose of this application is to provide a speech synthesis method, a training method for a diffusion model, an apparatus and device, which aim to improve the speech synthesis effect.

[0005] In a first aspect, this application provides a speech synthesis method, the speech synthesis method comprising:

[0006] Obtain the text information to be processed;

[0007] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0008] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0009] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0010] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0011] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum.

[0012] Secondly, this application also provides a training method for a diffusion model, the training method comprising:

[0013] Obtain a training sample set, which includes multiple text information to be processed and speech tags corresponding to each text information;

[0014] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0015] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0016] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0017] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0018] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum;

[0019] Based on the discriminator of the diffusion model, the model optimization parameters of the diffusion model are determined according to the target speech and speech tags;

[0020] Based on the model optimization parameters, the model parameters of the diffusion model are adjusted to obtain the optimized diffusion model.

[0021] Thirdly, this application also provides a speech synthesis device, the speech synthesis device comprising:

[0022] The text acquisition module is used to acquire the text information to be processed.

[0023] The encoding module is used to encode the text information based on the diffusion model sub-model to obtain the vector sequence corresponding to the text information;

[0024] The feature extraction module is used to perform acoustic feature extraction processing on the vector sequence based on the acoustic feature extraction sub-model of the diffusion model to obtain the first Mel spectrum corresponding to the vector sequence;

[0025] The text-speech alignment module is used to perform text-speech alignment processing on the first Mel spectrum based on the context-aware sub-model of the diffusion model and the vector sequence to obtain the corresponding second Mel spectrum.

[0026] The high-frequency compensation module is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum based on the high-frequency compensation diffusion sub-model of the diffusion model to obtain the corresponding target Mel spectrum.

[0027] The speech synthesis module is used to determine the target speech corresponding to the text information based on the diffusion sub-model of the diffusion model and the target Mel spectrum.

[0028] Fourthly, this application also provides a computer device, the computer device including a memory and a processor;

[0029] The memory is used to store computer programs;

[0030] The processor is configured to execute the computer program and, in executing the computer program, implement the steps of the speech synthesis method as described above, or implement the steps of the diffusion model training method as described above.

[0031] Fifthly, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the speech synthesis method described above, or the steps of the diffusion model training method described above.

[0032] This application provides a speech synthesis method, a training method for a diffusion model, an apparatus, and a device. The speech synthesis method includes: acquiring text information to be processed; encoding the text information using a text encoding sub-model based on a diffusion model to obtain a vector sequence corresponding to the text information; performing acoustic feature extraction processing on the vector sequence using an acoustic feature extraction sub-model based on a diffusion model to obtain a first Mel spectrum corresponding to the vector sequence; performing text-speech alignment processing on the first Mel spectrum based on a context-aware sub-model based on a diffusion model, to obtain a corresponding second Mel spectrum; performing residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum using a diffusion sub-model based on a diffusion model to obtain a corresponding target Mel spectrum; and determining the target speech corresponding to the text information based on the target Mel spectrum using a diffusion sub-model based on a diffusion model.

[0033] Given the text information to be processed, a diffusion model can be used to encode the text information, extract acoustic features from the corresponding vector sequence, perform text-to-speech alignment on the first Mel spectrum, perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum, and then determine the target speech corresponding to the text information based on the target Mel spectrum. In the process of determining the target speech using the diffusion model, the context-aware sub-model constrains the first Mel spectrum for text-to-speech alignment, which can improve the prosodic balance of the second Mel spectrum. This, in turn, can improve the prosodic balance of the target speech determined by the diffusion model, effectively enhancing the speech quality of the target speech and thus improving the speech synthesis effect. This improved speech synthesis effect is beneficial for promoting the development of financial and healthcare businesses. For example, for text information involved in financial transactions, such as payment information, sales information, leasing information, insurance policy information, pension information, etc., this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to more flexibly inform users of the text information involved in financial transactions, thereby promoting the development of financial businesses. For text information involved in healthcare services, such as health monitoring information, consultation results, health records, examination reports, etc., this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to users in a more flexible way, thereby promoting the development of healthcare services. Attached Figure Description

[0034] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of an application environment for a speech synthesis method according to an embodiment of this application;

[0036] Figure 2 This is a flowchart illustrating a speech synthesis method in one embodiment of this application;

[0037] Figure 3 This is a flowchart illustrating a training method for a diffusion model in one embodiment of this application;

[0038] Figure 4 This is a schematic diagram of a speech synthesis device in one embodiment of this application;

[0039] Figure 5This is a schematic diagram of the structure of a computer device according to one embodiment of this application;

[0040] Figure 6 This is another structural schematic diagram of a computer device in one embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] The speech synthesis method provided in this application embodiment can be applied to, for example, Figure 1In this application environment, the client communicates with the server via a network. The server can obtain the text information to be processed from the client, and encode the text information based on the text encoding sub-model of the diffusion model to obtain the corresponding vector sequence; perform acoustic feature extraction processing on the vector sequence based on the acoustic feature extraction sub-model of the diffusion model to obtain the first Mel spectrum corresponding to the vector sequence; perform text-speech alignment processing on the first Mel spectrum based on the context-aware sub-model of the diffusion model according to the vector sequence to obtain the corresponding second Mel spectrum; perform residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum based on the high-frequency compensation diffusion sub-model of the diffusion model to obtain the corresponding target Mel spectrum; and determine the target speech corresponding to the text information based on the target Mel spectrum. The server can then feed back the target speech to the client so that the client can play the target speech to inform the user of the text information in the form of voice. In this application, the diffusion model includes an artificial intelligence model that sequentially encodes, extracts acoustic features, performs text-to-speech alignment, performs residual learning, and extracts multi-scale acoustic features based on text information to obtain the corresponding target Mel spectrum. Then, the diffusion model, combined with the target Mel spectrum, determines the target speech corresponding to the text information. In the process of determining the target speech corresponding to the text information using the diffusion model, the context-aware sub-model constrains the first Mel spectrum for text-to-speech alignment, which can improve the prosodic balance of the second Mel spectrum. This, in turn, can improve the prosodic balance of the target speech determined by the diffusion model, effectively enhancing the speech quality of the target speech and thus improving the speech synthesis effect. This improved speech synthesis effect is beneficial for promoting the development of financial and healthcare businesses. For example, for text information involved in financial transactions, such as payment information, sales information, leasing information, insurance policy information, pension information, etc., this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to more flexibly inform users of the text information involved in financial transactions, thereby promoting the development of financial businesses. For textual information involved in healthcare services, such as health monitoring information, consultation results, health records, and examination reports, this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to users in a more flexible way, thereby promoting the development of healthcare services. However, this is not limited to this; the server can also deploy a diffusion model to the client, allowing the client to use this speech synthesis method to determine the target speech corresponding to the text information when it receives it. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.The present application will now be described in detail through specific embodiments.

[0043] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the speech synthesis method provided in this application embodiment includes the following steps:

[0044] S101: Obtain the text information to be processed.

[0045] The text information to be processed can vary depending on the specific business scenario, such as financial business scenarios or healthcare business scenarios.

[0046] For example, in the context of financial transactions, text information can include payment information, sales information, leasing information, insurance policy information, pension information, and so on. In the context of healthcare transactions, text information can include health monitoring information, consultation results, health records, examination reports, and so on.

[0047] For example, text messages can be sent to the computer device used by the user. However, there may be situations where it is inconvenient for the user to view the text message while using the computer device, such as when the user is busy with other tasks and has no time to look at the computer device, or when the user cannot clearly see the text message displayed on the computer device. In such cases, it is necessary to inform the user of the text message in a form other than text, such as voice.

[0048] If the text information to be processed is obtained, it can be used to determine the corresponding voice information, which will help improve the convenience of TTS processing of the text information.

[0049] S102: A text encoding sub-model based on the diffusion model, which encodes text information to obtain a vector sequence corresponding to the text information.

[0050] Once text information is obtained, it can be input into the diffusion model. The diffusion model, upon receiving the text information, can encode the text information using a text encoding sub-model to obtain a vector sequence corresponding to the text information. This vector sequence is the one required by the diffusion model for TTS processing of the text information, and can then be used by the diffusion model to subsequently determine the target speech corresponding to the text information.

[0051] Taking text information, including insurance policy information related to financial transactions, as an example, when the policy information is input into the text encoding sub-model of the diffusion model, the text encoding sub-model can encode the policy information to obtain a vector sequence corresponding to the policy information. This vector sequence can then be used by the diffusion model to subsequently determine the target speech corresponding to the policy information.

[0052] Taking textual information including consultation results related to healthcare as an example, when the consultation results are input into the text encoding sub-model of the diffusion model, the text encoding sub-model can encode the consultation results to obtain a vector sequence corresponding to the consultation results. This vector sequence can then be used by the diffusion model to subsequently determine the target speech corresponding to the consultation results. And so on.

[0053] When the text encoding sub-model of the diffusion model can be used to encode text information and obtain the vector sequence corresponding to the text information, it is beneficial to improve the ease of determining the vector sequence corresponding to the text information, so as to improve the ease of TTS processing of text information by the diffusion model in the future.

[0054] S103: Acoustic feature extraction sub-model based on diffusion model, which performs acoustic feature extraction processing on vector sequence to obtain the first Mel spectrum corresponding to the vector sequence.

[0055] In some implementations, when the vector sequence corresponding to the text information is determined, the vector sequence can be input into the acoustic feature extraction sub-model of the diffusion model to perform semantically guided diffusion on the vector sequence. For example, the acoustic feature extraction sub-model can perform acoustic feature extraction processing on the vector sequence according to the first step number to obtain the first Mel spectrum corresponding to the vector sequence. The first step number is less than or equal to a preset step number threshold. For example, the preset step number threshold may include 50 steps, but it is not limited to this, and no limitation is made here. Since the first step number is less than or equal to the preset step number threshold, the acoustic feature extraction sub-model can extract the corresponding acoustic features from the vector sequence with fewer steps to obtain the first Mel spectrum corresponding to the vector sequence. The first Mel spectrum can also be called a coarse-grained Mel spectrum. The coarse-grained Mel spectrum is equivalent to extracting more global and generalized spectral features from the vector sequence, while ignoring fine harmonic or formant details; furthermore, the coarse-grained Mel spectrum can reduce the dimensionality of the vector sequence, simplifying the input of the diffusion model; and the coarse-grained Mel spectrum can emphasize the overall shape of the spectral envelope rather than local peaks. The first Mel spectrum can be used by the diffusion model to subsequently determine the target speech corresponding to the text information. Understandably, since the vector sequence is determined by the diffusion model based on the text information, the first Mel spectrum corresponding to the vector sequence can also be called the first Mel spectrum corresponding to the text information.

[0056] Taking the vector sequence corresponding to text information, including the vector sequence corresponding to policy information involved in financial transactions, as an example, when the vector sequence corresponding to the policy information is input into the acoustic feature extraction sub-model of the diffusion model, the acoustic feature extraction sub-model can perform acoustic feature extraction processing on the vector sequence corresponding to the policy information to obtain the first Mel spectrum corresponding to the policy information. The first Mel spectrum corresponding to the policy information can be used by the diffusion model to subsequently determine the target speech corresponding to the policy information.

[0057] Taking the vector sequence corresponding to text information, including the vector sequence corresponding to consultation results in medical and health services, as an example, when the vector sequence corresponding to the consultation results is input into the acoustic feature extraction sub-model of the diffusion model, the acoustic feature extraction sub-model can perform acoustic feature extraction processing on the vector sequence corresponding to the consultation results to obtain the first Mel spectrum corresponding to the consultation results. The first Mel spectrum corresponding to the consultation results can be used by the diffusion model to subsequently determine the target speech corresponding to the consultation results. And so on.

[0058] When the acoustic feature extraction sub-model of the diffusion model can be used to perform acoustic feature extraction processing on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence, it is beneficial to improve the convenience of determining the first Mel spectrum corresponding to the vector sequence. The first Mel spectrum can be used by the diffusion model to subsequently determine the target speech corresponding to the text information, which is beneficial to improve the convenience of the diffusion model to perform TTS processing on the text information.

[0059] S104: A context-aware sub-model based on a diffusion model, which performs text-speech alignment processing on the first Mel spectrum according to the vector sequence to obtain the corresponding second Mel spectrum.

[0060] In some implementations, when the first Mel spectrum corresponding to the vector sequence is determined, the vector sequence and the first Mel spectrum can be input into the context-aware sub-model of the diffusion model, so that the context-aware sub-model can perform text-speech alignment processing on the first Mel spectrum based on the vector sequence to obtain the corresponding second Mel spectrum.

[0061] For example, in the process of the context-aware sub-model performing text-to-speech alignment on the first Mel spectrum based on the vector sequence to obtain the corresponding second Mel spectrum, the context-aware sub-model can perform phoneme segmentation on the vector sequence to obtain the phonemes included in the vector sequence. Accordingly, the context-aware sub-model can predict the duration of each phoneme included in the vector sequence to obtain the duration corresponding to each phoneme. Furthermore, the context-aware sub-model can perform prosodic modeling on each phoneme included in the vector sequence to determine the prosodic pattern corresponding to each phoneme. Having determined the duration of the Mel spectrum corresponding to each phoneme in the vector sequence and the prosodic pattern corresponding to each phoneme, the context-aware sub-model can perform text-to-speech alignment on the Mel spectrum corresponding to each phoneme in the first Mel spectrum based on the duration and prosodic pattern of each phoneme to obtain the second Mel spectrum.

[0062] For example, when predicting the duration of each phoneme in a vector sequence, the context-aware model can predict how many frames of Mel spectrum each phoneme should last based on the contextual semantics of the vector sequence, thus obtaining the duration of each phoneme. When modeling the prosody of each phoneme in a vector sequence, the context-aware model can infer the prosody of each phoneme based on the contextual semantics of the vector sequence, such as the presence of stress, intonation, pauses, etc., thereby determining the prosody of each phoneme. When performing text-speech alignment processing on the Mel spectrum corresponding to the first Mel spectrum of each phoneme based on its duration and prosody, the context-aware model can dynamically stretch or compress the corresponding Mel spectrum based on its duration and prosody, thus obtaining the second Mel spectrum.

[0063] Taking the first Mel spectrum corresponding to textual information, including the first Mel spectrum corresponding to policy information related to financial transactions, as an example, when the vector sequence corresponding to the policy information and the first Mel spectrum corresponding to the policy information are input into the context-aware sub-model of the diffusion model, the context-aware sub-model can determine the contextual semantics corresponding to the policy information based on the vector sequence. Then, based on the contextual semantics, it performs text-to-speech alignment processing on the first Mel spectrum corresponding to the policy information to obtain the second Mel spectrum corresponding to the policy information. The second Mel spectrum corresponding to the policy information can then be used by the diffusion model to subsequently determine the target speech corresponding to the policy information.

[0064] Taking the first Mel spectrum corresponding to textual information, including the first Mel spectrum corresponding to consultation results in healthcare business, as an example. When the vector sequence corresponding to the consultation result and the first Mel spectrum corresponding to the consultation result are input into the context-aware sub-model of the diffusion model, the context-aware sub-model can determine the contextual semantics corresponding to the consultation result based on the vector sequence. Then, based on the contextual semantics, it performs text-to-speech alignment processing on the first Mel spectrum corresponding to the consultation result to obtain the second Mel spectrum corresponding to the consultation result. The second Mel spectrum corresponding to the consultation result can be used by the diffusion model to subsequently determine the target speech corresponding to the consultation result. And so on.

[0065] In a diffusion model where the context-aware sub-model can perform text-to-speech alignment on the first Mel spectrum based on the vector sequence to obtain the corresponding second Mel spectrum, the context-aware sub-model can also perform text-to-speech alignment on the first Mel spectrum based on the contextual semantics corresponding to the vector sequence to obtain the second Mel spectrum. This is beneficial for improving the prosodic balance of the second Mel spectrum. The second Mel spectrum can then be used by the subsequent diffusion model to determine the target speech corresponding to the text information, which is beneficial for improving the prosodic balance of the target speech. In other words, the diffusion model can enhance the speech quality of the target speech, thereby improving the speech synthesis effect.

[0066] S105: A high-frequency compensated diffusion sub-model based on a diffusion model is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0067] In some implementations, when a second Mel spectrum corresponding to the text information is determined, the second Mel spectrum can be input into the high-frequency compensation diffusion sub-model of the diffusion model, and the high-frequency compensation diffusion sub-model can be used to perform detail enhancement diffusion on the second Mel spectrum. Since the second Mel spectrum is determined based on the first Mel spectrum, if the first Mel spectrum is a coarse-grained Mel spectrum, the second Mel spectrum is also a coarse-grained Mel spectrum.

[0068] For example, the high-frequency compensation diffusion sub-model can be conditioned on a second Mel spectrum and repair spectral details involved in the second Mel spectrum through residual learning. In an exemplary embodiment, the high-frequency compensation diffusion sub-model can be conditioned on a second Mel spectrum and focus on repairing high-frequency details through residual learning to determine the target Mel spectrum. High-frequency details can include frequency band details above 2kHz, such as voiceless fricatives and plosives like / s / , / sh / , / f / , / th / , / t / , / k / , and / p / . High-frequency details can improve consonant clarity, timbre, and brightness, etc.

[0069] Accordingly, the high-frequency compensation diffusion sub-model can use the second Mel spectrum as a condition and employ an encoder to perform multi-scale acoustic feature extraction on the second Mel spectrum, thereby extracting multi-scale acoustic features containing rich local details and global context from the second Mel spectrum. In an exemplary embodiment, the encoder may include a hybrid encoder, such as a convolutional neural network-transformer (CNN-Transformer) hybrid encoder. During the multi-scale acoustic feature extraction of the second Mel spectrum using the CNN-Transformer hybrid encoder, the high-frequency compensation diffusion sub-model can utilize the CNN encoder to capture the spectral texture within the second Mel spectrum, the short-term correlation between adjacent frames, and efficiently extract finer-grained phoneme features to coarser-grained phoneme features by stacking convolutional layers and downsampling, in order to determine the target Mel spectrum. Correspondingly, the self-attention mechanism involved in the Transformer encoder can also be used to model the long-range dependencies between spectra at any position in the second Mel spectrum, in order to determine the target Mel spectrum.

[0070] The high-frequency compensated diffusion sub-model, after performing residual learning and multi-scale acoustic feature extraction on the second Mel spectrum, can obtain the corresponding target Mel spectrum. The target Mel spectrum can also be called the fine-grained Mel spectrum.

[0071] Taking the second Mel spectrum corresponding to textual information, including the second Mel spectrum corresponding to policy information related to financial transactions, as an example, when the second Mel spectrum corresponding to the policy information is input into the high-frequency compensated diffusion sub-model of the diffusion model, the high-frequency compensated diffusion sub-model can perform residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum corresponding to the policy information to obtain the target Mel spectrum corresponding to the policy information. The target Mel spectrum corresponding to the policy information can then be used by the diffusion model to subsequently determine the target speech corresponding to the policy information.

[0072] Taking the second Mel spectrum corresponding to textual information, including the second Mel spectrum corresponding to consultation results in healthcare business, as an example, when the second Mel spectrum corresponding to the consultation result is input into the high-frequency compensated diffusion sub-model of the diffusion model, the high-frequency compensated diffusion sub-model can perform residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum corresponding to the consultation result to obtain the target Mel spectrum corresponding to the consultation result. The target Mel spectrum corresponding to the consultation result can be used by the diffusion model to subsequently determine the target speech corresponding to the consultation result. And so on.

[0073] In a diffusion model where the high-frequency compensated diffusion sub-model can be used for residual learning and multi-scale acoustic feature extraction of the second Mel spectrum to obtain the corresponding target Mel spectrum, the ease of determining the target Mel spectrum corresponding to the text information is improved. The target Mel spectrum can then be used by the diffusion model to determine the target speech corresponding to the text information, thus enhancing the ease of TTS processing of the text information by the diffusion model. Furthermore, since the target Mel spectrum is obtained by enhancing the details of the second Mel spectrum through the high-frequency compensated diffusion sub-model, the diffusion model's subsequent determination of the target speech corresponding to the text information based on the target Mel spectrum improves the speech synthesis effect.

[0074] S106: A diffusion sub-model based on the diffusion model, which determines the target speech corresponding to the text information according to the target Mel spectrum.

[0075] In some implementations, once the target Mel spectrum corresponding to the text information is determined, the target Mel spectrum can be input into the diffusion sub-model of the diffusion model. Upon obtaining the target Mel spectrum, the diffusion sub-model can perform forward and backward diffusion on the target Mel spectrum, thereby disrupting and reconstructing it. Then, the diffusion sub-model can use a vocoder to convert the reconstructed target Mel spectrum into a corresponding audio waveform, and determine the target speech corresponding to the text information based on the audio waveform.

[0076] Taking the target Mel spectrum corresponding to text information, including the target Mel spectrum corresponding to insurance policy information involved in financial business, as an example, when the target Mel spectrum corresponding to the insurance policy information is input into the diffusion sub-model of the diffusion model, the diffusion sub-model can output the target speech corresponding to the text information based on the target Mel spectrum corresponding to the insurance policy information.

[0077] Taking the target Mel spectrum corresponding to text information, including the target Mel spectrum corresponding to consultation results in healthcare business, as an example, when the target Mel spectrum corresponding to the consultation result is input into the diffusion sub-model of the diffusion model, the diffusion sub-model can output the target speech corresponding to the consultation result based on the target Mel spectrum corresponding to the consultation result. And so on.

[0078] When the diffusion sub-model of the diffusion model can determine the target speech corresponding to the text information based on the target Mel spectrum, it improves the ease of determining the target speech, thus enhancing the ease of TTS processing of the text information by the diffusion model. Since the target Mel spectrum is obtained through acoustic feature extraction, text-speech alignment, residual learning, and multi-scale acoustic feature extraction of the vector sequence corresponding to the text information, when the diffusion model determines the target speech based on the target Mel spectrum, it helps improve the prosodic balance of the target speech. This means the diffusion model can enhance the speech quality of the target speech, thereby improving the speech synthesis effect.

[0079] In some implementations, a text semantic coding network based on a text coding sub-model performs semantic feature encoding on the text information to obtain a semantic vector corresponding to the text information; a speech feature coding network based on a text coding sub-model performs speech feature encoding on the text information to obtain a speech feature vector corresponding to the text information; and a vector fusion network based on a text coding sub-model performs cross-modal fusion processing on the semantic vector and the speech feature vector to obtain a vector sequence corresponding to the text information.

[0080] For example, the text encoding sub-model of the diffusion model, upon acquiring text information, can perform different forms of encoding processing on the text information to obtain different types of vectors, based on the consideration of improving the expressiveness of the target speech corresponding to the subsequently determined text information. Accordingly, the text encoding sub-model can perform cross-modal fusion processing on different types of vectors to obtain the vector sequence corresponding to the text information.

[0081] The text encoding sub-model can perform different forms of encoding processing on text information, including semantic feature encoding processing and speech feature encoding processing.

[0082] The text encoding sub-model can utilize a text semantic encoding network to encode semantic features of text information, obtaining a semantic vector corresponding to the text information. This semantic vector can indicate at least one of the semantically related information of the text information, such as its grammatical structure and semantic focus. For example, the text semantic encoding network can include a Bert encoder. The text encoding sub-model can use the Bert encoder to identify the grammatical structure and semantic focus of the text information, obtaining the corresponding semantic vector. For instance, the Bert encoder might identify that the text information has at least one grammatical structure, such as an interrogative sentence, a declarative sentence, or an imperative sentence. Another example is that the Bert encoder might identify that the semantic focus of the text information includes at least one keyword. For the semantic focus of the text information, the Bert encoder can determine that at least one emphasis action, such as intonation or stress, is required.

[0083] The text encoding sub-model can utilize a speech feature encoding network to process text information using speech features, obtaining a speech feature vector corresponding to the text information. This speech feature vector can indicate at least one of the emotional polarity and intensity of the text information. For example, the speech feature encoding network can include a Whisper encoder. The text encoding sub-model can use the Whisper encoder to identify at least one of the emotional polarity and intensity of the text information, obtaining the corresponding speech feature vector. For instance, the Whisper encoder identifies at least one of the emotional polarities such as surprise, doubt, anger, and sadness in the text information. Furthermore, based on the identified emotional polarity, the Whisper encoder can determine the intensity of that emotional polarity, such as the intensity of surprise, doubt, anger, sadness, etc., without limitation.

[0084] In one exemplary implementation, the text encoding sub-model may include a Bert-Whisper co-encoder to simultaneously extract semantic vectors corresponding to text information and speech feature vectors corresponding to text information, without limitation.

[0085] In the text encoding sub-model, since the semantic vectors and speech feature vectors corresponding to the text information are vectors of different modalities (i.e., different types of vectors), the text encoding sub-model can utilize a vector fusion network to perform cross-modal fusion processing on the semantic vectors and speech feature vectors, thereby obtaining a vector sequence corresponding to the text information. For example, the vector fusion network can dynamically determine the weights of the semantic vectors and speech feature vectors, and then perform vector fusion processing based on these weights to obtain the corresponding vector sequence.

[0086] When the text semantic encoding network, speech feature encoding network, and vector fusion network of the text encoding sub-model are used to determine the vector sequence corresponding to the text information, the vector sequence can be used to indicate the grammatical structure and semantic focus of the text information, as well as at least one of the sentiment polarity and sentiment polarity intensity. The diffusion model then combines the vector sequence to determine the target speech corresponding to the text information. This, along with the grammatical structure, semantic focus, sentiment polarity, and / or sentiment polarity intensity indicated by the vector sequence, helps to determine the target speech, ensuring that the target speech reflects the text semantics, emotional intent, etc., thereby improving the speech synthesis effect.

[0087] In some implementations, a forward diffusion network based on a diffusion sub-model performs noise diffusion on the target Mel spectrum according to the text complexity corresponding to the text information, thereby obtaining a noise diffusion sample corresponding to the target Mel spectrum; a backward diffusion network based on a diffusion sub-model performs noise prediction on the noise diffusion sample according to the vector sequence, thereby obtaining the target speech corresponding to the noise diffusion sample.

[0088] For example, when the diffusion sub-model of the diffusion model obtains the target Mel spectrum, it can use a forward diffusion network to diffuse noise into the target Mel spectrum, gradually adding noise to obtain noise-diffused samples corresponding to the target Mel spectrum. During the noise diffusion process using the forward diffusion network, the diffusion sub-model can dynamically adjust the noise addition ratio based on the text complexity of the text information, thus diffusing noise into the target Mel spectrum according to the noise addition ratio to obtain the corresponding noise-diffused samples. For instance, if the text complexity is less than or equal to a preset complexity threshold, the text information can be determined to be simple text, and the forward diffusion network can use cosine scheduling to diffuse noise into the target Mel spectrum for rapid convergence. Conversely, if the text complexity is greater than a preset complexity threshold, the text information can be determined to be complex text, and the forward diffusion network can use exponential scheduling to diffuse noise into the target Mel spectrum to enhance detail generation. In the process of determining the noise diffusion sample corresponding to the target Mel spectrum based on the target Mel spectrum, the forward diffusion network can have different noise addition ratios for different target Mel spectra, or they can have the same noise addition ratio, which helps to improve the flexibility of determining the noise diffusion sample corresponding to the target Mel spectrum.

[0089] Given a noise diffusion sample corresponding to the target Mel spectrum, the forward diffusion network can output the noise diffusion sample to the inverse diffusion network of the diffusion sub-model. The inverse diffusion network can predict noise from the noise diffusion sample using the vector sequence as a condition, gradually reconstructing the target Mel spectrum from the noise diffusion sample to obtain the predicted Mel spectrum. Correspondingly, the inverse diffusion network can use a vocoder to convert the predicted Mel spectrum into an audio waveform and save the audio waveform as the target speech, i.e., the target speech corresponding to the noise diffusion sample. The target speech is also the target speech corresponding to the text information. The target speech can be saved as an audio file. Since the audio file is playable, users can play it according to their needs to listen to the target speech. When the diffusion model can utilize the inverse diffusion network of the diffusion sub-model, combined with the noise diffusion sample and the vector sequence corresponding to the text information, to determine the target speech corresponding to the text information, it improves the convenience of determining the target speech.

[0090] In some implementations, a lip-movement synchronization coding network based on a diffusion sub-model obtains the target lip shape parameters corresponding to the target lip spectrum, and updates the noise diffusion samples according to the target lip shape parameters to obtain the updated noise diffusion samples.

[0091] For example, when the diffusion sub-model of the diffusion model uses the forward diffusion network to determine the noise diffusion sample corresponding to the target Mel spectrum, it can update the noise diffusion sample based on the consideration of improving the expressiveness of the target speech corresponding to the text information determined subsequently. This update can be achieved by incorporating information related to the physical movement of the user's vocal organs, such as lip movement information, that may be involved in the target speech corresponding to the text information. This ensures that when the diffusion model subsequently determines the target speech corresponding to the text information based on the updated noise diffusion sample, the target speech can reflect the physical movement of the user's vocal organs, such as lip movement.

[0092] A lip-sync coding network can be introduced into the diffusion sub-model. This network allows the diffusion sub-model to obtain the target lip shape parameters corresponding to the target Mel spectrum. These target lip parameters can include 3D lip parameters. The target lip parameters corresponding to the target Mel spectrum can be used as diffusion conditions to update the noise diffusion samples, resulting in updated noise diffusion samples that reflect the user's lip movements. This improves the synchronization between lip movements and target speech during TTS processing of text information.

[0093] The inverse diffusion network based on the diffusion sub-model performs noise prediction on the updated noise diffusion samples according to the vector sequence, and obtains the target speech corresponding to the noise diffusion samples.

[0094] When the lip-sync coding network determines the updated noise diffusion sample, it can output the updated noise diffusion sample to the inverse diffusion network of the diffusion sub-model. The inverse diffusion network can use the vector sequence as a condition to perform noise prediction on the updated noise diffusion sample, gradually reconstructing the target Mel spectrum from the updated noise diffusion sample to obtain the predicted Mel spectrum. Correspondingly, the inverse diffusion network can use a vocoder to convert the predicted Mel spectrum into an audio waveform and save the audio waveform as the target speech, i.e., the target speech corresponding to the noise diffusion sample. Correspondingly, the target speech is also the target speech corresponding to the text information. When the diffusion model can use the inverse diffusion network of the diffusion sub-model, combined with the updated noise diffusion sample and the vector sequence corresponding to the text information, to determine the target speech corresponding to the text information, since the updated noise diffusion sample reflects the user's lip movements, the target speech corresponding to the text information determined by the diffusion model can be synchronized with the user's lip movements, i.e., the user's mouth shape, thus improving the naturalness of the target speech and consequently enhancing the speech synthesis effect.

[0095] In some implementations, based on a reverse diffusion network, noise prediction is performed on the noise diffusion samples according to the vector sequence and at least one of global semantic tokens and local tokens to obtain the corresponding target speech.

[0096] For example, when the backdiffusion network obtains the noise diffusion sample corresponding to the target Mel spectrum, it can perform noise prediction on the noise diffusion sample to obtain the corresponding target speech. Accordingly, during the noise prediction process of the noise diffusion sample, the backdiffusion network can perform at least one of the following on the target speech corresponding to the noise diffusion sample: consistency control and instantaneous spectral anomaly repair, thereby determining the corresponding target speech.

[0097] For example, during noise prediction of noise-spreading samples, the inverse diffusion network can control the consistency of global information related to the target speech corresponding to the noise-spreading samples based on vector sequences and global semantic tokens. Global information includes, for example, at least one of timbre, pitch, and emotional tone, but is not limited to these, and is not restricted here. Correspondingly, the inverse diffusion network can control the instantaneous spectral anomaly repair of local information related to the target speech corresponding to the noise-spreading samples based on vector sequences and local tokens. Local information includes, for example, at least one of specific pronunciations in each frame, instantaneous pitch, energy, and plosive distortion, but is not limited to these, and is not restricted here.

[0098] In one exemplary implementation, the diffusion model embeds a single-stream speech codec during the noise reduction stage. For example, the inverse diffusion network of the diffusion model embeds a single-stream speech codec. The single-stream speech codec can be used to provide at least one of global semantic tokens and local tokens for the inverse diffusion network to cooperate with vector sequences to predict noise in the noise-spreading samples, thereby determining the target speech corresponding to the noise-spreading samples. The single-stream speech codec includes, for example, a BiCodec encoder, but is not limited thereto.

[0099] Correspondingly, the reverse diffusion network can also predict noise in the updated noise diffusion sample based on the vector sequence and at least one of the global semantic token and local token, so as to obtain the target speech corresponding to the noise diffusion sample.

[0100] In cases where the inverse diffusion network can predict noise from noise-spreading samples based on vector sequences and at least one of global semantic tokens and local tokens, and obtain the target speech corresponding to the noise-spreading samples, the target speech can have the characteristics of global consistency and / or no instantaneous spectral anomalies, which is beneficial to improving the speech synthesis effect.

[0101] As can be seen in the above embodiments, for entities under different types of businesses, such as insurance business, financial business, medical and health business, etc., text information to be processed can be obtained; the text encoding sub-model based on the diffusion model encodes the text information to obtain the vector sequence corresponding to the text information; the acoustic feature extraction sub-model based on the diffusion model extracts acoustic features from the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence; the context-aware sub-model based on the diffusion model performs text-speech alignment processing on the first Mel spectrum according to the vector sequence to obtain the corresponding second Mel spectrum; the high-frequency compensation diffusion sub-model based on the diffusion model performs residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum to obtain the corresponding target Mel spectrum; the diffusion sub-model based on the diffusion model determines the target speech corresponding to the text information according to the target Mel spectrum.

[0102] Given the text information to be processed, a diffusion model can be used to encode the text information, extract acoustic features from the corresponding vector sequence, perform text-to-speech alignment on the first Mel spectrum, perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum, and then determine the target speech corresponding to the text information based on the target Mel spectrum. In the process of determining the target speech using the diffusion model, the context-aware sub-model constrains the first Mel spectrum for text-to-speech alignment, which can improve the prosodic balance of the second Mel spectrum. This, in turn, can improve the prosodic balance of the target speech determined by the diffusion model, effectively enhancing the speech quality of the target speech and thus improving the speech synthesis effect. This improved speech synthesis effect is beneficial for promoting the development of financial and healthcare businesses. For example, for text information involved in financial transactions, such as payment information, sales information, leasing information, insurance policy information, pension information, etc., this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to more flexibly inform users of the text information involved in financial transactions, thereby promoting the development of financial businesses. For text information involved in healthcare services, such as health monitoring information, consultation results, health records, examination reports, etc., this speech synthesis method can be used for TTS processing, and then the text information can be read aloud to users in a more flexible way, thereby promoting the development of healthcare services.

[0103] The speech synthesis method provided in this application can use a diffusion model to determine the target speech corresponding to the text information. In other words, the speech synthesis method can be implemented with the help of artificial intelligence (AI), which helps to reduce labor costs and improve the convenience of speech synthesis.

[0104] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0105] In one embodiment, a method for training a diffusion model is provided. The diffusion model trained by this method can be used in the aforementioned speech synthesis method. It should be noted that the diffusion model training method provided in this embodiment can be used on a computer device, and of course, on a server. For example, the server can obtain a training sample set from the computer device, process the training sample set according to the diffusion model training method, and obtain model optimization parameters for the diffusion model. Based on these optimization parameters, the model parameters of the diffusion model are adjusted to obtain an optimized diffusion model. For example, the server can send the optimized diffusion model to the computer device so that when the computer device obtains text information to be processed, it can determine the target speech corresponding to the text information using the optimized diffusion model; however, this is not limited to this method and is not restricted here.

[0106] In practice, computer devices include, but are not limited to, any of the following: mobile phones, tablets, laptops, and desktop computers; servers can be standalone servers, server clusters, or cloud servers that provide cloud computing services.

[0107] Please see Figure 3 As shown, Figure 3 A flowchart illustrating a training method for a diffusion model provided in this application embodiment includes the following steps:

[0108] S201: Obtain the training sample set, which includes multiple text information to be processed and the corresponding speech tags for each text information.

[0109] For example, when training a diffusion model, a training sample set can be obtained. The training sample set can include multiple text messages to be processed. Each text message in the training sample set can be assigned a corresponding speech tag. The training sample set can be used to subsequently train the diffusion model, enabling the diffusion model to learn how to determine the target speech corresponding to the text message by combining the text messages to be processed and their corresponding speech tags.

[0110] The training sample set includes text information covering various business scenarios, such as financial and healthcare. For example, in the financial business scenario, the text information could include payment information, sales information, leasing information, insurance policy information, pension information, and so on. In the healthcare business scenario, the text information could include health monitoring information, consultation results, health records, examination reports, and so on.

[0111] Of course, text information is not limited to this, and no restrictions are imposed here.

[0112] Thus, given the text information included in the training sample set and the corresponding speech tags, the training sample set can be used to train the diffusion model.

[0113] S202: A text encoding sub-model based on the diffusion model, which encodes text information to obtain a vector sequence corresponding to the text information.

[0114] For example, during training, a diffusion model needs to learn how to determine the vector sequence corresponding to text information. Based on this, text information can be input into the text encoding sub-model of the diffusion model, so that the text encoding sub-model can determine the vector sequence corresponding to the text information.

[0115] In this embodiment, the text encoding sub-model based on the diffusion model encodes the text information and obtains the vector sequence corresponding to the text information. The relevant description can be referred to the corresponding description in the previous embodiment, and will not be repeated here.

[0116] Thus, given a known vector sequence corresponding to the text information, the diffusion model can determine the target speech corresponding to the text information based on the vector sequence, and then optimize the model parameters of the diffusion model by combining the target speech and the speech tags corresponding to the text information, thereby improving the training effect of the diffusion model.

[0117] S203: Acoustic feature extraction sub-model based on diffusion model, which performs acoustic feature extraction processing on vector sequence to obtain the first Mel spectrum corresponding to the vector sequence.

[0118] For example, during training, the diffusion model needs to learn how to determine the first Mel spectrum corresponding to the vector sequence, i.e., the first Mel spectrum corresponding to the text information. Based on this, the vector sequence can be input into the acoustic feature extraction sub-model of the diffusion model, so that the acoustic feature extraction sub-model can determine the first Mel spectrum corresponding to the vector sequence.

[0119] In this embodiment, the acoustic feature extraction sub-model based on the diffusion model performs acoustic feature extraction processing on the vector sequence to obtain the relevant description of the first Mel spectrum corresponding to the vector sequence. The description can be referred to the corresponding description in the previous embodiment, and will not be repeated here.

[0120] Thus, given the first Mel spectrum corresponding to the vector sequence, the diffusion model can determine the target speech corresponding to the text information based on the first Mel spectrum corresponding to the text information. Then, by combining the target speech and the speech tags corresponding to the text information, the model parameters of the diffusion model can be optimized to improve the training effect of the diffusion model.

[0121] S204: A context-aware sub-model based on a diffusion model, which performs text-speech alignment processing on the first Mel spectrum according to the vector sequence to obtain the corresponding second Mel spectrum.

[0122] For example, during training, the diffusion model needs to learn how to determine the second Mel spectrum corresponding to the first Mel spectrum, i.e., the second Mel spectrum corresponding to the text information. Based on this, the vector sequence and the first Mel spectrum can be input into the context-aware sub-model of the diffusion model, so that the context-aware sub-model can constrain the first Mel spectrum to perform text-speech alignment based on the context indicated by the vector sequence, thereby determining the second Mel spectrum corresponding to the text information.

[0123] In this embodiment, the context-aware sub-model based on the diffusion model performs text-speech alignment processing on the first Mel spectrum according to the vector sequence to obtain the corresponding second Mel spectrum. The relevant description can be referred to the corresponding description in the previous embodiment, and will not be repeated here.

[0124] Thus, given the second Mel spectrum corresponding to the text information, the diffusion model can determine the target speech corresponding to the text information based on the second Mel spectrum, and then optimize the model parameters of the diffusion model by combining the target speech and the speech label corresponding to the text information, so as to improve the training effect of the diffusion model.

[0125] S205: A high-frequency compensated diffusion sub-model based on a diffusion model is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0126] For example, during training, the diffusion model needs to learn how to determine the target Mel spectrum corresponding to the second Mel spectrum, i.e., the target Mel spectrum corresponding to the text information. Based on this, the second Mel spectrum can be input into the high-frequency compensation diffusion sub-model of the diffusion model, so that the high-frequency compensation diffusion sub-model can determine the second Mel spectrum corresponding to the text information.

[0127] In this embodiment, a high-frequency compensated diffusion sub-model based on the diffusion model is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum to obtain the corresponding description of the target Mel spectrum. The relevant description can be referred to the corresponding description in the previous embodiment, and will not be repeated here.

[0128] Thus, given the target Mel spectrum corresponding to the text information, the diffusion model can determine the target speech corresponding to the text information based on the target Mel spectrum, and then optimize the model parameters of the diffusion model by combining the target speech and the speech label corresponding to the text information, so as to improve the training effect of the diffusion model.

[0129] S206: A diffusion sub-model based on the diffusion model, which determines the target speech corresponding to the text information based on the target Mel spectrum.

[0130] For example, during training, the diffusion model needs to learn how to determine the target speech corresponding to the target Mel spectrum, i.e., the target speech corresponding to the text information. Based on this, the target Mel spectrum can be input into the diffusion sub-model of the diffusion model, so that the diffusion sub-model can determine the target speech corresponding to the text information.

[0131] In this embodiment, the diffusion sub-model based on the diffusion model determines the relevant description of the target speech corresponding to the text information according to the target Mel spectrum. The description can be referred to the corresponding description in the previous embodiment, and will not be repeated here.

[0132] Thus, given the target speech corresponding to the text information, the diffusion model can integrate the target speech corresponding to the text information and the speech label corresponding to the text information to optimize the model parameters, thereby improving the training effect of the diffusion model.

[0133] S207: Based on the discriminator of the diffusion model, the model optimization parameters of the diffusion model are determined according to the target speech and speech tags.

[0134] For example, the diffusion model's framework can employ an Adversarial Diffusion Training Framework for Text-to-Speech (ADM-TTS). For instance, the diffusion model embeds a discriminator during the diffusion denoising process. The discriminator can include at least two of a spectrum discriminator, a waveform discriminator, and a fundamental frequency discriminator. The diffusion model can dynamically adjust the discriminative weights of at least two of the spectrum discriminator, waveform discriminator, and fundamental frequency discriminator through a meta-learning strategy. Given a target speech corresponding to the text information, the diffusion model can use the discriminator to evaluate the difference between the target speech and the speech label. For example, the discriminator can compare the predicted Mel spectrum corresponding to the target speech with the preset Mel spectrum corresponding to the speech label to determine whether there is a difference in at least two of the spectrum, waveform, and fundamental frequency, thereby determining whether there is a difference between the target speech and the speech label. The predicted Mel spectrum corresponding to the target speech can be determined during the diffusion model's noise prediction process on the noise diffusion sample corresponding to the target Mel spectrum. The preset Mel spectrum corresponding to the speech label can be predetermined. The predicted Mel spectrum can be used to determine whether it is the diffusion model's prediction result for the target speech corresponding to the text information. The preset Mel spectrum can be used to determine the actual result of the diffusion model on the target speech corresponding to the text information. The discriminator can determine whether the target speech is increasingly approaching the speech label based on the predicted Mel spectrum and the preset Mel spectrum. For example, the discriminator can determine the loss value between the target speech and the speech label.

[0135] The diffusion model can determine its optimization parameters based on the loss value. These optimization parameters can be used to adjust the model parameters later, allowing the diffusion model to more accurately predict the target speech corresponding to the text information, thus achieving the training objective.

[0136] S208: Adjust the model parameters of the diffusion model according to the model optimization parameters to obtain the optimized diffusion model.

[0137] For example, the diffusion model can adjust the model parameters of at least one of its sub-models, such as the text encoding sub-model, the acoustic feature extraction sub-model, the context-aware sub-model, the high-frequency compensation diffusion sub-model, and the diffusion sub-model itself, based on the model optimization parameters. However, it is not limited to this; the diffusion model can also adjust the model parameters of the networks, encoders, etc., included in each sub-model. No restrictions are placed here.

[0138] In one exemplary implementation, adversarial training can be used to drive stability optimization of the diffusion model during training and optimization. For example, the diffusion model can incorporate ADM-TTS, fusing progressive distillation with GAN constraints. Multiple discriminators (spectral discriminator, waveform discriminator, and fundamental frequency discriminator) are embedded during the diffusion denoising process, and discriminator weights are dynamically adjusted through a meta-learning strategy to suppress high-frequency artifacts and phase distortion. Furthermore, spectral normalization constraints are applied to the diffusion model to limit the risk of gradient explosion in the diffusion step size, thereby improving the training stability of the diffusion model. By using an adversarial training framework to improve the training convergence speed of the diffusion model by 35%, support continuous adjustment of pitch ±20% and speech rate 0.5-2.0 times, and achieve an emotion control accuracy exceeding 96%, the training stability and controllability of the diffusion model are further enhanced.

[0139] By adjusting the model parameters of the diffusion model, the optimized model can more accurately predict the target speech corresponding to the text information, thus achieving the training objective of the diffusion model. Correspondingly, when the optimized diffusion model's prediction of the target speech corresponding to the text information is closer to the actual result, the optimized diffusion model can more accurately convert the text information into target speech, which helps improve the accuracy of the diffusion model in determining the target speech corresponding to the text information, and consequently improves the speech synthesis effect.

[0140] In some implementations, after obtaining the optimized diffusion model, the process further includes: distilling the acoustic feature extraction sub-model and the context-aware sub-model in the optimized diffusion model according to the first step compression ratio to obtain the processed acoustic feature extraction sub-model and the processed context-aware sub-model; distilling the high-frequency compensation diffusion sub-model in the optimized diffusion model according to the second step compression ratio to obtain the processed high-frequency compensation diffusion sub-model; wherein the first step compression ratio is less than or equal to the second step compression ratio; and updating the optimized diffusion model according to the processed acoustic feature extraction sub-model, the processed context-aware sub-model, and the processed high-frequency compensation diffusion sub-model to obtain the updated diffusion model.

[0141] For example, in the process of training a diffusion model and obtaining an optimized diffusion model, the acoustic feature extraction sub-model, context-aware sub-model, and high-frequency compensation diffusion sub-model in the diffusion model can serve as a two-stage conditional diffusion model architecture to train how the diffusion model determines the target Mel spectrum corresponding to the text information. Exemplarily, the acoustic feature extraction sub-model and the context-aware sub-model can be used for the first stage of conditional diffusion, such as semantically guided diffusion. The high-frequency compensation diffusion sub-model can be used for the second stage of conditional diffusion, such as detail-enhancing diffusion. First stage: The acoustic feature extraction sub-model determines the first Mel spectrum corresponding to the text information, and the context-aware sub-model constrains text-speech alignment to improve prosodic balance, resulting in the second Mel spectrum. Second stage: Using the high-frequency compensation diffusion sub-network, conditioned on the second Mel spectrum, residual learning is used to focus on repairing high-frequency details, and a hybrid encoder is combined to extract multi-scale acoustic features, thereby determining the target Mel spectrum. Through two-stage diffusion and residual learning, the harmonic signal-to-noise ratio (HSNR) in the 2-8kHz frequency band is improved by 6.2dB, and the fricative sound clarity reaches human level (MOS≥4.5), which is beneficial to improving the high-frequency detail fidelity of the diffusion model when performing TTS processing on text information.

[0142] For the optimized diffusion model, a hierarchical distillation strategy can be adopted to distill the acoustic feature extraction sub-model, context-aware sub-model, and high-frequency compensation diffusion sub-model within the diffusion model. The first-step compression ratio is, for example, 4:1. The second-step compression ratio is, for example, 8:1. Of course, the first-step and second-step compression ratios are not limited to these; they are not restricted here. Through hierarchical distillation and dynamic scheduling, the diffusion model achieves a TTS inference speed of 0.18 seconds per sentence (within 100 characters), which is 5.7 times faster than the traditional diffusion model (TorToiSe), meeting the requirements of real-time interactive scenarios such as in-vehicle and AR applications, thus improving the real-time speech synthesis capabilities of the diffusion model.

[0143] By distilling the acoustic feature extraction sub-model, context-aware sub-model, and high-frequency compensation diffusion sub-model in the diffusion model, we obtain the processed acoustic feature extraction sub-model, processed context-aware sub-model, and processed high-frequency compensation diffusion sub-model. This process is then used to update and optimize the diffusion model, resulting in an updated diffusion model. The updated diffusion model is more efficient at determining the target speech corresponding to the text information than the optimized diffusion model, which helps improve the speech synthesis efficiency of the diffusion model and thus reduces the inference latency when performing TTS processing on text information.

[0144] Of course, this is not the only approach. During the distillation process of the acoustic feature extraction sub-model, context-aware sub-model, and high-frequency compensation diffusion sub-model in the diffusion model, knowledge transfer processing can be performed on the diffusion model to improve the inference speed of the target speech corresponding to the text information while maintaining the speech quality, resulting in an updated diffusion model. Correspondingly, the updated diffusion model can be deployed to the client. For example, FP16 quantization and dynamic axis sparsity can be used during the deployment phase to reduce the memory footprint of the diffusion model. By performing knowledge transfer on the diffusion model, even with 10 minutes of low-quality data, the naturalness of speech (MOS) of the target speech obtained by the diffusion model after TTS processing of the text information remains above 4.1, thus improving the robustness of the diffusion model in low-resource scenarios.

[0145] The training method for the diffusion model provided in the above embodiments involves: acquiring a training sample set, which includes multiple text information to be processed and corresponding speech tags for each text information; using a text encoding sub-model based on the diffusion model to encode the text information to obtain a vector sequence corresponding to the text information; using an acoustic feature extraction sub-model based on the diffusion model to extract acoustic features from the vector sequence to obtain a first Mel spectrum corresponding to the vector sequence; using a context-aware sub-model based on the diffusion model to perform text-speech alignment processing on the first Mel spectrum according to the vector sequence to obtain a corresponding second Mel spectrum; using a high-frequency compensation diffusion sub-model based on the diffusion model to perform residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum to obtain a corresponding target Mel spectrum; using a diffusion sub-model based on the diffusion model to determine the target speech corresponding to the text information based on the target Mel spectrum; using a discriminator based on the diffusion model to determine the model optimization parameters of the diffusion model according to the target speech and speech tags; and adjusting the model parameters of the diffusion model according to the model optimization parameters to obtain an optimized diffusion model, thereby improving the training effect of the diffusion model. Subsequently, the diffusion model, such as the optimized diffusion model or the updated diffusion model, can be applied to the client, which will help improve the convenience and accuracy of determining the target speech corresponding to the text information obtained by the client, and thus help improve the speech synthesis effect.

[0146] In one embodiment, a speech synthesis device is provided, which corresponds one-to-one with the speech synthesis methods described in the above embodiments. For example... Figure 4 As shown, the speech synthesis device includes a text acquisition module 110, an encoding module 120, a feature extraction module 130, a text-to-speech alignment module 140, a high-frequency compensation module 150, and a speech synthesis module 160. Detailed descriptions of each functional module are as follows:

[0147] Text acquisition module 110 is used to acquire text information to be processed;

[0148] The encoding module 120 is used to encode the text information based on the diffusion model sub-model to obtain the vector sequence corresponding to the text information;

[0149] Feature extraction module 130 is used to perform acoustic feature extraction processing on the vector sequence based on the acoustic feature extraction sub-model of the diffusion model to obtain the first Mel spectrum corresponding to the vector sequence;

[0150] The text-speech alignment module 140 is used to perform text-speech alignment processing on the first Mel spectrum based on the context-aware sub-model of the diffusion model and the vector sequence to obtain the corresponding second Mel spectrum.

[0151] The high-frequency compensation module 150 is used to perform residual learning and multi-scale acoustic feature extraction processing on the second Mel spectrum based on the high-frequency compensation diffusion sub-model of the diffusion model to obtain the corresponding target Mel spectrum.

[0152] The speech synthesis module 160 is used to determine the target speech corresponding to the text information based on the diffusion sub-model of the diffusion model and the target Mel spectrum.

[0153] In one embodiment, the encoding module 120 is configured to:

[0154] Based on the text encoding sub-model, the text semantic encoding network performs semantic feature encoding processing on the text information to obtain the semantic vector corresponding to the text information;

[0155] Based on the speech feature coding network of the text coding sub-model, the text information is processed by speech feature coding to obtain the speech feature vector corresponding to the text information.

[0156] Based on the vector fusion network of the text encoding sub-model, cross-modal fusion processing is performed on the semantic vector and the speech feature vector to obtain the vector sequence corresponding to the text information.

[0157] In one embodiment, the speech synthesis module 160 is used for:

[0158] Based on the forward diffusion network of the diffusion sub-model, noise diffusion is performed on the target Mel spectrum according to the text complexity corresponding to the text information to obtain the noise diffusion sample corresponding to the target Mel spectrum;

[0159] Based on the inverse diffusion network of the diffusion sub-model, noise prediction is performed on the noise diffusion sample according to the vector sequence to obtain the target speech corresponding to the noise diffusion sample.

[0160] In one embodiment, the speech synthesis device is used for:

[0161] Based on the lip movement synchronization coding network of the diffusion sub-model, the target lip shape parameters corresponding to the target Mel spectrum are obtained, and the noise diffusion sample is updated according to the target lip shape parameters to obtain the updated noise diffusion sample;

[0162] The inverse diffusion network based on the diffusion sub-model performs noise prediction on the noise diffusion samples according to the vector sequence to obtain the target speech corresponding to the noise diffusion samples, including:

[0163] Based on the inverse diffusion network of the diffusion sub-model, noise prediction is performed on the updated noise diffusion sample according to the vector sequence to obtain the target speech corresponding to the noise diffusion sample.

[0164] In one embodiment, the speech synthesis device is used for:

[0165] Based on the reverse diffusion network, noise prediction is performed on the noise diffusion samples according to the vector sequence and at least one of the global semantic token and local token, to obtain the corresponding target speech.

[0166] Based on the speech synthesis apparatus provided in this application, computer devices can determine the target speech corresponding to acquired text information, thereby improving the convenience and accuracy of TTS processing of text information and ultimately enhancing the speech synthesis effect. This improved speech synthesis effect is beneficial for promoting the development of financial and healthcare businesses.

[0167] For specific limitations regarding the speech synthesis device, please refer to the limitations on the speech synthesis method above, which will not be repeated here. Each module in the aforementioned speech synthesis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.

[0168] In another embodiment, a training apparatus for a diffusion model is provided, which corresponds one-to-one with the training methods for the diffusion model in the above embodiments. The training apparatus includes a training sample acquisition module, an encoding training module, a feature extraction training module, a text-to-speech alignment training module, a high-frequency compensation training module, a speech synthesis module, an optimization parameter determination module, and a model optimization module. Detailed descriptions of each functional module are as follows:

[0169] The training sample acquisition module is used to acquire a training sample set, which includes multiple text information to be processed and speech tags corresponding to each text information.

[0170] The encoding training module is used to encode the text information based on the diffusion model's text encoding sub-model to obtain the vector sequence corresponding to the text information;

[0171] The feature extraction training module is used to perform acoustic feature extraction processing on the vector sequence based on the acoustic feature extraction sub-model of the diffusion model to obtain the first Mel spectrum corresponding to the vector sequence;

[0172] The text-speech alignment training module is used to perform text-speech alignment processing on the first Mel spectrum based on the context-aware sub-model of the diffusion model and the vector sequence to obtain the corresponding second Mel spectrum.

[0173] The high-frequency compensation training module is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum based on the high-frequency compensation diffusion sub-model of the diffusion model to obtain the corresponding target Mel spectrum.

[0174] The speech synthesis module is used to determine the target speech corresponding to the text information based on the diffusion sub-model of the diffusion model and the target Mel spectrum.

[0175] The optimization parameter determination module is used to determine the model optimization parameters of the diffusion model based on the discriminator of the diffusion model, according to the target speech and speech tags.

[0176] The model optimization module is used to adjust the model parameters of the diffusion model according to the model optimization parameters to obtain the optimized diffusion model.

[0177] In one embodiment, the training device is used for:

[0178] Based on the compression ratio in the first step, the acoustic feature extraction sub-model and the context-aware sub-model in the optimized diffusion model are distilled to obtain the processed acoustic feature extraction sub-model and the processed context-aware model.

[0179] According to the second step compression ratio, the high-frequency compensation diffusion sub-model in the optimized diffusion model is distilled to obtain the processed high-frequency compensation diffusion sub-model; the first step compression ratio is less than or equal to the second step compression ratio.

[0180] Based on the processed acoustic feature extraction sub-model, the processed context-aware sub-model, and the processed high-frequency compensation diffusion sub-model, the optimized diffusion model is updated to obtain the updated diffusion model.

[0181] Based on the diffusion model training device provided in this application, computer equipment can determine the optimized or updated diffusion model through the diffusion model training device, and then perform TTS processing on text information to determine the target speech corresponding to the text information, thereby improving the convenience and accuracy of TTS processing of text information and thus improving the speech synthesis effect. The improved speech synthesis effect is beneficial to promoting the development of financial and healthcare businesses.

[0182] Specific limitations regarding the training apparatus for the diffusion model can be found in the limitations on the training method for the diffusion model described above, and will not be repeated here. Each module in the aforementioned training apparatus for the diffusion model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0183] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a speech synthesis method server-side function, or the functions or steps of a diffusion model training method server-side function.

[0184] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a speech synthesis method, or client-side functions or steps of a diffusion model training method.

[0185] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0186] Obtain the text information to be processed;

[0187] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0188] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0189] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0190] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0191] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum.

[0192] In another embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0193] Obtain a training sample set, which includes multiple text information to be processed and speech tags corresponding to each text information;

[0194] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0195] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0196] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0197] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0198] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum;

[0199] Based on the discriminator of the diffusion model, the model optimization parameters of the diffusion model are determined according to the target speech and speech tags;

[0200] Based on the model optimization parameters, the model parameters of the diffusion model are adjusted to obtain the optimized diffusion model.

[0201] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0202] Text information to be processed;

[0203] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0204] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0205] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0206] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0207] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum.

[0208] In another embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0209] Obtain a training sample set, which includes multiple text information to be processed and speech tags corresponding to each text information;

[0210] A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information.

[0211] Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence;

[0212] Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum;

[0213] Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum.

[0214] Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum;

[0215] Based on the discriminator of the diffusion model, the model optimization parameters of the diffusion model are determined according to the target speech and speech tags;

[0216] Based on the model optimization parameters, the model parameters of the diffusion model are adjusted to obtain the optimized diffusion model.

[0217] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0218] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0219] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0220] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A speech synthesis method, characterized in that, The speech synthesis method includes: Obtain the text information to be processed; A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information. Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence; Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum; Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum. Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum.

2. The speech synthesis method according to claim 1, characterized in that, The text encoding sub-model based on the diffusion model encodes the text information to obtain a vector sequence corresponding to the text information, including: Based on the text encoding sub-model, the text semantic encoding network performs semantic feature encoding processing on the text information to obtain the semantic vector corresponding to the text information; Based on the speech feature coding network of the text coding sub-model, the text information is processed by speech feature coding to obtain the speech feature vector corresponding to the text information. Based on the vector fusion network of the text encoding sub-model, cross-modal fusion processing is performed on the semantic vector and the speech feature vector to obtain the vector sequence corresponding to the text information.

3. The speech synthesis method according to any one of claims 1 to 2, characterized in that, The diffusion sub-model based on the diffusion model determines the target speech corresponding to the text information according to the target Mel spectrum, including: Based on the forward diffusion network of the diffusion sub-model, noise diffusion is performed on the target Mel spectrum according to the text complexity corresponding to the text information to obtain the noise diffusion sample corresponding to the target Mel spectrum; Based on the inverse diffusion network of the diffusion sub-model, noise prediction is performed on the noise diffusion sample according to the vector sequence to obtain the target speech corresponding to the noise diffusion sample.

4. The speech synthesis method according to claim 3, characterized in that, After obtaining the noise diffusion sample corresponding to the target Mel spectrum, the method further includes: Based on the lip movement synchronization coding network of the diffusion sub-model, the target lip shape parameters corresponding to the target Mel spectrum are obtained, and the noise diffusion sample is updated according to the target lip shape parameters to obtain the updated noise diffusion sample; The inverse diffusion network based on the diffusion sub-model performs noise prediction on the noise diffusion samples according to the vector sequence to obtain the target speech corresponding to the noise diffusion samples, including: Based on the inverse diffusion network of the diffusion sub-model, noise prediction is performed on the updated noise diffusion sample according to the vector sequence to obtain the target speech corresponding to the noise diffusion sample.

5. The speech synthesis method according to claim 3, characterized in that, The inverse diffusion network based on the diffusion sub-model performs noise prediction on the noise diffusion samples according to the vector sequence to obtain the target speech corresponding to the noise diffusion samples, including: Based on the reverse diffusion network, noise prediction is performed on the noise diffusion samples according to the vector sequence and at least one of the global semantic token and local token, to obtain the corresponding target speech.

6. A training method for a diffusion model, characterized in that, The training method includes: Obtain a training sample set, which includes multiple text information to be processed and speech tags corresponding to each text information; A text encoding sub-model based on a diffusion model is used to encode the text information to obtain a vector sequence corresponding to the text information. Based on the acoustic feature extraction sub-model of the diffusion model, acoustic feature extraction processing is performed on the vector sequence to obtain the first Mel spectrum corresponding to the vector sequence; Based on the context-aware sub-model of the diffusion model, the first Mel spectrum is processed by text-speech alignment according to the vector sequence to obtain the corresponding second Mel spectrum; Based on the high-frequency compensated diffusion sub-model of the diffusion model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain the corresponding target Mel spectrum. Based on the diffusion sub-model of the diffusion model, the target speech corresponding to the text information is determined according to the target Mel spectrum; Based on the discriminator of the diffusion model, the model optimization parameters of the diffusion model are determined according to the target speech and speech tags; Based on the model optimization parameters, the model parameters of the diffusion model are adjusted to obtain the optimized diffusion model.

7. The training method according to claim 6, characterized in that, Following the obtained optimized diffusion model, the following is also included: Based on the compression ratio in the first step, the acoustic feature extraction sub-model and the context-aware sub-model in the optimized diffusion model are distilled to obtain the processed acoustic feature extraction sub-model and the processed context-aware model. According to the second step compression ratio, the high-frequency compensation diffusion sub-model in the optimized diffusion model is distilled to obtain the processed high-frequency compensation diffusion sub-model; the first step compression ratio is less than or equal to the second step compression ratio. Based on the processed acoustic feature extraction sub-model, the processed context-aware sub-model, and the processed high-frequency compensation diffusion sub-model, the optimized diffusion model is updated to obtain the updated diffusion model.

8. A speech synthesis device, characterized in that, The speech synthesis device includes: The text acquisition module is used to acquire the text information to be processed. The encoding module is used to encode the text information based on the diffusion model sub-model to obtain the vector sequence corresponding to the text information; The feature extraction module is used to perform acoustic feature extraction processing on the vector sequence based on the acoustic feature extraction sub-model of the diffusion model to obtain the first Mel spectrum corresponding to the vector sequence; The text-speech alignment module is used to perform text-speech alignment processing on the first Mel spectrum based on the context-aware sub-model of the diffusion model and the vector sequence to obtain the corresponding second Mel spectrum. The high-frequency compensation module is used to perform residual learning and multi-scale acoustic feature extraction on the second Mel spectrum based on the high-frequency compensation diffusion sub-model of the diffusion model to obtain the corresponding target Mel spectrum. The speech synthesis module is used to determine the target speech corresponding to the text information based on the diffusion sub-model of the diffusion model and the target Mel spectrum.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the speech synthesis method as described in any one of claims 1 to 5, or implement the training method of the diffusion model as described in any one of claims 6 to 7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the speech synthesis method as described in any one of claims 1 to 5, or the steps of the training method of the diffusion model as described in any one of claims 6 to 7.

Citation Information

Cited By

  • Method and device for improving speech synthesis speed of Cauchy denoising diffusion probability model

    CN121999754A