TTS generation method and device based on potential diffusion model, equipment and medium
Through the TTS generation method based on the latent diffusion model, the speech is generated using text encoder, latent spatial mapping and denoising subnetwork, the problems of low generation efficiency and poor speech quality in the prior art are solved, and efficient and natural speech synthesis is achieved.
Patent Information
- Application Number
- CN202510670789.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-08
AI Technical Summary
The existing TTS synthesis technology has problems such as low generation efficiency, poor speech quality and poor scalability, especially in low resource environments, which are insufficient stability and speech naturalness.
Using a method based on the potential diffusion model, the context feature vector is generated through the text encoder, and the target denoising subnet of the latent space mapping module is input to the latent diffusion model for denoising processing, and the target voice is generated by combining the waveform generator.
It improves the efficiency of speech generation, the generated speech is more natural and smooth, reduces noise and distortion, and can flexibly adapt to the generation needs of different languages and contexts.
Smart Images

Figure CN120452414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence speech processing and medical technology, and in particular to a TTS generation method, device, equipment and medium based on a potential diffusion model. Background Art
[0002] With the rapid development of artificial intelligence technology, text-to-speech (TTS) synthesis technology, as the core link of human-computer interaction, has shown broad application prospects in many fields such as medical and health reminders, news broadcasting, virtual digital humans, and financial customer service.
[0003] However, despite significant progress in TTS synthesis technology in recent years, significant shortcomings remain. Currently, existing technologies include generation methods based on end-to-end models, streaming generation, and diffusion models. Among them, end-to-end neural network TTS systems (such as Tacotron 2) can directly generate Mel spectrograms from text and obtain speech waveforms through a vocoder, but they rely on large data sets and high computing power, and have poor stability in complex contexts or low-resource environments. Diffusion models (such as DDPM) perform well in the field of image generation, but suffer from slow generation and high computational overhead when used for speech generation. Generative adversarial networks (such as WaveGAN) rely on adversarial training between generators and discriminators to generate speech, but training is unstable, making it difficult to capture long-term dependencies, and the speech is easily distorted and unnatural. Therefore, existing TTS synthesis suffers from low generation efficiency, susceptibility to noise and distortion, poor model scalability, and potential degradation of speech quality. Summary of the Invention
[0004] The embodiments of the present invention provide a TTS generation method, apparatus, device and medium based on a latent diffusion model, aiming to solve the problems of low speech generation efficiency, poor speech quality and poor scalability in the prior art.
[0005] In a first aspect, an embodiment of the present invention provides a method for generating a TTS based on a latent diffusion model, the method comprising:
[0006] Obtaining a target text sequence, inputting the target text sequence into a text encoder for encoding, and generating a context feature vector corresponding to the target text sequence;
[0007] Performing dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector;
[0008] Inputting the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, thereby generating speech features corresponding to the first low-dimensional feature vector;
[0009] The speech features are converted using a waveform generator to generate target speech corresponding to the speech features.
[0010] In a second aspect, an embodiment of the present invention further provides a TTS generation device based on a potential diffusion model, the device comprising:
[0011] A text sequence encoding unit, configured to obtain a target text sequence, input the target text sequence into a text encoder for encoding, and generate a context feature vector corresponding to the target text sequence;
[0012] a latent space mapping unit, configured to perform dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector;
[0013] a diffusion denoising unit, configured to input the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, and generate speech features corresponding to the first low-dimensional feature vector;
[0014] The target speech generating unit is used to convert the speech features using a waveform generator to generate a target speech corresponding to the speech features.
[0015] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the TTS generation method based on the potential diffusion model as described in the first aspect is implemented.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the TTS generation method based on the potential diffusion model described in the first aspect.
[0017] The embodiment of the present invention provides a TTS generation method, device, equipment and medium based on a latent diffusion model, which obtains a target text sequence, inputs the target text sequence into a text encoder for encoding, and generates a context feature vector; performs dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector; inputs the first low-dimensional feature vector into a target denoising subnetwork in the latent diffusion model for denoising processing to generate speech features; uses a waveform generator to convert the speech features to generate a target speech, thereby reducing the amount of high-dimensional space calculations and improving speech generation efficiency by performing dimensionality reduction processing on the context feature vector in the latent space; and makes the generated speech more natural and smoother through the denoising characteristics of the latent extension model, reduces noise and distortion, improves speech generation quality, and can flexibly adapt to the generation requirements of different languages and contexts. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A schematic diagram of a flow chart of a TTS generation method based on a potential diffusion model provided by an embodiment of the present invention;
[0020] Figure 2 A schematic block diagram of a TTS generation device based on a potential diffusion model provided by an embodiment of the present invention;
[0021] Figure 3 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0024] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0025] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0026] The embodiment of the present invention provides a TTS generation method based on a latent diffusion model that is specifically applicable to various application scenarios involving text to speech, such as news broadcasting, virtual digital people, financial customer service, medical health reminders, etc., and the TTS generation method based on the latent diffusion model can be applied to a server. The server obtains a target text sequence, inputs the target text sequence into a text encoder for encoding, and generates a context feature vector corresponding to the target text sequence; the context feature vector is subjected to dimensionality reduction processing through a latent space mapping module to obtain a first low-dimensional feature vector; the first low-dimensional feature vector is input into a target denoising subnetwork in the latent diffusion model for denoising processing to generate speech features corresponding to the first low-dimensional feature vector; the speech features are converted using a waveform generator to generate a target speech corresponding to the speech features. The server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0027] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0028] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0029] The present invention is described in detail below through specific examples.
[0030] See also Figure 1 , Figure 1 FIG is a flow chart of a TTS generation method based on a potential diffusion model provided by an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110-S140.
[0031] S110 , obtaining a target text sequence, inputting the target text sequence into a text encoder for encoding, and generating a context feature vector corresponding to the target text sequence.
[0032] In this embodiment, in order to convert the target text sequence into a numerical representation that can be understood by the server, after obtaining the target text sequence, this embodiment inputs the target text sequence into a text encoder for encoding, so as to effectively capture the context information in the target text sequence to generate a context feature vector corresponding to the target text sequence; wherein the context feature vector includes semantic information such as time, task, action, etc.; and the text encoder is preferably a Transformer encoder.
[0033] For example, if the target text information is “Remind me to measure my blood pressure at 8 o’clock tomorrow morning”, the context feature vector contains semantic information such as “remind”, “8 o’clock”, and “measure blood pressure”.
[0034] In one embodiment, inputting the target text sequence into a text encoder for encoding to generate a context feature vector corresponding to the target text sequence includes:
[0035] Convert the target text sequence according to a preset word embedding table to obtain an embedding vector corresponding to the target text sequence;
[0036] fusing the embedding vectors through a cross-attention module of the text encoder to obtain first fused information;
[0037] The first fusion information is subjected to feature extraction using multiple multi-head attention modules in the text encoder to obtain the context feature vector.
[0038] In this embodiment, the target text sequence is first converted through a preset word embedding table to obtain an embedding vector corresponding to the target text sequence, so as to capture the semantic information of the target text sequence and facilitate the subsequent model to understand the text semantics; and, through the cross-attention module, the association weights between the embedding vectors corresponding to the target text sequence are dynamically adjusted to capture the dependency between different parts of the text, thereby enhancing semantic fusion; for example: in the target text information "Remind me to measure blood pressure at 8 o'clock tomorrow morning", the cross-attention module will enhance the association between "remind" and "measure blood pressure", while weakening the direct association between "tomorrow" and "blood pressure"; finally, through the multi-head attention module, the semantic features of the first fusion information are extracted from different angles and integrated into a global context feature vector; for example: in the target text information "Remind me to measure blood pressure at 8 o'clock tomorrow morning", the multi-head attention module will extract the time feature: [tomorrow, 8 o'clock in the morning], the task feature: [measure blood pressure] and the action feature: [remind], and finally integrate the time feature, task feature and action feature into a context feature vector.
[0039] In one embodiment, before obtaining the target text sequence, the method further includes:
[0040] In response to a text input instruction, acquiring target text information according to the text input instruction;
[0041] The target text information is formatted to generate the target text sequence.
[0042] In this embodiment, the target text information refers to target text information input by a user from a user terminal via voice or text, etc.; wherein the user terminal can be various personal computers, laptops, smartphones, tablet computers, portable wearable devices, etc. Specifically, the server responds to the text input command sent by the user terminal and obtains the corresponding target text information according to the text input command; wherein the target text information can be text information such as medical health reminders (such as medication, vital sign monitoring, etc.), or text information such as financial customer service (such as account inquiries, financial consulting, risk assessment, etc.), without specific limitation here. For example: the target text information is "Remind me to measure my blood pressure at 8 o'clock tomorrow morning and record the data." In addition, after obtaining the target text information, in order to ensure the consistency of the text input, the target text information needs to be formatted to convert it into a target text sequence with a unified text sequence format; for example, "8 o'clock tomorrow morning" → the standardized time format "YYYY-MM-DD 08:00".
[0043] S120 , performing dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector.
[0044] In this embodiment, in order to reduce the amount of calculation and noise, this embodiment performs dimensionality reduction processing on the context feature vector through a latent space mapping module, thereby compressing the feature dimension and retaining the core information; wherein, the latent space mapping module is used to map the high-dimensional context feature vector (such as 2048 dimensions) to a low-dimensional space (such as 256 dimensions) while retaining semantic associations (such as the strong correlation between "reminder" and "measure blood pressure"); and, in this embodiment, the latent space mapping module is preferably a variational autoencoder (VAE), which is not specifically limited here.
[0045] S130: Input the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, to generate speech features corresponding to the first low-dimensional feature vector.
[0046] In this embodiment, in order to improve the quality of speech generation, this embodiment inputs the first low-dimensional feature vector into the target denoising subnetwork in the latent diffusion model to identify and remove the noise in the first low-dimensional feature vector, and generates clear speech features corresponding to the denoised feature vector, so that while accurately generating speech features corresponding to the target text information, noise and distortion are reduced, and the generated speech is further made more natural and smooth; wherein, the target denoising subnetwork is preferably a U-Net neural network in this embodiment.
[0047] In one embodiment, the step S130 includes:
[0048] Utilizing the target denoising sub-network to gradually denoise the low-dimensional feature vector to obtain a denoised feature vector;
[0049] Performing feature extraction on the noise reduction feature vector to obtain feature information;
[0050] The noise reduction feature vector and the feature information are integrated to generate the speech feature.
[0051] In this embodiment, the generation of speech features corresponding to the first low-dimensional feature vector is specifically carried out through the target denoising sub-network of the latent diffusion model, which gradually restores the input low-dimensional feature vector from the noise to a clear feature representation, that is, the denoised feature vector; and extracts key information (such as frequency, pitch, rhythm, etc.) from the denoised feature vector to obtain feature information; finally, the denoised feature vector and the extracted feature information are integrated to generate speech features; wherein, the speech features can be Mel-spectrograms.
[0052] In one embodiment, before step S130, the method further includes:
[0053] Acquire pre-stored historical text information from a local database, pre-process the historical text information based on a pre-processing strategy, and generate a second low-dimensional feature vector;
[0054] Using the second low-dimensional feature vector as a real sample data set; wherein the real sample data set includes at least one real sample data;
[0055] gradually adding Gaussian noise to the second low-dimensional feature vector in the latent space according to a preset noise strategy to obtain a noise data set;
[0056] Pairing each step of noise data in the noise data set with the real sample data corresponding to each step of noise data to obtain a training pair data set;
[0057] The training pair data set is input into the denoising sub-network to be trained in the potential diffusion model for training until the denoising loss function in the denoising sub-network to be trained meets the preset training stop condition, then the model training is stopped and the target denoising sub-network is obtained.
[0058] In this embodiment, first, the pre-stored historical text information is obtained from the local database, wherein the historical text information may be historical medical health reminder information, etc.; and the historical text information is pre-processed based on the pre-processing strategy to generate a second low-dimensional feature vector corresponding to the historical text; wherein the operation of pre-processing the historical text information based on the pre-processing strategy includes format conversion, encoding and dimensionality reduction of the historical text information, and the specific process of generating the second low-dimensional feature vector is similar to the above-mentioned steps S110 and S120, which will not be described here one by one. In addition, Gaussian noise is gradually added to the second low-dimensional feature vector in the latent space according to the preset noise strategy until a noise data set (Z0, Z1, ..., Z T ); and pairing each noise data in the noise data set with the real sample data corresponding to each noise data to obtain a training pair data set for training the denoising sub-network to be trained; wherein the training pair data set includes a plurality of training data pairs; and the real sample data corresponding to each noise data refers to the original data without noise, that is, each second low-dimensional feature vector in the second low-dimensional feature vector; and, noise data Z is obtained from the noise data set. T Initially, the training pair data set is input into the denoising sub-network to be trained for training, so as to reversely output the noise prediction value of each step of adding noise; and in the training process, by continuously adjusting the parameters of the denoising sub-network to be trained, the denoising loss function in the denoising sub-network to be trained satisfies the preset training stop condition, that is, the error between the noise prediction value and the actual added noise is minimized, thereby stopping the model training and obtaining the target denoising sub-network, so that the first low-dimensional feature vector can be denoised by the target denoising sub-network in the future, thereby accurately restoring the corresponding speech features.
[0059] S140: Convert the speech features using a waveform generator to generate target speech corresponding to the speech features.
[0060] In this embodiment, after generating the speech features corresponding to the first low-dimensional feature vector, since the speech features are in digital form that computers can understand and process, rather than sounds that humans can directly hear, this embodiment converts these digital speech features into analog audio waveform signals, i.e., the speech heard in daily life, by using a waveform generator. Furthermore, the waveform generator in this embodiment preferably uses a high-quality vocoder such as HiFi-GAN, VOCOS vocoder, or BIGVGAN, so that the speech features are restored to a high-fidelity audio waveform, thereby making the generated speech more natural and fluent.
[0061] In one embodiment, the step S140 includes:
[0062] Converting the speech feature into a first audio waveform using the waveform generator;
[0063] performing low-pass filtering on the first audio waveform to obtain a second audio waveform;
[0064] The phase information of the second audio waveform is optimized according to a preset phase recovery algorithm to obtain a target audio waveform, and the target audio waveform is output as the target speech.
[0065] In this embodiment, the waveform generator receives input speech features and converts them into an initial waveform shape, namely, a first audio waveform. The speech features include information such as pitch, intensity, timbre, and duration. Furthermore, after obtaining the first audio waveform, it is low-pass filtered to remove high-frequency interference, making the first audio waveform smoother and purer, thereby obtaining a second audio waveform that retains the primary audio features. Subsequently, because the aforementioned steps may affect phase information, resulting in a decrease in audio quality, the phase information of the second audio waveform is optimized according to a preset phase recovery algorithm to make the phase relationship of the second audio waveform on the time axis more accurate and reasonable, thereby reducing phase distortion and improving audio clarity. The phase recovery algorithm may be a Griffin-Lim algorithm. Finally, the optimized target audio waveform is output as the target speech, for example, in a medical health reminder scenario such as "Measure blood pressure at 8:00 AM tomorrow." The waveform is played to a target user via a smart device, allowing the user to clearly hear the reminder and perform corresponding health management actions based on the reminder.
[0066] It can be seen from the above technical solutions that the present invention obtains a target text sequence, inputs the target text sequence into a text encoder for encoding, and generates a context feature vector; performs dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector; inputs the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising processing to generate speech features; utilizes a waveform generator to convert the speech features to generate a target speech, thereby reducing the amount of high-dimensional space calculations and improving speech generation efficiency by performing dimensionality reduction processing on the context feature vector in the latent space; and through the denoising characteristics of the latent extension model, the generated speech is made more natural and smooth, noise and distortion are reduced, and the speech generation quality is improved, while being able to flexibly adapt to the generation requirements of different languages and contexts.
[0067] Figure 2 Schematic block diagram of a TTS generation device based on a potential diffusion model provided by an embodiment of the present invention. Figure 2 As shown, the present invention also provides a TTS generation device based on a potential diffusion model; the TTS generation device based on a potential diffusion model includes a unit for executing the TTS generation method based on a potential diffusion model, specifically, refer to Figure 2 The TTS generation device 100 based on the latent diffusion model includes a text sequence encoding unit 110, a latent space mapping unit 120, a diffusion denoising unit 130 and a target speech generation unit 140.
[0068] The text sequence encoding unit 110 is configured to obtain a target text sequence, input the target text sequence into a text encoder for encoding, and generate a context feature vector corresponding to the target text sequence.
[0069] In this embodiment, in order to convert the target text sequence into a numerical representation that can be understood by the server, after obtaining the target text sequence, this embodiment inputs the target text sequence into a text encoder for encoding, so as to effectively capture the context information in the target text sequence to generate a context feature vector corresponding to the target text sequence; wherein the context feature vector includes semantic information such as time, task, action, etc.; and the text encoder is preferably a Transformer encoder.
[0070] For example, the target text information is “Remind me to measure my blood pressure at 8 o'clock tomorrow morning”, and the context feature vector contains semantic information such as “remind”, “8 o'clock”, and “measure blood pressure”.
[0071] In one embodiment, the step of inputting the target text sequence into a text encoder for encoding to generate a context feature vector corresponding to the target text sequence is specifically used to:
[0072] Convert the target text sequence according to a preset word embedding table to obtain an embedding vector corresponding to the target text sequence;
[0073] fusing the embedding vectors through a cross-attention module of the text encoder to obtain first fused information;
[0074] The first fusion information is subjected to feature extraction using multiple multi-head attention modules in the text encoder to obtain the context feature vector.
[0075] In this embodiment, the target text sequence is first converted through a preset word embedding table to obtain an embedding vector corresponding to the target text sequence, so as to capture the semantic information of the target text sequence and facilitate the subsequent model to understand the text semantics; and, through the cross-attention module, the association weights between the embedding vectors corresponding to the target text sequence are dynamically adjusted to capture the dependency between different parts of the text, thereby enhancing semantic fusion; for example: in the target text information "Remind me to measure blood pressure at 8 o'clock tomorrow morning", the cross-attention module will enhance the association between "remind" and "measure blood pressure", while weakening the direct association between "tomorrow" and "blood pressure"; finally, through the multi-head attention module, the semantic features of the first fusion information are extracted from different angles and integrated into a global context feature vector; for example: in the target text information "Remind me to measure blood pressure at 8 o'clock tomorrow morning", the multi-head attention module will extract the time feature: [tomorrow, 8 o'clock in the morning], the task feature: [measure blood pressure] and the action feature: [remind], and finally integrate the time feature, task feature and action feature into a context feature vector.
[0076] In one embodiment, before obtaining the target text sequence, the method is further used to:
[0077] In response to a text input instruction, acquiring target text information according to the text input instruction;
[0078] The target text information is formatted to generate the target text sequence.
[0079] In this embodiment, the target text information refers to target text information input by a user from a user terminal via voice or text, etc.; wherein the user terminal can be various personal computers, laptops, smartphones, tablet computers, portable wearable devices, etc. Specifically, the server responds to the text input command sent by the user terminal and obtains the corresponding target text information according to the text input command; wherein the target text information can be text information such as medical health reminders (such as medication, vital sign monitoring, etc.), or text information such as financial customer service (such as account inquiries, financial consulting, risk assessment, etc.), without specific limitation here. For example: the target text information is "Remind me to measure my blood pressure at 8 o'clock tomorrow morning and record the data." In addition, after obtaining the target text information, in order to ensure the consistency of the text input, the target text information needs to be formatted to convert it into a target text sequence with a unified text sequence format; for example, "8 o'clock tomorrow morning" → the standardized time format "YYYY-MM-DD 08:00".
[0080] The latent space mapping unit 120 is configured to perform dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector.
[0081] In this embodiment, in order to reduce the amount of calculation and noise, this embodiment performs dimensionality reduction processing on the context feature vector through a latent space mapping module, thereby compressing the feature dimension and retaining the core information; wherein, the latent space mapping module is used to map the high-dimensional context feature vector (such as 2048 dimensions) to a low-dimensional space (such as 256 dimensions) while retaining semantic associations (such as the strong correlation between "reminder" and "measure blood pressure"); and, in this embodiment, the latent space mapping module is preferably a variational autoencoder (VAE), which is not specifically limited here.
[0082] The diffusion denoising unit 130 is configured to input the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, thereby generating speech features corresponding to the first low-dimensional feature vector.
[0083] In this embodiment, in order to improve the quality of speech generation, this embodiment inputs the first low-dimensional feature vector into the target denoising subnetwork in the latent diffusion model to identify and remove the noise in the first low-dimensional feature vector, and generates clear speech features corresponding to the denoised feature vector, so that while accurately generating speech features corresponding to the target text information, noise and distortion are reduced, and the generated speech is further made more natural and smooth; wherein, the target denoising subnetwork is preferably a U-Net neural network in this embodiment.
[0084] In one embodiment, the diffusion denoising unit 130 is specifically configured to:
[0085] Utilizing the target denoising sub-network to gradually denoise the low-dimensional feature vector to obtain a denoised feature vector;
[0086] Performing feature extraction on the noise reduction feature vector to obtain feature information;
[0087] The noise reduction feature vector and the feature information are integrated to generate the speech feature.
[0088] In this embodiment, the generation of speech features corresponding to the first low-dimensional feature vector is specifically carried out through the target denoising sub-network of the latent diffusion model, which gradually restores the input low-dimensional feature vector from the noise to a clear feature representation, that is, the denoised feature vector; and extracts key information (such as frequency, pitch, rhythm, etc.) from the denoised feature vector to obtain feature information; finally, the denoised feature vector and the extracted feature information are integrated to generate speech features; wherein, the speech features can be Mel-spectrograms.
[0089] In one embodiment, before the diffusion denoising unit 130, the unit is further configured to:
[0090] Acquire pre-stored historical text information from a local database, pre-process the historical text information based on a pre-processing strategy, and generate a second low-dimensional feature vector;
[0091] Using the second low-dimensional feature vector as a real sample data set; wherein the real sample data set includes at least one real sample data;
[0092] gradually adding Gaussian noise to the second low-dimensional feature vector in the latent space according to a preset noise strategy to obtain a noise data set;
[0093] Pairing each step of noise data in the noise data set with the real sample data corresponding to each step of noise data to obtain a training pair data set;
[0094] The training pair data set is input into the denoising sub-network to be trained in the potential diffusion model for training until the denoising loss function in the denoising sub-network to be trained meets the preset training stop condition, then the model training is stopped and the target denoising sub-network is obtained.
[0095] In this embodiment, first, the pre-stored historical text information is obtained from the local database, wherein the historical text information may be historical medical health reminder information, etc.; and the historical text information is pre-processed based on the pre-processing strategy to generate a second low-dimensional feature vector corresponding to the historical text; wherein the operation of pre-processing the historical text information based on the pre-processing strategy includes format conversion, encoding and dimensionality reduction of the historical text information, and the specific process of generating the second low-dimensional feature vector is similar to the above-mentioned steps S110 and S120, which will not be described here one by one. In addition, Gaussian noise is gradually added to the second low-dimensional feature vector in the latent space according to the preset noise strategy until a noise data set (Z0, Z1, ..., Z T ); and pairing each noise data in the noise data set with the real sample data corresponding to each noise data to obtain a training pair data set for training the denoising sub-network to be trained; wherein the training pair data set includes a plurality of training data pairs; and the real sample data corresponding to each noise data refers to the original data without noise, that is, each second low-dimensional feature vector in the second low-dimensional feature vector; and, noise data Z is obtained from the noise data set. T Initially, the training pair data set is input into the denoising sub-network to be trained for training, so as to reversely output the noise prediction value of each step of adding noise; and in the training process, by continuously adjusting the parameters of the denoising sub-network to be trained, the denoising loss function in the denoising sub-network to be trained satisfies the preset training stop condition, that is, the error between the noise prediction value and the actual added noise is minimized, thereby stopping the model training and obtaining the target denoising sub-network, so that the first low-dimensional feature vector can be denoised by the target denoising sub-network in the future, thereby accurately restoring the corresponding speech features.
[0096] The target speech generating unit 140 is configured to convert the speech features using a waveform generator to generate a target speech corresponding to the speech features.
[0097] In this embodiment, after generating the speech features corresponding to the first low-dimensional feature vector, since the speech features are in digital form that computers can understand and process, rather than sounds that humans can directly hear, this embodiment converts these digital speech features into analog audio waveform signals, i.e., the speech heard in daily life, by using a waveform generator. Furthermore, the waveform generator in this embodiment preferably uses a high-quality vocoder such as HiFi-GAN, VOCOS vocoder, or BIGVGAN, so that the speech features are restored to a high-fidelity audio waveform, thereby making the generated speech more natural and fluent.
[0098] In one embodiment, the target speech generation unit 140 is specifically configured to:
[0099] Converting the speech feature into a first audio waveform using the waveform generator;
[0100] performing low-pass filtering on the first audio waveform to obtain a second audio waveform;
[0101] The phase information of the second audio waveform is optimized according to a preset phase recovery algorithm to obtain a target audio waveform, and the target audio waveform is output as the target speech.
[0102] In this embodiment, the waveform generator receives input speech features and converts them into an initial waveform shape, namely, a first audio waveform. The speech features include information such as pitch, intensity, timbre, and duration. Furthermore, after obtaining the first audio waveform, it is low-pass filtered to remove high-frequency interference, making the first audio waveform smoother and purer, thereby obtaining a second audio waveform that retains the primary audio features. Subsequently, because the aforementioned steps may affect phase information, resulting in a decrease in audio quality, the phase information of the second audio waveform is optimized according to a preset phase recovery algorithm to make the phase relationship of the second audio waveform on the time axis more accurate and reasonable, thereby reducing phase distortion and improving audio clarity. The phase recovery algorithm may be a Griffin-Lim algorithm. Finally, the optimized target audio waveform is output as the target speech, for example, in a medical health reminder scenario such as "Measure blood pressure at 8:00 AM tomorrow." The waveform is played to a target user via a smart device, allowing the user to clearly hear the reminder and perform corresponding health management actions based on the reminder.
[0103] It can be seen from the above technical solutions that the present invention obtains a target text sequence, inputs the target text sequence into a text encoder for encoding, and generates a context feature vector; performs dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector; inputs the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising processing to generate speech features; utilizes a waveform generator to convert the speech features to generate a target speech, thereby reducing the amount of high-dimensional space calculations and improving speech generation efficiency by performing dimensionality reduction processing on the context feature vector in the latent space; and through the denoising characteristics of the latent extension model, the generated speech is made more natural and smooth, noise and distortion are reduced, and the speech generation quality is improved, while being able to flexibly adapt to the generation requirements of different languages and contexts.
[0104] The TTS generation device based on the potential diffusion model can be implemented in the form of a computer program. The computer program can be used in Figure 3 Runs on the computer equipment shown.
[0105] See also Figure 3 , Figure 3 is a schematic block diagram of a computer device provided in an embodiment of the present invention. Computer device 500 is a server, which can be a standalone server or a server cluster consisting of multiple servers. Computer device 500 can also be a transmitter, which can be a communication-capable electronic device such as a smartphone, tablet computer, laptop computer, desktop computer, personal digital assistant, or wearable device.
[0106] See Figure 3 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a storage medium 503 and an internal memory 504 .
[0107] The storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 may execute a TTS generation method based on a latent diffusion model.
[0108] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0109] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503 . When the computer program 5032 is executed by the processor 502 , the processor 502 can execute the TTS generation method based on the potential diffusion model.
[0110] The network interface 505 is used for network communication, such as providing data information transmission. Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0111] The processor 502 is configured to run a computer program 5032 stored in the memory to implement the TTS generation method based on the potential diffusion model disclosed in the embodiment of the present invention.
[0112] Those skilled in the art will understand that Figure 3 The embodiment of the computer device shown in the figure does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structure and function of the memory and processor are the same as those in the figure. Figure 3 The embodiments shown are consistent and will not be described again here.
[0113] It should be understood that in the embodiment of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0114] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when executed by a processor, the computer program implements the method for generating text-to-speech (TTS) based on a potential diffusion model disclosed in an embodiment of the present invention.
[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the devices, systems and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0116] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, or units with the same function may be combined into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, systems or units, or may be an electrical, mechanical or other form of connection.
[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0118] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0120] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A TTS generation method based on a latent diffusion model, characterized in that: The method comprises: Obtaining a target text sequence, inputting the target text sequence into a text encoder for encoding, and generating a context feature vector corresponding to the target text sequence; Performing dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector; Inputting the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, thereby generating speech features corresponding to the first low-dimensional feature vector; The speech features are converted using a waveform generator to generate target speech corresponding to the speech features.
2. The TTS generation method based on the latent diffusion model according to claim 1, characterized in that: Inputting the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising to generate speech features corresponding to the first low-dimensional feature vector includes: Utilizing the target denoising sub-network to gradually denoise the low-dimensional feature vector to obtain a denoised feature vector; Performing feature extraction on the noise reduction feature vector to obtain feature information; The noise reduction feature vector and the feature information are integrated to generate the speech feature.
3. The TTS generation method based on the latent diffusion model according to claim 2, characterized in that: Before inputting the first low-dimensional feature vector into the target denoising subnetwork in the latent diffusion model for denoising and generating speech features corresponding to the first low-dimensional feature vector, the method further includes: Acquire pre-stored historical text information from a local database, pre-process the historical text information based on a pre-processing strategy, and generate a second low-dimensional feature vector; Using the second low-dimensional feature vector as a real sample data set; wherein the real sample data set includes at least one real sample data; gradually adding Gaussian noise to the second low-dimensional feature vector in the latent space according to a preset noise strategy to obtain a noise data set; Pairing each step of noise data in the noise data set with the real sample data corresponding to each step of noise data to obtain a training pair data set; The training pair data set is input into the denoising sub-network to be trained in the potential diffusion model for training until the denoising loss function in the denoising sub-network to be trained meets the preset training stop condition, then the model training is stopped and the target denoising sub-network is obtained.
4. The TTS generation method based on the latent diffusion model according to claim 1, characterized in that: The converting the speech feature by using a waveform generator to generate a target speech corresponding to the speech feature includes: Converting the speech feature into a first audio waveform using the waveform generator; performing low-pass filtering on the first audio waveform to obtain a second audio waveform; The phase information of the second audio waveform is optimized according to a preset phase recovery algorithm to obtain a target audio waveform, and the target audio waveform is output as the target speech.
5. The TTS generation method based on the latent diffusion model according to claim 1, characterized in that: Inputting the target text sequence into a text encoder for encoding to generate a context feature vector corresponding to the target text sequence includes: Convert the target text sequence according to a preset word embedding table to obtain an embedding vector corresponding to the target text sequence; fusing the embedding vectors through a cross-attention module of the text encoder to obtain first fused information; The first fusion information is subjected to feature extraction using multiple multi-head attention modules in the text encoder to obtain the context feature vector.
6. The TTS generation method based on the latent diffusion model according to claim 5, characterized in that: Before obtaining the target text sequence, the method further includes: In response to a text input instruction, acquiring target text information according to the text input instruction; The target text information is formatted to generate the target text sequence.
7. The TTS generation method based on the latent diffusion model according to claim 4, characterized in that: The waveform generator includes a HiFi-GAN vocoder, a VOCOS vocoder, or a BIGVGAN vocoder.
8. A TTS generation device based on a latent diffusion model, characterized in that: The device comprises: A text sequence encoding unit, configured to obtain a target text sequence, input the target text sequence into a text encoder for encoding, and generate a context feature vector corresponding to the target text sequence; a latent space mapping unit, configured to perform dimensionality reduction processing on the context feature vector through a latent space mapping module to obtain a first low-dimensional feature vector; a diffusion denoising unit, configured to input the first low-dimensional feature vector into a target denoising subnetwork in a latent diffusion model for denoising, thereby generating speech features corresponding to the first low-dimensional feature vector; The target speech generating unit is used to convert the speech features using a waveform generator to generate a target speech corresponding to the speech features.
9. A computer device, characterized in that: The computer device includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the TTS generation method based on the potential diffusion model according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the method for generating TTS based on a latent diffusion model according to any one of claims 1 to 7 can be implemented.