Voice generation method and device based on self-guiding diffusion model, equipment and medium

By leveraging the collaborative work of a self-guided diffusion model, hierarchical and refined speech features are generated, addressing the accuracy and reliability issues of text-to-speech technology in the fields of fintech and healthcare and elderly care, and improving the reliability of speech generation in intelligent customer service systems.

CN121884758APending Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing text-to-speech technologies struggle to effectively handle complex technical terms, numbers, and specific rhythms in the fintech and healthcare/elderly care sectors, resulting in insufficient accuracy and reliability in speech generation, which may lead to customer disputes or medication safety risks.

Method used

A self-guided diffusion model is adopted, which generates semantic tag sequences by working together with a main semantic prediction model and a weakened semantic guidance model. Then, the Mel spectrogram is optimized by using coarse-grained and fine-grained diffusion models. Finally, the target speech information is generated through feature transformation technology, realizing hierarchical and refined speech feature generation.

Benefits of technology

It improves the reliability of text-to-speech conversion, reduces the probability of erroneous semantic tags, and ensures the accuracy of key information and the clarity of speech, making it suitable for intelligent customer service systems in the fields of fintech, healthcare, and elderly care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884758A_ABST
    Figure CN121884758A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech semantics, and discloses a speech generation method and device based on a self-guiding diffusion model, equipment and a medium, and the method comprises the steps: generating a semantic mark sequence; generating a Mel spectrogram according to the coarse-grained diffusion model, the fine-grained diffusion model and the semantic marking sequence; and generating target voice information through a feature conversion technology and the Mel spectrogram. By means of the mode, the weakening semantic guiding model and the main semantic prediction model work cooperatively, self-guiding enhancement is carried out in the semantic mark generation process, and the generation probability of wrong semantic marks is reduced. Through self-guiding optimization of coarse-grained and fine-grained two-stage diffusion models, hierarchical refined generation of speech features is realized, fine noise easily generated by a traditional single diffusion model is avoided, and the method can be applied to the business fields of financial science and technology, medical treatment, health, old-age care and the like, and has a wide application prospect. And the reliability of converting the text information into the voice information by the intelligent customer service system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech and semantic technology, and in particular to a speech generation method, apparatus, device and medium based on a self-guided diffusion model. Background Technology

[0002] In today's era of rapid advancements in artificial intelligence, text-to-speech systems, as a key interface for human-computer interaction, have been widely applied across various industries. Particularly in the critical sectors of fintech and healthcare / elderly care—two areas vital to public well-being and social efficiency—there is a tremendous demand for high-quality, highly reliable, and expressive speech synthesis technology.

[0003] In the fintech field, text-to-speech technology is widely used in scenarios such as intelligent customer service outbound calls, risk transaction confirmation, financial product promotion, bill reminders, and investor education. These scenarios place extremely high demands on the accuracy, professionalism, and credibility of the generated speech content. For example, when automatically broadcasting stock prices or transaction details, a misread or repetition of a single number can lead to serious customer disputes and financial losses; when promoting financial products, monotonous, mechanical, or noisy speech can significantly reduce user trust and willingness to purchase. However, most current business systems rely on traditional or seed-model-based hierarchical text-to-speech technologies, which are inadequate when dealing with the complex technical terminology, numbers, and specific rhythms in financial texts.

[0004] In the fields of healthcare and elderly care, text-to-speech technology plays an increasingly important role, such as in intelligent online consultation and triage, medication reminders and guidance, broadcasting chronic disease management advice, and providing information broadcasting services for visually impaired or elderly users. This field has specific requirements for the clarity, naturalness, and emotional appeal of the voice. For example, when broadcasting medication instructions to the elderly, the voice must be absolutely clear and unambiguous; any omission of content or background noise could lead to medication safety risks.

[0005] Therefore, improving the reliability of intelligent customer service systems in converting text information into voice information in business areas such as fintech, healthcare, and elderly care has become an urgent technical problem to be solved. Summary of the Invention

[0006] This application provides a speech generation method, apparatus, device, and medium based on a self-guided diffusion model to improve the reliability of intelligent customer service systems in converting text information into speech information.

[0007] In a first aspect, this application provides a speech generation method based on a self-guided diffusion model, the method comprising: Obtain the text sequence to be converted, and generate a semantic tag sequence using the main semantic prediction model, the weakened semantic guidance model, and the text sequence to be converted; A Mel spectrogram is generated based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence. The target speech information is generated by using a preset feature conversion technique and the Mel spectrogram.

[0008] Secondly, this application also provides a speech generation device based on a self-guided diffusion model, the device comprising: The semantic tag sequence generation module is used to obtain the text sequence to be converted and generate a semantic tag sequence through the main semantic prediction model, the weakened semantic guidance model and the text sequence to be converted; The Mel spectrogram generation module is used to generate a Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence. The target speech information generation module is used to generate target speech information through preset feature conversion technology and the Mel spectrogram.

[0009] Thirdly, this application also provides a computer device, the computer device including a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the speech generation method based on the self-guided diffusion model as described above.

[0010] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the speech generation method based on the self-guided diffusion model as described above.

[0011] This application discloses a speech generation method, apparatus, device, and medium based on a self-guided diffusion model. The speech generation method includes acquiring a text sequence to be converted, generating a semantic tag sequence using a main semantic prediction model, a weakened semantic guidance model, and the text sequence; generating a Mel spectrogram based on a coarse-grained diffusion model, a fine-grained diffusion model, and the semantic tag sequence; and generating target speech information using a preset feature conversion technique and the Mel spectrogram. Through this approach, the application utilizes the collaborative work of a weakened semantic guidance model and a main semantic prediction model to perform self-guided enhancement during the semantic tag generation process, reducing the probability of generating erroneous semantic tags. By optimizing the self-guided flow through coarse-grained and fine-grained diffusion models, hierarchical and refined speech feature generation is achieved, avoiding the subtle noise easily generated by traditional single diffusion models. This improves the reliability of converting text information into speech information in intelligent customer service systems in fields such as fintech, healthcare, and elderly care. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic flowchart of a speech generation method based on a self-guided diffusion model provided in an embodiment of this application; Figure 2 A schematic block diagram of a speech generation device based on a self-guided diffusion model provided for embodiments of this application; Figure 3 A schematic block diagram of the structure of a computer device provided for an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0016] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0017] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0018] This application provides a speech generation method, apparatus, device, and medium based on a self-guided diffusion model. The speech generation method based on this model can be applied to intelligent customer service systems. By coordinating a weakened semantic guidance model with a main semantic prediction model, self-guided enhancement is performed during the semantic tag generation process, reducing the probability of erroneous semantic tags. Through self-guided optimization of coarse-grained and fine-grained two-level diffusion models, hierarchical and refined speech feature generation is achieved, avoiding the subtle noise easily generated by traditional single diffusion models. This improves the reliability of converting text information into speech information in intelligent customer service systems in fields such as fintech, healthcare, and elderly care.

[0019] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0020] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a speech generation method based on a self-guided diffusion model, provided in an embodiment of this application. This speech generation method based on the self-guided diffusion model can be applied to intelligent customer service systems to improve the reliability of converting text information into speech information in business areas such as fintech, healthcare, and elderly care.

[0021] like Figure 1 As shown, the speech generation method based on the self-guided diffusion model specifically includes steps S10 to S30.

[0022] Step S10: Obtain the text sequence to be converted, and generate a semantic tag sequence using the main semantic prediction model, the weakened semantic guidance model, and the text sequence to be converted; Specifically, the original input text sequence is obtained, and the text is preprocessed for standardization, including text normalization (converting numbers and symbols into Chinese characters), word segmentation, etc. A pre-trained text encoder is used to convert the text sequence into a series of text feature vectors T.

[0023] The processed text sequence T is simultaneously input into the parallel-deployed main semantic prediction model. and weakened semantic guidance model Both models are autoregressive models used to predict semantically labeled sequences. .

[0024] Main semantic prediction model It is a large autoregressive Transformer model whose task is to predict the probability distribution of the next semantic tag based on the preceding context. Weakened semantic guidance model It is a model with the same or similar structure as the main model, but with fewer parameters or less training (such as fewer training rounds), and its prediction results are more uncertain and noisy than those of the main model.

[0025] When predicting the i-th semantic tag, the main model and the guiding model output probability distributions respectively. and By fusing the results using the following formula, a sharper and more accurate guided distribution can be obtained:

[0026] in, It is a guiding strength coefficient greater than 1.

[0027] From the merged distribution The semantic tag is obtained by sampling in the middle, and this tag is used as the input for the next time step until a complete semantic tag sequence is generated. .

[0028] Step S20: Generate a Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence; Specifically, the semantic tag sequence S is transformed into a condition vector c through a projection layer, which will serve as the global condition for the diffusion model.

[0029] Starting with random Gaussian noise, T denoising iterations are performed. At each iteration t, the current noise and the condition vector c are input in parallel into the coarse-grained master diffusion model. and coarse-grained guided models Their noise predictions were obtained respectively. and The final noise prediction is calculated using a self-guided fusion formula: ,in .

[0030] The denoising process begins again with random Gaussian noise. At each iteration t, the current noise, the conditional vector c, and the coarse-grained features are input into the fine-grained master diffusion model. and fine-grained guided models To obtain their predictions and The final noise prediction is calculated using the fusion formula: Independent strength coefficients are used for the coarse-grained and fine-grained stages. and It allows for layered and refined control of rhythmic contours and acoustic details, and through iterative denoising, ultimately generates a high-resolution, high-quality Mel spectrogram M.

[0031] Step S30: Generate target speech information using preset feature conversion technology and the Mel spectrogram.

[0032] Specifically, a high-performance neural vocoder is pre-trained or loaded. This vocoder is preferably a non-autoregressive feedforward neural network model, trained to directly map a Mel spectrogram to a time-domain waveform. The Mel spectrogram M is directly input into the vocoder, which converts the spectrogram into a corresponding time-domain speech waveform signal through one or a few forward propagations. The output is the target speech information, i.e., a playable audio file or data stream.

[0033] The intensity predictor is a differentiable function: The static hyperparameters that originally required manual preset will be replaced. Transformed into dynamic values ​​inferred in real time by the model based on the input content. The predictor It can be jointly optimized with the entire text-to-speech system, with the training objective being to minimize the reconstruction loss (such as L1 loss) between the final speech output and the real speech:

[0034] in This represents the entire text-to-speech generation process.

[0035] This embodiment discloses a speech generation method based on a self-guided diffusion model. The method includes acquiring a text sequence to be converted, generating a semantic tag sequence using a main semantic prediction model, a weakened semantic guidance model, and the text sequence; generating a Mel spectrogram based on a coarse-grained diffusion model, a fine-grained diffusion model, and the semantic tag sequence; and generating target speech information using a preset feature conversion technique and the Mel spectrogram. Through this approach, the weakened semantic guidance model and the main semantic prediction model work together to perform self-guided enhancement during the semantic tag generation process, reducing the probability of generating erroneous semantic tags. By optimizing the self-guided diffusion model through coarse-grained and fine-grained approaches, hierarchical and refined speech feature generation is achieved, avoiding the subtle noise easily generated by traditional single diffusion models. This improves the reliability of converting text information into speech information in intelligent customer service systems in fields such as fintech, healthcare, and elderly care.

[0036] based on Figure 1 In the illustrated embodiment, step S10 includes: The main output probability distribution of each semantic tag in the initial text sequence is generated by the main semantic prediction model. The weakened semantic guidance model generates the guidance output probability distribution of each semantic tag in the initial text sequence. The main output probability distribution and the guided output probability distribution are weighted and fused based on a preset semantic guidance strength coefficient to generate a self-guided probability distribution; Sample from the self-guided probability distribution to determine the current semantic tag, and generate the semantic tag sequence based on each of the current semantic tags.

[0037] Specifically, the system receives a preprocessed sequence of financial text, such as "Please confirm the transfer of 50,000 yuan to account number ending in 9988". This text sequence is simultaneously input into a main semantic prediction model and a weakened semantic guidance model. Based on its large parameter size and training data, the main model generates an accurate and high-confidence main output probability distribution for each semantic marker (such as key numbers and financial terms like "9988" and "50,000"). Meanwhile, the weakened guidance model (e.g., a version with fewer parameters or insufficient training) generates a relatively fuzzy and more uncertain guidance output probability distribution based on the same input.

[0038] A preset semantic guidance strength coefficient with a value greater than 1 (e.g., set to 2.5) is invoked, and the two distributions are weighted and fused in the logarithmic probability space. The current semantic tag is determined by sampling from the optimized self-guiding probability distribution.

[0039] In one embodiment, in the healthcare and elderly care business field, when inputting medical guidance text such as "Take one antihypertensive pill after breakfast every day, please do not forget," the main semantic prediction model, based on its powerful language understanding capabilities trained on a large-scale medical text corpus, generates a highly deterministic probability distribution for each semantic tag. For key dosage information such as "one pill," it points to the correct dosage representation with extremely high confidence. Meanwhile, the weakened semantic guidance model, due to its limited model capacity, produces a relatively ambiguous probability distribution when faced with professional medical terminology, potentially assigning similar probability weights to "one pill," "two pills," or even incorrect expressions.

[0040] By employing a high preset semantic guidance strength coefficient (considering the stringent accuracy requirements of medical information), the weighted calculation of the two probability distributions through a self-guided fusion algorithm significantly enhances the main model's confidence in correct labeling while effectively suppressing the uncertainty introduced by the guided model. Sampling from the optimized self-guided probability distribution allows for the near-deterministic selection of correct semantic labels, ensuring the generated semantic label sequence is completely accurate in key medical information such as "drug dosage" and "dosage time." This eliminates speech errors caused by semantic comprehension biases at the source, providing strong protection for patient safety.

[0041] In one embodiment, for the fintech business, a standardized sequence of financial text, such as "Your credit card bill payment date is the 15th of each month," is acquired and converted into a high-dimensional vector representation using a text encoder, initiating an autoregressive semantic tag prediction process. For each semantic tag location to be predicted, the following operations are performed synchronously: The text vectors and the generated tag sequence are input into the main semantic prediction model trained on a massive financial corpus to obtain its confidence distribution for each candidate tag; By feeding the same input into a weakened semantic guidance model with the same structure but half the number of parameters, a probability distribution with higher uncertainty in its output can be obtained. The dynamic strength predictor calculates the semantic guidance strength coefficient in real time based on the complexity of the current text (such as whether it contains numbers, technical terms, etc.). For key financial data, the coefficient will be automatically increased to above 2.0. The logarithms of the two probability distributions are weighted and fused using preset coefficients, with the main model distribution having a weight greater than 1 and the guide model distribution having a negative weight. This operation significantly enhances the high-confidence predictions of the main model and suppresses uncertainty. The current marker is determined by sampling from the optimized probability distribution, ensuring the accurate generation of key information such as numbers, dates, and amounts.

[0042] In a specific embodiment, before generating the self-guided probability distribution by weighting and fusing the main output probability distribution and the guided output probability distribution based on a preset semantic guidance strength coefficient, the process includes: The preset semantic guidance intensity coefficient corresponding to the coarse-grained diffusion model is predicted by a preset intensity prediction algorithm based on the text sequence to be converted and / or the semantic tag sequence.

[0043] Specifically, the original text sequence T to be converted (e.g., "Your credit card bill is due tomorrow and you need to pay 5,280 yuan") or the intermediate semantic tag sequence S generated by the semantic prediction model is used as the input to the intensity predictor.

[0044] If the input is text T, it is first converted into a dense sequence of word vectors through an embedding layer.

[0045] If the input is a semantically labeled sequence S, it can be used directly or adjusted through a linear projection layer. In practical applications, the semantically labeled sequence S is a better input choice.

[0046] The prepared input feature sequence is fed into a lightweight encoder to capture its global contextual information, generating a context vector that represents the overall complexity and potential risk of the input text. For example, for text containing long numbers, technical terms, or complex sentence structures, this vector should reflect its "high difficulty" characteristic.

[0047] The generated context vector is mapped through a fully connected layer, ultimately outputting a scalar value. The fully connected layer typically ends with a linear layer and may use an activation function to constrain the output value to a reasonable range of positive numbers (e.g., between 1.0 and 5.0), serving as a coarse-grained guiding strength coefficient for dynamic prediction.

[0048] based on Figure 1 In the illustrated embodiment, before step S20, the following steps are included: Configure a self-guided diffusion model for the coarse-grained diffusion model to generate a coarse-grained guided model; Configure a self-guided diffusion model for the fine-grained diffusion model to generate a fine-grained guided model.

[0049] Specifically, the guiding model maintains architectural isomorphism with the main model but has weakened performance, serving as a coarse-grained main diffusion model. Configuration of coarse-grained guiding model And with fine-grained master diffusion models Paired fine-grained guided model The model maintains consistency with the main model in its basic building blocks (such as residual layers and attention mechanisms) and conditional input methods to ensure that they make predictions in the same feature space.

[0050] Systematically weaken its performance through one or more of the following strategies: 1. Reduce model capacity, for example, halve the number of residual channels in the main model; 2. Introduce regularization constraints, add constraints to its weights, limit its fitting ability, so that when the guiding model understands the same conditional information (such as semantic labels, coarse-grained spectrum), it will produce higher uncertainty, more fuzzy or suboptimal predictions, thereby providing an effective "opposite comparison" anchor point for the optimization direction of the main model.

[0051] During training, the guiding model receives the exact same inputs and conditions as the main model, but its training objective is designed to share the final reconstruction loss with the main model, rather than being subject to strong constraints independently. Therefore, the guiding model can learn the data distribution, but due to its weakened structure, it cannot achieve the accuracy of the main model. During inference, these two guiding models are tightly integrated into the generation process of their corresponding main models: in each denoising step of coarse-grained and fine-grained diffusion, the guiding model and the main model infer forward in parallel, and their output is used to perform self-guided fusion computation, thereby correcting the generation path of the main model in real time and step-by-step, guiding it away from low-quality regions, and ultimately co-generating high-quality, high-fidelity acoustic features.

[0052] In a specific embodiment, step S20 includes: Coarse-grained acoustic features are generated based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence; The Mel spectrum is generated based on the fine-grained diffusion model, the fine-grained guidance model, and the coarse-grained acoustic features.

[0053] Specifically, the semantic tag sequence is used as a conditional input and fed into the coarse-grained main diffusion model and its corresponding coarse-grained guided model in parallel. In each denoising iteration, the two models output noise predictions respectively. They are weighted and fused by a preset guided intensity coefficient to form a guided noise prediction and perform denoising operation, ultimately generating coarse-grained acoustic features containing global prosodic contours.

[0054] The coarse-grained acoustic features and semantic tag sequences are concatenated and used as conditional inputs to the fine-grained main diffusion model and its guiding model. In each denoising step, a self-guided fusion strategy is also adopted, using the fine-grained guiding intensity coefficient to fuse and denoise the predictions of the two models. Through this hierarchical processing, a Mel spectrogram with rich details and pure sound quality is finally output.

[0055] In a specific embodiment, generating coarse-grained acoustic features based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence includes: In each denoising time step, coarse-grained main noise prediction information is generated by the coarse-grained diffusion model based on the semantic tag sequence, and coarse-grained guided noise prediction information is generated by the coarse-grained guided model based on the semantic tag sequence. The coarse-grained main noise prediction information and the coarse-grained guided noise prediction information are weighted and fused according to a preset coarse-grained guiding intensity coefficient to generate the coarse-grained acoustic features.

[0056] Specifically, the current noisy acoustic feature state and semantic tag sequence are input together into a coarse-grained diffusion model, which generates coarse-grained main noise prediction information for the target noise. Meanwhile, the same input is fed into a weakened coarse-grained guidance model, which generates a coarse-grained guidance noise prediction information with higher uncertainty. The preset coarse-grained guiding intensity system is invoked, and the self-guided fusion formula is applied. Weighted fusion is performed. This process is repeated iteratively until the diffusion process ends, ultimately outputting clean, coarse-grained acoustic features that contain the global prosodic contour.

[0057] Based on any of the above embodiments, in this embodiment, step S30 includes: Extract the Mel scale amplitude features from the Mel spectrogram; The preset feature conversion technology is used to perform phase reconstruction, upsampling and waveform synthesis on the Mel scale amplitude features to generate a time-domain signal corresponding to the Mel spectrum. The time-domain signal is converted into an audio signal to generate the target speech information.

[0058] Specifically, the Mel scale amplitude feature matrix is ​​directly obtained from the generated Mel spectrogram. This feature matrix is ​​then input into a pre-trained non-autoregressive neural vocoder. As the core of the preset feature conversion technology, the non-autoregressive neural vocoder first performs phase reconstruction on the amplitude features through its internal neural network to complete the phase information necessary for the time domain signal.

[0059] Upsampling is performed by transposed convolution or interpolation layers to extend the length of the feature sequence to the sampling rate dimension of the target audio. The output layer of the vocoder directly synthesizes the processed features into a complete time-domain waveform signal. This time-domain signal is then converted into a playable audio signal through a digital-to-analog converter or audio interface, thereby generating the final target speech information.

[0060] Please see Figure 2 , Figure 2 This application provides a schematic block diagram of a speech generation apparatus based on a self-guided diffusion model, which is used to execute the aforementioned speech generation method based on a self-guided diffusion model. The speech generation apparatus based on the self-guided diffusion model can be configured on a server.

[0061] like Figure 2 As shown, the speech generation device 400 based on the self-guided diffusion model includes: The semantic tag sequence generation module 410 is used to obtain the text sequence to be converted and generate a semantic tag sequence through the main semantic prediction model, the weakened semantic guidance model and the text sequence to be converted; Mel spectrogram generation module 420 is used to generate a Mel spectrogram based on a coarse-grained diffusion model, a fine-grained diffusion model, and the semantic tag sequence; The target speech information generation module 430 is used to generate target speech information through preset feature conversion technology and the Mel spectrogram.

[0062] Furthermore, the semantic tag sequence generation module 410 includes: The main output probability distribution generation unit is used to generate the main output probability distribution of each semantic tag in the initial text sequence through the main semantic prediction model. The guided output probability distribution generation unit is used to generate the guided output probability distribution of each semantic tag in the initial text sequence through the weakened semantic guidance model; The self-guided probability distribution generation unit is used to weight and fuse the main output probability distribution and the guided output probability distribution based on a preset semantic guidance strength coefficient to generate a self-guided probability distribution; A semantic tag sequence generation unit is used to sample from the self-guided probability distribution, determine the current semantic tag, and generate the semantic tag sequence based on each of the current semantic tags.

[0063] Furthermore, the semantic tag sequence generation module 410 includes: The guidance intensity coefficient prediction unit is used to predict the preset semantic guidance intensity coefficient corresponding to the coarse-grained diffusion model based on the text sequence to be converted and / or the semantic tag sequence using a preset intensity prediction algorithm.

[0064] Furthermore, the speech generation device 400 based on the self-guided diffusion model includes: A coarse-grained guided model generation module is used to configure a self-guided diffusion model for the coarse-grained diffusion model and generate a coarse-grained guided model. The fine-grained guided model generation module is used to configure a self-guided diffusion model for the fine-grained diffusion model and generate a fine-grained guided model.

[0065] Furthermore, the Mel spectrogram generation module 420 includes: A coarse-grained acoustic feature generation unit is used to generate coarse-grained acoustic features based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence; The Mel spectrogram generation unit is used to generate the Mel spectrogram based on the fine-grained diffusion model, the fine-grained guidance model, and the coarse-grained acoustic features.

[0066] Furthermore, the coarse-grained acoustic feature generation unit includes: The coarse-grained guided noise prediction information generation subunit is used to generate coarse-grained main noise prediction information according to the semantic tag sequence through the coarse-grained diffusion model in each denoising time step, and to generate coarse-grained guided noise prediction information according to the semantic tag sequence through the coarse-grained guided model. The coarse-grained acoustic feature generation subunit is used to weight and fuse the coarse-grained main noise prediction information and the coarse-grained guided noise prediction information according to a preset coarse-grained guided intensity coefficient to generate the coarse-grained acoustic features.

[0067] Furthermore, the target voice information generation module 430 includes: The Mel scale amplitude feature extraction unit is used to extract the Mel scale amplitude features of the Mel spectrum. The time-domain signal generation unit is used to perform phase reconstruction processing, upsampling processing, and waveform synthesis processing on the Mel scale amplitude features through the preset feature conversion technology to generate a time-domain signal corresponding to the Mel spectrum. The target speech information generation unit is used to convert the time-domain signal into an audio signal to generate the target speech information.

[0068] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0069] The aforementioned device can be implemented as a computer program, which can be used in, for example... Figure 3 It runs on the computer device shown.

[0070] Please see Figure 3 , Figure 3 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.

[0071] See Figure 3 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0072] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any speech generation method based on a self-guided diffusion model.

[0073] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0074] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any speech generation method based on a self-guided diffusion model.

[0075] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0076] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0077] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Obtain the text sequence to be converted, and generate a semantic tag sequence using the main semantic prediction model, the weakened semantic guidance model, and the text sequence to be converted; A Mel spectrogram is generated based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence. The target speech information is generated by using a preset feature conversion technique and the Mel spectrogram.

[0078] In one embodiment, a semantic tag sequence is generated using a main semantic prediction model, a weakened semantic guidance model, and the text sequence to be converted, for the purpose of: The main output probability distribution of each semantic tag in the initial text sequence is generated by the main semantic prediction model. The weakened semantic guidance model generates the guidance output probability distribution of each semantic tag in the initial text sequence. The main output probability distribution and the guided output probability distribution are weighted and fused based on a preset semantic guidance strength coefficient to generate a self-guided probability distribution; Sample from the self-guided probability distribution to determine the current semantic tag, and generate the semantic tag sequence based on each of the current semantic tags.

[0079] In one embodiment, before generating the self-guiding probability distribution by weighting and fusing the main output probability distribution and the guiding output probability distribution based on a preset semantic guidance strength coefficient, the following is used to achieve: The preset semantic guidance intensity coefficient corresponding to the coarse-grained diffusion model is predicted by a preset intensity prediction algorithm based on the text sequence to be converted and / or the semantic tag sequence.

[0080] In one embodiment, before generating the Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence, the following is performed: Configure a self-guided diffusion model for the coarse-grained diffusion model to generate a coarse-grained guided model; Configure a self-guided diffusion model for the fine-grained diffusion model to generate a fine-grained guided model.

[0081] In one embodiment, a Mel spectrogram is generated based on a coarse-grained diffusion model, a fine-grained diffusion model, and the semantic tag sequence to achieve: Coarse-grained acoustic features are generated based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence; The Mel spectrum is generated based on the fine-grained diffusion model, the fine-grained guidance model, and the coarse-grained acoustic features.

[0082] In one embodiment, coarse-grained acoustic features are generated based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence, for the purpose of: In each denoising time step, coarse-grained main noise prediction information is generated by the coarse-grained diffusion model based on the semantic tag sequence, and coarse-grained guided noise prediction information is generated by the coarse-grained guided model based on the semantic tag sequence. The coarse-grained main noise prediction information and the coarse-grained guided noise prediction information are weighted and fused according to a preset coarse-grained guiding intensity coefficient to generate the coarse-grained acoustic features.

[0083] In one embodiment, target speech information is generated using a preset feature transformation technique and the Mel spectrogram, for the purpose of: Extract the Mel scale amplitude features from the Mel spectrogram; The preset feature conversion technology is used to perform phase reconstruction, upsampling and waveform synthesis on the Mel scale amplitude features to generate a time-domain signal corresponding to the Mel spectrum. The time-domain signal is converted into an audio signal to generate the target speech information.

[0084] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the speech generation methods based on the self-guided diffusion model provided in the embodiments of this application.

[0085] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A speech generation method based on a self-guided diffusion model, characterized by, include: Obtain the text sequence to be converted, and generate a semantic tag sequence using the main semantic prediction model, the weakened semantic guidance model, and the text sequence to be converted; A Mel spectrogram is generated based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence. The target speech information is generated by using a preset feature conversion technique and the Mel spectrogram.

2. The speech generation method based on a self-bootstrapping diffusion model according to claim 1, wherein, The step of generating a semantically labeled sequence using the main semantic prediction model, the weakened semantic guidance model, and the text sequence to be converted includes: The main output probability distribution of each semantic tag in the initial text sequence is generated by the main semantic prediction model. The weakened semantic guidance model generates the guidance output probability distribution of each semantic tag in the initial text sequence. The main output probability distribution and the guided output probability distribution are weighted and fused based on a preset semantic guidance strength coefficient to generate a self-guided probability distribution; Sample from the self-guided probability distribution to determine the current semantic tag, and generate the semantic tag sequence based on each of the current semantic tags.

3. The speech generation method based on a self-guided diffusion model according to claim 2, characterized in that, Before generating the self-guided probability distribution by weighting and fusing the main output probability distribution and the guided output probability distribution based on a preset semantic guidance strength coefficient, the process includes: The preset semantic guidance intensity coefficient corresponding to the coarse-grained diffusion model is predicted by a preset intensity prediction algorithm based on the text sequence to be converted and / or the semantic tag sequence.

4. The speech generation method based on a self-guided diffusion model according to claim 1, characterized in that, Before generating the Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence, the following steps are included: Configure a self-guided diffusion model for the coarse-grained diffusion model to generate a coarse-grained guided model; Configure a self-guided diffusion model for the fine-grained diffusion model to generate a fine-grained guided model.

5. The speech generation method based on a self-guided diffusion model according to claim 4, characterized in that, The step of generating the Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence includes: Coarse-grained acoustic features are generated based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence; The Mel spectrum is generated based on the fine-grained diffusion model, the fine-grained guidance model, and the coarse-grained acoustic features.

6. The speech generation method based on a self-guided diffusion model according to claim 5, characterized in that, The step of generating coarse-grained acoustic features based on the coarse-grained diffusion model, the coarse-grained guidance model, and the semantic tag sequence includes: In each denoising time step, coarse-grained main noise prediction information is generated by the coarse-grained diffusion model based on the semantic tag sequence, and coarse-grained guided noise prediction information is generated by the coarse-grained guided model based on the semantic tag sequence. The coarse-grained main noise prediction information and the coarse-grained guided noise prediction information are weighted and fused according to a preset coarse-grained guiding intensity coefficient to generate the coarse-grained acoustic features.

7. The speech generation method based on a self-guided diffusion model according to any one of claims 1 to 6, characterized in that, The generation of target speech information through preset feature conversion technology and the Mel spectrogram includes: Extract the Mel scale amplitude features from the Mel spectrogram; The preset feature conversion technology is used to perform phase reconstruction, upsampling and waveform synthesis on the Mel scale amplitude features to generate a time-domain signal corresponding to the Mel spectrum. The time-domain signal is converted into an audio signal to generate the target speech information.

8. A speech generation device based on a self-guided diffusion model, characterized in that, include: The semantic tag sequence generation module is used to obtain the text sequence to be converted and generate a semantic tag sequence through the main semantic prediction model, the weakened semantic guidance model and the text sequence to be converted; The Mel spectrogram generation module is used to generate a Mel spectrogram based on the coarse-grained diffusion model, the fine-grained diffusion model, and the semantic tag sequence. The target speech information generation module is used to generate target speech information through preset feature conversion technology and the Mel spectrogram.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and, in executing the computer program, implement the speech generation method based on the self-guided diffusion model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to implement the speech generation method based on the self-guided diffusion model as described in any one of claims 1 to 7.