Method and device for improving speech synthesis speed of Cauchy denoising diffusion probability model

By introducing the Cauchy denoising diffusion probability model, defining and optimizing the relevant network structure and sampling method, the problem of slow speech synthesis speed was solved, achieving efficient speech synthesis, meeting real-time requirements and improving speech quality.

CN121999754APending Publication Date: 2026-05-08ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-02-02
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing denoising diffusion probability models have too slow sampling speeds in speech synthesis, making it difficult to meet real-time requirements, especially when dealing with imbalanced speech data.

Method used

A Cauchy denoising diffusion probability model is introduced. By defining the Cauchy prior square scale long table, the Cauchy posterior square scale long table, the Cauchy single-step diffusion operation, and the Cauchy multi-step diffusion operation, a denoising neural network and a Cauchy square scale mapping neural network are constructed. Grid search and fast sampling methods are used to optimize the model parameters to improve sampling efficiency.

Benefits of technology

It significantly improves the speech synthesis speed, meets real-time requirements, and at the same time alleviates the performance degradation caused by the speed increase of speech synthesis, thereby improving the quality of synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999754A_ABST
    Figure CN121999754A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for improving speech synthesis speed of a Cauchy denoising diffusion probability model. The method comprises the following steps: (1) defining a speech synthesis-oriented Cauchy denoising diffusion probability model; (2) defining a loss function of the speech synthesis-oriented Cauchy denoising diffusion probability model; (3) constructing and optimizing a denoising neural network oriented to speech synthesis; (4) defining a loss function of the speech synthesis-oriented Cauchy square scale mapping network; (5) constructing and optimizing a voice synthesis-oriented Cauchy square scale mapping network; (6) defining and executing a voice synthesis-oriented Cauchy square scale optimal short table search method; and (7) defining and executing a voice synthesis-oriented Cauchy fast sampling method. According to the invention, the quality and diversity of the synthesized speech can be effectively improved on the premise of meeting the real-time performance of speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, and in particular to a method and apparatus for improving the speech synthesis speed of the Cauchy denoising diffusion probability model. Background Technology

[0002] Deep generative models have achieved significant breakthroughs and excellent performance in the field of speech synthesis. Current mainstream models in this area include stream-based generative models, adversarial neural network-based generative models, and deep generative models based on denoising diffusion probability models. Benefiting from the stable training and diverse generated samples of denoising diffusion probability models, this type of deep generative model is gradually leading the latest trends and developments in the field of speech synthesis.

[0003] For example, Chinese patent document CN120998174A discloses a speech synthesis method, a training method for a diffusion model, an apparatus, and a device. Based on a text encoding sub-model of the diffusion model, text information is encoded to obtain a vector sequence; based on an acoustic feature extraction sub-model, acoustic features are extracted from the vector sequence to obtain a first Mel spectrum; based on a context-aware sub-model, text-speech alignment is performed on the first Mel spectrum according to the vector sequence to obtain a second Mel spectrum; based on a high-frequency compensation diffusion sub-model, residual learning and multi-scale acoustic feature extraction are performed on the second Mel spectrum to obtain a target Mel spectrum; based on the diffusion sub-model and the target Mel spectrum, the target speech is determined to improve the speech synthesis effect.

[0004] Due to the objective constraints of diffusion process theory, current research on denoising diffusion probability models for speech synthesis mainly focuses on the addition and removal of Gaussian noise. In order to address the problem of imbalanced speech data, researchers have proposed a denoising diffusion probability model based on the addition and removal of heavy-tailed noise, which effectively improves the quality and diversity of synthesized speech.

[0005] For example, Chinese patent document CN119049446A discloses a speech synthesis method and device based on the Cauchy denoising probability diffusion model. Cauchy noise is introduced into the denoising probability diffusion model to realize the training and sampling of the diffusion model, and finally completes the speech synthesis, which improves the robustness of the speech synthesis method and effectively improves the quality of synthesized speech.

[0006] One limitation of denoising diffusion probability models is their slow sampling rate, which translates to difficulty in achieving real-time speech synthesis. On one hand, current denoising diffusion probability models for fast speech synthesis focus on Gaussian noise, making it difficult to address the challenges posed by imbalanced speech data. On the other hand, due to changes in underlying theory, designing methods and devices for denoising diffusion probability models based on heavy-tailed noise for fast speech synthesis faces challenges.

[0007] For example, Chinese patent document CN120877701A discloses a system and method for improving the speed of diffusion model speech synthesis. It can generate acoustic features with fewer iterations, pass the acoustic features to the trained or fine-tuned vocoder to synthesize speech signals, improve the speed of speech synthesis, and generate high-quality speech signals. Summary of the Invention

[0008] This invention provides a method and apparatus for improving the speech synthesis speed of the Cauchy denoising diffusion probability model, which can significantly improve the speech synthesis speed and enhance the quality of synthesized speech.

[0009] A method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model includes the following steps: (1) Define a Cauchy denoising diffusion probability model for speech synthesis, including a Cauchy prior square scale long table, a Cauchy posterior square scale long table, a Cauchy single-step diffusion operation, and a Cauchy multi-step diffusion operation. (2) Construct a denoising neural network to predict Cauchy noise and Cauchy posterior squared scale for all diffusion steps; define the first loss function based on the Cauchy prior squared scale table and the Cauchy posterior squared scale table, and optimize the parameters of the denoising neural network using the first loss function; (3) Define a Cauchy squared scale mapping neural network, use a denoising neural network to predict Cauchy noise as the true value, and construct a second loss function to optimize the parameters of the Cauchy squared scale mapping neural network; (4) Define the best short table search method of Cauchy square scale. Based on the optimized Cauchy square scale mapping neural network, the random and deterministic iterative sampling methods of the Cauchy denoising diffusion probability model are used, and the grid search method is adopted to search for the best Cauchy prior square scale short table. (5) Define the Cauchy fast sampling method, calculate the Cauchy prior square scale short table corresponding to the Cauchy posterior square scale short table in an approximate manner, define the single-step sampling operation based on the Cauchy prior square scale short table and the Cauchy posterior square scale short table, and execute the single-step sampling operation in an iterative manner to achieve fast speech synthesis.

[0010] This invention introduces a sampling module (containing a Cauchy squared-scale mapping neural network, a Cauchy squared-scale optimal short table search method, and a Cauchy fast sampling method) based on the Cauchy denoising diffusion probability model to achieve fast speech synthesis using the Cauchy denoising diffusion probability model.

[0011] In step (1), the Cauchy prior square scale length table is defined as follows: ; in, and Let represent the number of diffusion steps of the two Gaussian denoising diffusion probability models in the prior square-scale long table. The prior squared scale value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model on a Cauchy prior square scale long table is... The prior squared scale value at that time; The Cauchy posterior square scale length table is defined as follows: ; in, and Let represent the number of diffusion steps in the posterior squared scale long table for the two Gaussian denoised diffusion probability models. The posterior squared scale value at time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy posterior squared scale table is... The posterior squared scale value at time; The Cauchy single-step diffusion operation is defined as follows: ; in, Indicates the input voice signal; and These represent the number of diffusion steps in the Cauchy prior square-scale long table, respectively. and The voice signal at that time; The Cauchy multi-step diffusion operation is defined as follows: ; ; in, This indicates the number of diffusion steps in the prior square-scale long table. Standard Cauchy noise sampled at time, The square root of the cumulative residual scale when the number of diffusion steps is t; In step (2), the first loss function is defined based on the Cauchy prior squared scale table and the Cauchy posterior squared scale table, specifically as follows: ; ; ; ; ; in, Indicates the input voice signal; Represents the first loss function; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The loss function value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The noise prediction loss value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The posterior squared scale predicted loss value at that time; This parameter represents the trade-off between the noise prediction loss value and the posterior squared scale prediction loss value. This indicates that the number of diffusion steps of the denoising neural network in the Cauchy prior square-scale long table is . Cauchy noise prediction value at that time; This indicates the number of diffusion steps in the Cauchy prior square-scale long table. The true value of Cauchy noise at that time; This indicates that the number of diffusion steps of the denoising neural network in the Cauchy prior square-scale long table is . Cauchy's posterior squared scale prediction at time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The true value of the posterior squared scale at time.

[0012] In step (3), the Cauchy squared-scale mapping neural network is defined, including: defining the Cauchy prior squared-scale mapping method, the iterative operation of the Cauchy prior squared-scale short table, and the second loss function of the Cauchy squared-scale mapping neural network; specifically as follows: Define the Cauchy prior squared scale mapping method: ; in, Represents the length of Cauchy's a priori square scale; Represents the short table of Cauchy's a priori square scale; This indicates the number of diffusion steps in the Cauchy prior square-scale long table. The voice signal at that time; This indicates the number of diffusion steps in the short table of Cauchy's prior square scale. The voice signal at that time; Define the iteration operation of the Cauchy prior square scale short table: ; ; in, and represent the prior square scale values ​​in the Cauchy prior square scale short table when the number of diffusion steps is n and n+1, respectively; For the Cauchy squared-scale mapping neural network, the prior squared-scale iterative operation is used during training. This represents a deep neural network that receives a speech signal after a diffusion step of t. As input, predict the ratio of the prior square scale values ​​of two adjacent diffusion steps in the Cauchy prior square scale short table; Define the second loss function for the Cauchy squared-scale mapping neural network: ; ; in, This represents the Cauchy noise prediction value of the optimized denoising neural network when the number of diffusion steps is t in the prior square scale table; This represents the true value of Cauchy noise randomly sampled in a short table with n diffusion steps in the prior square scale.

[0013] The specific process of step (4) is as follows: The definition of the Cauchy squared scale mapping iteration is ; ; in, This indicates that the optimized Cauchy squared-scale mapping neural network, given... The ratio of adjacent elements in the short table of Cauchy square scale predicted in real time; Indicates that in a given and At that time, the calculation was based on the optimized Cauchy squared scale mapping neural network. The formula; Indicates that in a given and Time calculation The formula; Based on the aforementioned calculations and The definition of random and deterministic iterative sampling methods ; in, and These represent the speech signals with diffusion steps of n and n+1 in the short table of Cauchy's a prior square scale, respectively; and These represent the Cauchy noise value and Cauchy squared scale value predicted by the optimized Cauchy denoising diffusion probability model, respectively. and At that time, deterministic sampling and random sampling were used respectively; Given initial value and Through the aforementioned iterative calculation , and To achieve speech synthesis; divided into grids. and The range of values ​​is used to perform speech synthesis and evaluate the quality of the synthesized speech, thereby achieving the search for the optimal Cauchy prior square scale short table.

[0014] In step (5), the Cauchy a prior square scale short table corresponding to the Cauchy posterior square scale short table is calculated in an approximate manner, using the following formula: ; in, and The optimal Cauchy square scale short table search method is given; This represents the approximate Cauchy posterior squared scale short table.

[0015] In step (5), a single-step sampling operation is defined based on the Cauchy prior square scale short table and the Cauchy posterior square scale short table, specifically as follows: ; in, and These represent the speech signals with diffusion steps of n-1 and n in the short table of Cauchy's a prior square scale, respectively; This represents the noise prediction value of the optimized Cauchy denoising diffusion probability model; This represents the true value of Cauchy noise randomly sampled in the prior square-scale short table when the number of diffusion steps is n; using the Mel spectrogram as a conditional input, single-step sampling operations are continuously executed to achieve fast speech synthesis.

[0016] An apparatus for improving the speech synthesis speed of a Cauchy denoising diffusion probability model is characterized by comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the aforementioned method for improving the speech synthesis speed of a Cauchy denoising diffusion probability model.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention can effectively improve the speech synthesis speed based on the Cauchy denoising diffusion probability model, easily meet the real-time requirements of speech synthesis, and has significant effects on the implementation of the technology.

[0018] 2. This invention is only applicable to the Cauchy denoising diffusion probability model for speech synthesis, which effectively alleviates the performance degradation caused by the speed increase of speech synthesis while improving the speech synthesis speed. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a method for improving the speech synthesis speed of a Cauchy denoising diffusion probability model according to an embodiment of the present invention.

[0021] Figure 2 The graph shows a comparison of the performance degradation of the proposed method for fast speech synthesis under different datasets. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.

[0024] This embodiment uses the widely used English speech corpus LJSpeech and the VCTK dataset. The LJSpeech dataset contains 13,100 audio files recorded by a female speaker, with a total duration of approximately 24 hours. In this embodiment, 100 audio files are randomly selected as the test set, and the remaining 13,000 audio files are used as the training set. The VCTK dataset contains 44,455 audio files recorded by 110 speakers, with a total duration of approximately 44 hours. This embodiment uses the VCTK dataset to create two test sets: the first test set, VCTK1, randomly selects 5 speakers with the same accent; the second test set, VCTK2, randomly selects 11 speakers with different accents. It is important to note that in this embodiment, 10 audio files are randomly selected from each test speaker in VCTK1 and VCTK2 to form the test set, and all audio files from the remaining speakers form the training set.

[0025] like Figure 1 As shown, a method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model includes the following steps: S01 defines a Cauchy denoising diffusion probability model for speech synthesis.

[0026] Define the prior square scale length table for the Cauchy denoising diffusion probability model for speech synthesis, the posterior square scale length table for the Cauchy denoising diffusion probability model, and the single-step diffusion operation and multi-step diffusion operation for the Cauchy denoising diffusion probability model.

[0027] S02 defines the first loss function for the Cauchy denoising diffusion probability model for speech synthesis.

[0028] Based on the prior squared scale length table and the posterior squared scale length table of the Cauchy denoising diffusion probability model calculated in S01, define the first loss function and calculate the true Cauchy noise and posterior squared scale for all diffusion steps.

[0029] S03, Construct and optimize a denoising neural network for speech synthesis.

[0030] A denoising neural network is constructed. Based on this neural network, the Cauchy noise and posterior squared scale of all diffusion steps are predicted, as well as the true Cauchy noise and posterior squared scale of all diffusion steps calculated in S02. The parameters of the denoising neural network are then optimized using the aforementioned first loss function.

[0031] S04 defines the loss function for Cauchy squared-scale mapping networks for speech synthesis.

[0032] Based on the definition of the Cauchy denoising diffusion probability model in S01 and the calculated Cauchy prior square scale long table, we define the Cauchy prior square scale short table, the Cauchy prior square scale mapping method, and the iterative operation of the Cauchy prior square scale short table, and then define the second loss function of the Cauchy square scale mapping neural network.

[0033] S05, Construct and optimize Cauchy square-scale mapping network for speech synthesis.

[0034] A Cauchy squared scale mapping neural network is constructed. The optimized denoising neural network in S03 is used to calculate Cauchy noise and posterior squared scale as the true values. The parameters of the Cauchy squared scale mapping neural network are optimized using the aforementioned second loss function.

[0035] S06 defines and executes a Cauchy square-scale optimal short table search method for speech synthesis.

[0036] Given an optimized Cauchy denoising diffusion probability model and an optimized Cauchy square-scale mapping neural network, a grid search method is used to search for the best-performing short table of Cauchy prior square-scale maps, employing both random and deterministic sampling methods.

[0037] S07 defines and executes the Cauchy fast sampling method for speech synthesis.

[0038] The Cauchy prior squared scale short table is calculated in an approximate manner, and a single-step sampling operation is defined based on the Cauchy prior squared scale short table and the Cauchy posterior squared scale short table. The single-step sampling operation is executed in an iterative manner to achieve fast speech synthesis.

[0039] In this embodiment of the invention, model training specifically includes the following steps: (1) Audio data processing: The Mel spectrogram of the 80-band frequency range was extracted using the short-time Fourier transform as the conditional input. The parameters of the short-time Fourier transform were set as follows: transform length was set to 1024, jump length was set to 256, window length was set to 1024, minimum frequency was set to 0, and maximum frequency was set to 8000. When training the denoising neural network, 62 frames were randomly selected from the Mel spectrogram as input.

[0040] (2) Constructing a denoising neural network: The denoising neural network in this embodiment consists of a time-series mapping module, a downsampling module, and an upsampling module.

[0041] (2-1) Timing Mapping Module: This module is a neural network containing two SiLU nonlinear layers, each containing 512 neurons.

[0042] (2-2) Downsampling Module: The first component of this module is a convolutional layer with a kernel size of 7 and 32 channels. The second component is a convolutional layer with a kernel size of 3, a spread factor, and a padding size of 1. The third and fourth components are similar to the second component, but with spread factors of 2 and padding sizes of 4, respectively. Skip connections exist between the components of this module, implemented using linear convolutions and nearest neighbor differences.

[0043] (2-3) Upsampling Module: This module consists of three LVC convolutional modules with upsampling rates of 8, 8 and 4 respectively. Each LVC convolutional module contains three LVC convolutional layers with 256 neurons. The kernel size and number of channels of the kernel predictor are 3 and 64 respectively.

[0044] (3) Calculate the first loss function of the Cauchy denoising diffusion probability model for speech synthesis.

[0045] (3-1) Calculate the true Cauchy noise and posterior squared scale for all diffusion steps.

[0046] (3-2) Predict Cauchy noise and posterior squared scale for all diffusion steps.

[0047] (3-3) Calculate the loss function value for a given number of diffusion steps.

[0048] (4) Optimize the parameters of the Cauchy denoising diffusion probability model: The AdamW optimizer is used to optimize the parameters of the denoising neural network. The betas parameter of AdamW is set to (0.9, 0.98), the batch size of the samples is 64, and the learning rate is 0.0002. The optimization process uses parameter regularization and gradient cutoff techniques, with a gradient cutoff value of 1.

[0049] (5) Constructing a Cauchy squared scale mapping network for speech synthesis: The denoising neural network in this embodiment is implemented by the GALR neural network architecture. The GALR network contains two modules. The input and hidden dimensions of the Bi-LSTM in the module are 128, the window size and the segmentation length are 8 and 32, respectively. In addition, the number of attention heads is 8 and the dropout probability is 0.1.

[0050] (6) Calculate the loss function of the Cauchy squared-scale mapping network for speech synthesis.

[0051] (6-1) Calculate Cauchy noise as the true value using the optimized denoising neural network.

[0052] (6-2) Calculate and predict Cauchy noise based on Cauchy square scale mapping network.

[0053] (6-3) Calculate the second loss function value of the Cauchy squared-scale mapping network.

[0054] (7) Optimize the parameters of the Cauchy squared scale mapping network: The AdamW optimizer is used to optimize the parameters of the Cauchy squared scale mapping network. The betas parameter of AdamW is set to (0.9, 0.995), the batch size of the samples is 32, and the learning rate is 0.00001. The optimization process uses parameter regularization and gradient cutoff techniques, with a gradient cutoff value of 15.

[0055] This embodiment compares the fast speech synthesis method of the present invention with other fast speech synthesis methods on multiple datasets, including LJSpeech, VCTK1, and VCTK2. Evaluation metrics include PESQ, STOI, MCD, and RTF. The comparison results are shown in Table 1.

[0056] Table 1 The values ​​of PESQ, STOI, and MCD show that the proposed method for improving the speech synthesis speed using the Cauchy denoising diffusion probability model (Fast Cauchy Diffusion) achieves better overall performance than other speech synthesis methods. The RTF value indicates that although the proposed method does not reach the speed of other fast speech synthesis methods, it easily meets the real-time requirements of speech synthesis. Furthermore, as... Figure 2 As shown, compared with other denoising diffusion probability model speech synthesis methods, the method proposed in this invention can significantly alleviate the problem of speech synthesis quality degradation caused by the number of sampling steps.

[0057] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for improving the speech synthesis speed of a Cauchy denoising diffusion probability model, characterized in that, Includes the following steps: (1) Define a Cauchy denoising diffusion probability model for speech synthesis, including a Cauchy prior square scale long table, a Cauchy posterior square scale long table, a Cauchy single-step diffusion operation, and a Cauchy multi-step diffusion operation. (2) Construct a denoising neural network to predict Cauchy noise and Cauchy posterior squared scale for all diffusion steps; The first loss function is defined based on the Cauchy prior squared scale length table and the Cauchy posterior squared scale length table, and the parameters of the denoising neural network are optimized using the first loss function. (3) Define a Cauchy squared scale mapping neural network, use a denoising neural network to predict Cauchy noise as the true value, and construct a second loss function to optimize the parameters of the Cauchy squared scale mapping neural network; (4) Define the best short table search method of Cauchy square scale. Based on the optimized Cauchy square scale mapping neural network, the random and deterministic iterative sampling methods of the Cauchy denoising diffusion probability model are used, and the grid search method is adopted to search for the best Cauchy prior square scale short table. (5) Define the Cauchy fast sampling method, calculate the Cauchy prior square scale short table corresponding to the Cauchy posterior square scale short table in an approximate manner, define the single-step sampling operation based on the Cauchy prior square scale short table and the Cauchy posterior square scale short table, and execute the single-step sampling operation in an iterative manner to achieve fast speech synthesis.

2. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 1, characterized in that, In step (1), the Cauchy prior square scale length table is defined as follows: ; in, and Let represent the number of diffusion steps of the two Gaussian denoising diffusion probability models in the prior square-scale long table. The prior squared scale value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model on a Cauchy prior square scale long table is... The prior squared scale value at that time; The Cauchy posterior square scale length table is defined as follows: ; in, and Let represent the number of diffusion steps in the posterior squared scale long table for the two Gaussian denoised diffusion probability models. The posterior squared scale value at time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy posterior squared scale table is... The posterior squared scale value at time; The Cauchy single-step diffusion operation is defined as follows: ; in, Indicates the input voice signal; and These represent the number of diffusion steps in the Cauchy prior square-scale long table, respectively. and The voice signal at that time; The Cauchy multi-step diffusion operation is defined as follows: ; ; in, This indicates the number of diffusion steps in the prior square-scale long table. Standard Cauchy noise sampled at time, This represents the square root of the cumulative residual scale when the number of diffusion steps is t.

3. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 1, characterized in that, In step (2), the first loss function is defined based on the Cauchy prior squared scale table and the Cauchy posterior squared scale table, specifically as follows: ; ; ; ; ; in, Indicates the input voice signal; Represents the first loss function; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The loss function value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The noise prediction loss value at that time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The posterior squared scale predicted loss value at that time; This parameter represents the trade-off between the noise prediction loss value and the posterior squared scale prediction loss value. This indicates that the number of diffusion steps of the denoising neural network in the Cauchy prior square-scale long table is . Cauchy noise prediction value at that time; This indicates the number of diffusion steps in the Cauchy prior square-scale long table. The true value of Cauchy noise at that time; This indicates that the number of diffusion steps of the denoising neural network in the Cauchy prior square-scale long table is . Cauchy's posterior squared scale prediction at time; This indicates that the number of diffusion steps in the Cauchy denoising diffusion probability model within the Cauchy prior square scale long table is... The true value of the posterior squared scale at time.

4. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 3, characterized in that, In step (3), the Cauchy squared-scale mapping neural network is defined, including: defining the Cauchy prior squared-scale mapping method, the iterative operation of the Cauchy prior squared-scale short table, and the second loss function of the Cauchy squared-scale mapping neural network; specifically as follows: Define the Cauchy prior squared scale mapping method: ; in, Represents the length of Cauchy's a priori square scale; Represents the short table of Cauchy's a priori square scale; This indicates the number of diffusion steps in the Cauchy prior square-scale long table. The voice signal at that time; This indicates the number of diffusion steps in the short table of Cauchy's prior square scale. The voice signal at that time; Define the iteration operation of the Cauchy prior square scale short table: ; ; in, and represent the prior square scale values ​​in the Cauchy prior square scale short table when the number of diffusion steps is n and n+1, respectively; For the Cauchy squared-scale mapping neural network, the prior squared-scale iterative operation is used during training. This represents a deep neural network that receives a speech signal after a diffusion step of t. As input, predict the ratio of the prior square scale values ​​of two adjacent diffusion steps in the Cauchy prior square scale short table; Define the second loss function for the Cauchy squared-scale mapping neural network: ; ; in, This represents the Cauchy noise prediction value of the optimized denoising neural network when the number of diffusion steps is t in the prior square scale table; This represents the true value of Cauchy noise randomly sampled in a short table with n diffusion steps in the prior square scale.

5. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 1, characterized in that, The specific process of step (4) is as follows: The definition of the Cauchy squared scale mapping iteration is ; ; in, This indicates that the optimized Cauchy squared-scale mapping neural network, given... The ratio of adjacent elements in the short table of Cauchy square scale predicted in real time; Indicates that in a given and At that time, the calculation was based on the optimized Cauchy squared scale mapping neural network. The formula; Indicates that in a given and Time calculation The formula; Based on the aforementioned calculations and The definition of random and deterministic iterative sampling methods ; in, and These represent the speech signals with diffusion steps of n and n+1 in the short table of Cauchy's a prior square scale, respectively; and These represent the Cauchy noise value and Cauchy squared scale value predicted by the optimized Cauchy denoising diffusion probability model, respectively. and At that time, deterministic sampling and random sampling were used respectively; Given initial value and Through the aforementioned iterative calculation , and To achieve speech synthesis; divided into grids. and The range of values ​​is used to perform speech synthesis and evaluate the quality of the synthesized speech, thereby achieving the search for the optimal Cauchy prior square scale short table.

6. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 1, characterized in that, In step (5), the Cauchy a prior square scale short table corresponding to the Cauchy posterior square scale short table is calculated in an approximate manner, using the following formula: ; in, and The optimal Cauchy square scale short table search method is given; This represents the approximate Cauchy posterior squared scale short table.

7. The method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model according to claim 1, characterized in that, In step (5), a single-step sampling operation is defined based on the Cauchy prior square scale short table and the Cauchy posterior square scale short table, specifically as follows: ; in, and These represent the speech signals with diffusion steps of n-1 and n in the short table of Cauchy's a prior square scale, respectively; This represents the noise prediction value of the optimized Cauchy denoising diffusion probability model; This represents the true value of Cauchy noise randomly sampled in the prior square-scale short table when the number of diffusion steps is n; using the Mel spectrogram as a conditional input, single-step sampling operations are continuously executed to achieve fast speech synthesis.

8. An apparatus for improving the speech synthesis speed of a Cauchy denoising diffusion probability model, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method for improving the speech synthesis speed of the Cauchy denoising diffusion probability model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice synthesis method and device based on Cauchy denoising probability diffusion model

    CN119049446A

  • System and method for improving speech synthesis speed of diffusion model

    CN120877701A

  • Speech synthesis method, diffusion model training method, device and equipment

    CN120998174A