Acoustic model post-processing method, server and readable memory based on probability diffusion model
By using the acoustic model post-processing method of probability diffusion model in speech synthesis technology, the predicted spectrum is optimized, which solves the problem of low acoustic spectrum quality in the prior art, and improves the naturalness and spectrum details of speech synthesis.
Patent Information
- Application Number
- CN202111652872.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-30
AI Technical Summary
In the existing speech synthesis technology, the acoustic spectrum quality is not high, resulting in the low naturalness of the synthesized speech waveform file.
The acoustic model post-processing method based on the probability diffusion model is used to optimize the predicted spectrum in detail, and through the model training and inference process, an optimized spectrum is closer to the real spectrum.
The natural performance of synthetic speech is improved, the generated speech waveform files are of higher quality and richer spectrum details.
Smart Images

Figure CN114512114B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology. Aiming at the problem that the quality of speech synthesis needs to be improved, a method for post-processing an acoustic model based on a probability diffusion model is proposed. The method improves the naturalness of synthesized speech by optimizing the details of the predicted spectrum. Background Art
[0002] With the continuous development of artificial intelligence technology, great progress has been made in various fields based on artificial intelligence technology. Speech synthesis is an important direction in artificial intelligence technology. Speech synthesis specifically refers to the synthesis from text to audio. The goal of speech synthesis technology is to enable machines to speak like humans. High-quality speech synthesis technology can meet the needs of human-computer interaction in today's society.
[0003] Existing speech synthesis technology first converts text into linguistic features, combines them with speaker features for feature fusion, and then uses deep learning to encode and decode them to obtain acoustic spectra or directly obtain synthesized speech waveform files. However, the quality of the obtained acoustic spectra is not high enough, so the synthesized speech waveform files will also have a lower degree of naturalness. Summary of the invention
[0004] Based on this, the present invention aims at the shortcomings of existing speech synthesis technology and proposes an acoustic model post-processing method based on a probability diffusion model. This method can process the acoustic model after obtaining a predicted spectrum. The result obtained after the processing has richer spectral details and is closer to the real spectrum, thereby improving the naturalness of the synthesized speech waveform file.
[0005] To achieve the above-mentioned object of the invention, the present invention adopts the following technical scheme, a method for post-processing an acoustic model based on a probability diffusion model, comprising the following steps:
[0006] Step 1: Model training: Use the server to train the probability diffusion model, optimize the parameters of the probability diffusion model by reducing the loss function until the model converges, and obtain the weight of the probability diffusion model; including the following sub-steps:
[0007] S11. Using a server to perform feature extraction on the text and the corresponding audio in a specific English data set, specifically: extracting linguistic features from the text, extracting the corresponding pitch and intensity from the audio, and using specific acoustic parameters to extract a Mel spectrum from the audio as a real spectrum, and extracting a phoneme sequence in combination with the audio and linguistic features, and then combining the Mel spectrum and the phoneme sequence to obtain duration information using an alignment search method;
[0008] S12, using the server to model an acoustic model for the phoneme sequence, duration information, pitch and intensity, and Mel spectrum obtained in step S11; the modeling process of the acoustic model includes: firstly encoding the phoneme sequence, then aligning the phoneme sequence to the length of the Mel spectrum using the duration information, then adding corresponding pitch and intensity features, then decoding the sequence, and finally obtaining a generated predicted spectrum;
[0009] S13. Use the server to calculate the probability diffusion model for the extracted real spectrum. Specifically, by designing a specific noise coefficient that varies with the number of diffusion steps, the data distribution of the conditional probability corresponding to the noise coefficient under different diffusion steps is calculated. The data distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as After the probability transfer of all the set diffusion steps, the latent vector sampling points of the diffusion space are obtained;
[0010] S14, using the server to train the probability diffusion model, specifically: calculating the loss function L according to the diffusion space latent vector sampling points calculated in step S13 and the predicted spectrum generated in step S12 m , and the probability distribution corresponding to the probability diffusion model under different diffusion steps is calculated through the noise estimation network in the probability diffusion model. The probability distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as According to the mean value obtained in step S13 and variance Calculate the loss function L n , the loss function L m and L n According to the preset weight λ m and λ n Weighted, unified training is performed on the server, and the weighted loss function calculation formula is L = L m λ m +L n λ n ;
[0011] S15. Use the server to calculate the gradient according to the loss function L. Specifically, the gradient refers to the partial derivative of the loss function expression with respect to the learnable parameters in the model. Then, the parameters of the probability diffusion model are updated by back-propagation using the Adam optimizer, so that the loss function L is reduced until convergence. At this point, the training is completed, and the weights of the trained probability diffusion model are obtained.
[0012] Step 2: Model inference: Based on the model weights obtained in the training phase, the server uses the input predicted spectrum to optimize the spectrum, which includes the following sub-steps:
[0013] S21. Take the predicted spectrum as the mean of the probability diffusion space, and the variance of the multivariate random Gaussian distribution as the variance to obtain the sampling points of the diffusion space, and then perform reverse probability diffusion on the sampling points. Specifically, for each diffusion step, use the trained noise estimation network weights to obtain the mean of the distribution under the current step number t. and variance Then the probability distribution is transferred step by step, and finally the optimized prediction spectrum is obtained after the preset number of steps are completed;
[0014] S22. Using the server to reconstruct the waveform of the optimized predicted spectrum to obtain a speech waveform file.
[0015] Furthermore, the acoustic parameters taken in step S11 are as follows: the minimum frequency and the maximum frequency are 0 and 8000 Hz respectively, the sampling rate is 22050 Hz, the number of points calculated by FFT is 1025, and the extracted Mel spectrum is set to 80 dimensions.
[0016] The alignment search method is designed to take three phonemes as a unit, the next unit contains the last two phonemes of the previous unit, calculate the correlation coefficient between this search unit and each frame of the Mel spectrum, and use a dynamic programming algorithm to set the item with the largest correlation coefficient to 1. The final duration information obtained is a natural number sequence of the same length as the phoneme sequence, and the sum of the natural number sequence matches the corresponding Mel spectrum length.
[0017] Furthermore, in step S12, the phoneme sequence is encoded using a multi-head attention mechanism and a one-dimensional convolution operation as an encoding layer, with a total of 4 layers of encoding structure design;
[0018] The method of using the duration information to align the phoneme sequence to the length of the Mel spectrum includes: the duration information represents the utterance duration corresponding to a certain vector in the encoded latent sequence, the corresponding vector is repeated a certain number of times according to the duration, and then the same operation is performed on all vectors in the sequence, and they are connected in series in the time dimension to obtain a latent sequence with the same length as the Mel spectrum;
[0019] The sequence decoding is also based on a decoder structure combining a 4-layer multi-head attention mechanism and a one-dimensional convolution, and after decoding, a linear layer is used to map the dimension of the latent sequence to the dimension of the Mel spectrum to generate a predicted spectrum.
[0020] Furthermore, the diffusion step number in step S13 is 1000 steps, and the noise coefficient β that varies with the diffusion step number t is t The calculation formula is β t =β0+(β1-β0)t, where β0 and β1 are 0.05 and 20 respectively.
[0021] The mean and variance calculation formulas of the data distribution of the conditional probability are:
[0022]
[0023] Where x t-1 is the hidden vector of the previous step of the current diffusion step, and I represents the identity matrix.
[0024] Furthermore, the noise estimation network described in step S14 uses two 4-layer two-dimensional convolution structures. In the first 4-layer two-dimensional convolution structure, maximum pooling-based downsampling is used between each layer to downsample the hidden sequence by one time; in the second 4-layer two-dimensional convolution structure, nearest neighbor interpolation-based upsampling is used between each layer to upsample the hidden sequence by one time, and in the upsampling calculation of each layer, the current result and the downsampled hidden sequence of the corresponding resolution are connected in parallel in the time dimension, and finally a linear layer is used to map the output dimension to the dimension of the Mel spectrum.
[0025] The weight λ m and λ n The values are 1 and 0.4 respectively.
[0026] Furthermore, the number of training steps until convergence in step S15 is 1000.
[0027] Furthermore, in step S21, when using the trained model for reverse probability diffusion, the preset number of steps is 100, and the obtained optimized predicted spectrum is closer to the real spectrum, so the quality of the acoustic model synthesized spectrum is further improved through post-processing.
[0028] Furthermore, the waveform reconstruction described in step S22 is specifically as follows: first initialize a phase spectrum, obtain a waveform file through inverse short-time Fourier transform (ISTFT) according to the optimized predicted spectrum and the initialized phase spectrum, and then calculate the short-time Fourier transform (STFT) of the waveform file to obtain a new spectrum and phase spectrum, then discard the new spectrum, and use the optimized predicted spectrum and the new phase spectrum to synthesize the waveform again, and repeat this process many times until a more natural speech waveform file is obtained, the sampling rate of the obtained speech waveform file is 22050Hz, and the output format is wav.
[0029] The present invention further provides a server, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0030] The present invention also provides a readable memory having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0031] The advantages and beneficial effects of the present invention are as follows:
[0032] (1) Through the design based on the probability diffusion model, the model learns the difference between the predicted spectrum and the real spectrum, so that it can synthesize an optimized spectrum with higher quality;
[0033] (2) Through reasonable model and parameter design and model training, the synthesized optimized spectrum can produce a more natural speech result. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow chart of the steps of the method of the present invention;
[0035] Figure 2 For the implementation effect of the present invention, the upper and lower figures are respectively the predicted spectrum and the optimized spectrum;
[0036] Figure 3 FIG. 1 is a schematic diagram of the internal structure of a server in an embodiment. DETAILED DESCRIPTION
[0037] In order to better explain the method of the present invention, the technical solution of the present invention is further described below by specific embodiments in conjunction with the accompanying drawings, so that the method is clearer. Through the content disclosed in this specification, those skilled in the art can easily understand the effect of the present invention. The present invention can also be implemented or applied by other specific implementation methods, and the details of this specification can also be modified based on different viewpoints, and various modifications or changes are made without departing from the essence of the present invention. It should be noted that, in the absence of conflict, the following embodiments and the features therein can be further combined.
[0038] In detail, the present invention proposes an acoustic model post-processing method based on a probability diffusion model, including the following specific implementation contents:
[0039] Step 1: Model training: Use the server to train the probability diffusion model, optimize the parameters of the probability diffusion model by reducing the loss function until the model converges, and obtain the weight of the probability diffusion model; including the following sub-steps:
[0040] S11. Using a server to perform feature extraction on the text and the corresponding audio in a specific English data set, specifically: extracting linguistic features from the text, extracting the corresponding pitch and intensity from the audio, and using specific acoustic parameters to extract a Mel spectrum from the audio as a real spectrum, and extracting a phoneme sequence in combination with the audio and linguistic features, and then combining the Mel spectrum and the phoneme sequence to obtain duration information using an alignment search method;
[0041] S12, using the server to model an acoustic model for the phoneme sequence, duration information, pitch and intensity, and Mel spectrum obtained in step S11; the modeling process of the acoustic model includes: firstly encoding the phoneme sequence, then aligning the phoneme sequence to the length of the Mel spectrum using the duration information, then adding corresponding pitch and intensity features, then decoding the sequence, and finally obtaining a generated predicted spectrum;
[0042] S13. Use the server to calculate the probability diffusion model for the extracted real spectrum. Specifically, by designing a specific noise coefficient that varies with the number of diffusion steps, the data distribution of the conditional probability corresponding to the noise coefficient under different diffusion steps is calculated. The data distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as After the probability transfer of all the set diffusion steps, the latent vector sampling points of the diffusion space are obtained;
[0043] S14, using the server to train the probability diffusion model, specifically: calculating the loss function L according to the diffusion space latent vector sampling points calculated in step S13 and the predicted spectrum generated in step S12 m , and the probability distribution corresponding to the probability diffusion model under different diffusion steps is calculated through the noise estimation network in the probability diffusion model. The probability distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as According to the mean value obtained in step S13 and variance Calculate the loss function L n , the loss function L m and L n According to the preset weight λ m and λ n Weighted, unified training is performed on the server, and the weighted loss function calculation formula is L = L m λ m +L n λ n ;
[0044] S15. Use the server to calculate the gradient according to the loss function L. Specifically, the gradient refers to the partial derivative of the loss function expression with respect to the learnable parameters in the model. Then, the parameters of the probability diffusion model are updated by back-propagation using the Adam optimizer, so that the loss function L is reduced until convergence. At this point, the training is completed, and the weights of the trained probability diffusion model are obtained.
[0045] Step 2: Model inference: Based on the model weights obtained in the training phase, the server uses the input predicted spectrum to optimize the spectrum, which includes the following sub-steps:
[0046] S21. Take the predicted spectrum as the mean of the probability diffusion space, and the variance of the multivariate random Gaussian distribution as the variance to obtain the sampling points of the diffusion space, and then perform reverse probability diffusion on the sampling points. Specifically, for each diffusion step, use the trained noise estimation network weights to obtain the mean of the distribution under the current step number t. and variance Then the probability distribution is transferred step by step, and finally the optimized prediction spectrum is obtained after the preset number of steps are completed;
[0047] S22. Using the server to reconstruct the waveform of the optimized predicted spectrum to obtain a speech waveform file.
[0048] The acoustic parameters taken in step S11 are as follows: the minimum frequency and the maximum frequency are 0 and 8000 Hz respectively, the sampling rate is 22050 Hz, the number of points calculated by FFT is 1025, and the extracted Mel spectrum is set to 80 dimensions.
[0049] The specific English dataset is a public academic dataset in the field of speech, containing 13,100 single-speaker audio clips of reading passages from 7 non-fiction books. The length of the clips ranges from 1 second to 10 seconds, with a total length of about 24 hours.
[0050] The alignment search method is designed to take three phonemes as a unit, with the next unit containing the last two phonemes of the previous unit. The correlation coefficient between this search unit and each frame of the Mel spectrum is calculated, and the item with the largest correlation coefficient is set to 1 using a dynamic programming algorithm. The final duration information obtained is a natural number sequence of the same length as the phoneme sequence, and the sum of the natural number sequence matches the corresponding Mel spectrum length.
[0051] The calculation of the correlation coefficient is specifically to calculate the Euclidean distances between the three phonemes on the search unit and the current frame of Mel spectrum, and then add the three calculated Euclidean distances to obtain the correlation coefficient between the search unit and the frame of Mel spectrum.
[0052] The phoneme sequence encoding in step S12 uses a multi-head attention mechanism and a one-dimensional convolution operation as an encoding layer, and a total of 4 layers of encoding structure design are adopted;
[0053] The method of using the duration information to align the phoneme sequence to the length of the Mel spectrum includes: the duration information represents the utterance duration corresponding to a certain vector in the encoded latent sequence, the corresponding vector is repeated a certain number of times according to the duration, and then the same operation is performed on all vectors in the sequence, and they are connected in series in the time dimension to obtain a latent sequence with the same length as the Mel spectrum;
[0054] The sequence decoding is also based on a decoder structure combining a 4-layer multi-head attention mechanism and a one-dimensional convolution, and after decoding, a linear layer is used to map the dimension of the latent sequence to the dimension of the Mel spectrum to generate a predicted spectrum.
[0055] The number of diffusion steps in step S13 is 1000 steps, and the noise coefficient β that varies with the number of diffusion steps t is t The calculation formula is β t =β0+(β1-β0)t, where β0 and β1 are 0.05 and 20 respectively.
[0056] The mean and variance calculation formulas of the data distribution of the conditional probability are:
[0057]
[0058] Where x t-1 is the hidden vector of the previous step of the current diffusion step, and I represents the identity matrix.
[0059] The noise estimation network described in step S14 uses two 4-layer two-dimensional convolution structures. In the first 4-layer two-dimensional convolution structure, maximum pooling-based downsampling is used between each layer to downsample the hidden sequence by one time; in the second 4-layer two-dimensional convolution structure, nearest neighbor interpolation-based upsampling is used between each layer to upsample the hidden sequence by one time, and when calculating the upsampling of each layer, the current result and the downsampled hidden sequence of the corresponding resolution are connected in parallel in the time dimension, and finally a linear layer is used to map the output dimension to the dimension of the Mel spectrum.
[0060] The noise estimation network is described in detail: Assume that the input sequence length is N, and the input hidden sequence is represented as x = (x1, x2, ..., x N ), then in the first 4-layer 2D convolution structure, the output of each downsampled layer is represented as x 1 =(x1,x2,…,x N / 2 ), x 2 =(x1,x2,…,x N / 4 ), x 3 =(x1,x2,…,x N / 8 ) and x 4 =(x1,x2,…,x N / 16 ), where the length of each downsampled hidden sequence will be reduced to the original That is, the corresponding lengths are and In the second 4-layer 2D convolution structure, the first up-sampled input is y 4 =(x1,x2,…,x N / 16 ), the output is y 3 =(x1,x2,…,x N / 8), then each layer of upsampling calculation first converts y i and x i Connect in parallel in the time dimension, and then upsample the output to y i-1 , after 4 upsampling, the output is y 1 , and then pass the linear layer to get the output Mel spectrum.
[0061] The weight λ m and λ n The values are 1 and 0.4 respectively.
[0062] The number of training steps until convergence in step S15 is 1000.
[0063] In step S21, when using the trained model for reverse probability diffusion, the preset number of steps is 100, and the obtained optimized predicted spectrum is closer to the real spectrum. Therefore, the quality of the synthesized spectrum of the acoustic model is further improved through post-processing.
[0064] The waveform reconstruction described in step S22 is as follows: first, a phase spectrum is initialized, and a waveform file is obtained by inverse short-time Fourier transform (ISTFT) according to the optimized predicted spectrum and the initialized phase spectrum, and then the short-time Fourier transform (STFT) is calculated on the waveform file to obtain a new spectrum and phase spectrum, and then the new spectrum is discarded, and the waveform is synthesized again using the optimized predicted spectrum and the new phase spectrum, and this process is repeated many times until a more natural speech waveform file is obtained. The sampling rate of the obtained speech waveform file is 22050Hz, and the output format is wav.
[0065] like Figure 1 As shown, the input text is first converted into text features, namely the phoneme sequence, and also includes pitch and intensity features extracted based on text and audio. The text features are then input into the encoder to obtain an encoded latent sequence, where the length of the latent sequence is consistent with the length of the text encoding. Next, the latent sequence obtained after text encoding is aligned through a bypass, that is, the duration information is extracted, that is, a specific vector is repeated the number of times according to the duration, and the latent vector aligned with the spectrum sequence is input into the decoder to obtain a predicted spectrum generated after decoding.
[0066] like Figure 2 As shown in the figure, the predicted spectrum is relatively smooth. Although it has most of the information of the synthesized speech, it lacks some fine-grained spectral features, which will reduce the naturalness performance in the synthesized speech.
[0067] like Figure 1As shown in the figure, the predicted spectrum is used as the mean of the diffusion space, the sampling points are obtained through the specific data distribution in the diffusion space, and then a total of 100 steps of stepwise reverse probability diffusion calculation are performed. In the probability diffusion, the trained noise estimation network is used to calculate the transfer probability of each step distribution, and finally the synthesized optimized spectrum is obtained.
[0068] like Figure 2 As shown in the figure below, the optimized spectrum is richer in details than the predicted spectrum. The spectrum details make the synthesized speech more natural.
[0069] Figure 3 FIG. 1 is a schematic diagram of the internal structure of a server in an embodiment. Figure 3 As shown, the server includes a processor and a memory connected via a system bus. The processor is used to provide computing and control capabilities to support the operation of the entire server. The memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The computer program can be executed by the processor to implement an acoustic model post-processing method based on a probability diffusion model provided in the following embodiments. The internal memory provides a cached operating environment for the operating system computer program in the non-volatile storage medium. The server can be a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc.
[0070] The embodiment of the present invention further provides a computer-readable storage medium, one or more non-volatile computer-readable storage media containing computer-executable instructions, which, when executed by one or more processors, enable the processors to perform the steps of the acoustic model post-processing method based on the probability diffusion model.
[0071] A computer program product comprising instructions, when running on a computer, causes the computer to perform an acoustic model post-processing method based on a probability diffusion model.
[0072] Any reference to memory, storage, database or other medium used in embodiments of the present invention may include non-volatile and / or volatile memory. Suitable non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM), which is used as an external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0073] The above embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the present application. For those skilled in the art, the present invention can have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An acoustic model post-processing method based on a probability diffusion model, characterized in that: The following steps are involved: Step 1: Model training: Use the server to train the probability diffusion model, optimize the parameters of the probability diffusion model by reducing the loss function until the model converges, and obtain the weight of the probability diffusion model; It includes the following sub-steps: S11. Using a server to perform feature extraction on the text and the corresponding audio in a specific English data set, specifically: extracting linguistic features from the text, extracting the corresponding pitch and intensity from the audio, and using specific acoustic parameters to extract a Mel spectrum from the audio as a real spectrum, and extracting a phoneme sequence in combination with the audio and linguistic features, and then combining the Mel spectrum and the phoneme sequence to obtain duration information using an alignment search method; S12, using the server to model an acoustic model for the phoneme sequence, duration information, pitch and intensity, and Mel spectrum obtained in step S11; the modeling process of the acoustic model includes: firstly encoding the phoneme sequence, then aligning the phoneme sequence to the length of the Mel spectrum using the duration information, then adding corresponding pitch and intensity features, then decoding the sequence, and finally obtaining a generated predicted spectrum; S13. Use the server to calculate the probability diffusion model for the extracted real spectrum. Specifically, by designing a specific noise coefficient that varies with the number of diffusion steps, the data distribution of the conditional probability corresponding to the noise coefficient under different diffusion steps is calculated. The data distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as After the probability transfer of all the set diffusion steps, the latent vector sampling points of the diffusion space are obtained; S14, using the server to train the probability diffusion model, specifically: calculating the loss function L according to the diffusion space latent vector sampling points calculated in step S13 and the predicted spectrum generated in step S12 m , and the probability distribution corresponding to the probability diffusion model under different diffusion steps is calculated through the noise estimation network in the probability diffusion model. The probability distribution conforms to the multivariate random Gaussian distribution, and the mean is recorded as Variance is recorded as According to the mean value obtained in step S13 and variance Calculate the loss function L n , the loss function L m and L n According to the preset weight λ m and λ n Weighted, unified training is performed on the server, and the weighted loss function calculation formula is L = L m λ m +L n λ n ; S15. Using the server to calculate the gradient according to the loss function L, specifically: using the Adam optimizer to back-propagate and update the parameters of the probability diffusion model, so that the loss function L is reduced until convergence, the training is completed, and the weight of the trained probability diffusion model is obtained; Step 2: Model inference: Based on the model weights obtained in the training phase, the server uses the input predicted spectrum to optimize the spectrum. It includes the following sub-steps: S21. Take the predicted spectrum as the mean of the probability diffusion space, and the variance of the multivariate random Gaussian distribution as the variance to obtain the sampling points of the diffusion space, and then perform reverse probability diffusion on the sampling points. Specifically, for each diffusion step, use the trained noise estimation network weights to obtain the mean of the distribution under the current step number t. and variance Then the probability distribution is transferred step by step, and finally the optimized prediction spectrum is obtained after the preset number of steps are completed; S22. Using the server to reconstruct the waveform of the optimized predicted spectrum to obtain a speech waveform file.
2. The acoustic model post-processing method based on the probability diffusion model according to claim 1, characterized in that: The acoustic parameters taken in step S11 are as follows: the minimum frequency and the maximum frequency are 0 and 8000 Hz respectively, the sampling rate is 22050 Hz, the number of points calculated by FFT is 1025, and the Mel spectrum extraction is set to 80 dimensions; The alignment search method is designed to take three phonemes as a unit, the next unit contains the last two phonemes of the previous unit, calculate the correlation coefficient between this search unit and each frame of the Mel spectrum, and use a dynamic programming algorithm to set the item with the largest correlation coefficient to 1. The final duration information obtained is a natural number sequence of the same length as the phoneme sequence, and the sum of the natural number sequence matches the corresponding Mel spectrum length.
3. The acoustic model post-processing method based on the probability diffusion model according to claim 1, characterized in that: The phoneme sequence encoding in step S12 uses a multi-head attention mechanism and a one-dimensional convolution operation as an encoding layer, and a total of 4 layers of encoding structure design are adopted; The method of using the duration information to align the phoneme sequence to the length of the Mel spectrum includes: the duration information represents the utterance duration corresponding to a certain vector in the encoded latent sequence, the corresponding vector is repeated a certain number of times according to the duration, and then the same operation is performed on all vectors in the sequence, and they are connected in series in the time dimension to obtain a latent sequence with the same length as the Mel spectrum; The sequence decoding is also based on a decoder structure combining a 4-layer multi-head attention mechanism and a one-dimensional convolution, and after decoding, a linear layer is used to map the dimension of the latent sequence to the dimension of the Mel spectrum to generate a predicted spectrum.
4. The acoustic model post-processing method based on the probability diffusion model according to claim 1, characterized in that: The diffusion step number in step S13 is 1000 steps, and the noise coefficient β that varies with the diffusion step number t is t The calculation formula is β t =β0+(β1-β0)t, where β0 and β1 are 0.05 and 20 respectively; The mean and variance calculation formulas of the data distribution of the conditional probability are: Where x t-1 is the hidden vector of the previous step of the current diffusion step, and I represents the identity matrix.
5. The acoustic model post-processing method based on the probability diffusion model according to claim 1, characterized in that: The noise estimation network described in step S14 uses two 4-layer two-dimensional convolution structures. In the first 4-layer two-dimensional convolution structure, maximum pooling-based downsampling is used between each layer to downsample the hidden sequence by one time; in the second 4-layer two-dimensional convolution structure, nearest neighbor interpolation-based upsampling is used between each layer to upsample the hidden sequence by one time, and when calculating the upsampling of each layer, the current result and the downsampled hidden sequence of the corresponding resolution are connected in parallel in the time dimension, and finally a linear layer is used to map the output dimension to the dimension of the Mel spectrum.
6. The acoustic model post-processing method based on probability diffusion model according to claim 1, characterized in that: The number of training steps until convergence in step S15 is 1000.
7. The acoustic model post-processing method based on the probability diffusion model according to claim 1, characterized in that: In step S21, when using the trained model for reverse probability diffusion, the preset number of steps is 100, and the obtained optimized predicted spectrum is closer to the real spectrum. Therefore, the quality of the synthesized spectrum of the acoustic model is further improved through post-processing.
8. The acoustic model post-processing method based on probability diffusion model according to claim 1, characterized in that: The waveform reconstruction described in step S22 is specifically as follows: first, a phase spectrum is initialized, and a waveform file is obtained by inverse short-time Fourier transform based on the optimized predicted spectrum and the initialized phase spectrum, and then the short-time Fourier transform is calculated on the waveform file to obtain a new spectrum and phase spectrum, and then the new spectrum is discarded, and the waveform is synthesized again using the optimized predicted spectrum and the new phase spectrum, and this process is repeated many times until a more natural speech waveform file is obtained.
9. A server comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the computer program is executed by the processor, the processor is caused to perform the steps of the acoustic model post-processing method based on the probability diffusion model according to any one of claims 1 to 8.
10. A readable memory having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the acoustic model post-processing method based on the probability diffusion model as claimed in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Singing synthesis method and device, computer equipment and storage medium
CN113421544A
Voice transcription recognition training decoding method based on fast jump decoding and system
CN113488028A