Speech signal processing method and related device

The far-field speech signal is processed through a two-step diffusion model, which first reduces noise and then repairs the high-frequency components. This solves the problem of insufficient clarity and fullness of far-field speech signals, improves the listening experience, and reduces algorithm complexity and device power consumption.

WO2025200819A1PCT designated stage Publication Date: 2025-10-02HONOR DEVICE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/077005
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-02-12
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

When collecting voice signals from a long distance or in the far field, the clarity and fullness of the voice signals are insufficient, especially the high-frequency components are drowned out by noise, making it difficult for existing technologies to effectively improve the listening experience.

Method used

A two-step processing method is adopted: first, noise reduction is performed through the first diffusion model, and then the high-frequency components are repaired using the second diffusion model. The speech signal before noise reduction is used as the condition of the conditional diffusion model to reduce interference and improve the repair effect.

Benefits of technology

It significantly improves the clarity and fullness of far-field voice signals, reduces algorithm complexity and power consumption of terminal devices, and has better processing effects than a single model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025077005_02102025_PF_FP_ABST
    Figure CN2025077005_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applied to the field of audio. Provided are a speech signal processing method and a related device. In the method, a speech signal to be processed (in particular a far-field speech signal) is subjected to noise reduction on the basis of a diffusion model, and a high-frequency component is then recovered on the basis of a diffusion model, so as to obtain a target speech signal. By means of noise reduction, the target speech signal can be made perceptually clearer; and by means of recovery, the target speech signal can be made perceptually fuller. Compared with noise reduction processing alone, the target speech signal is perceptually fuller. While recovery processing is performed, a speech signal before noise reduction is performed is used as a condition of a conditional diffusion model, i.e., as reference information for recovering the high-frequency component, such that the interference with recovery processing that may be caused by noise reduction processing can be reduced, thereby improving a recovery effect. The far-field speech signal is processed respectively using two diffusion models. Compared with performing processing using one model, in the case that the same effect is achieved, the complexity of an algorithm can be reduced, and the computing power can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

A voice signal processing method and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 26, 2024, with application number 202410350758.1 and application name “A Speech Signal Processing Method and Related Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to a voice signal processing method and related equipment. Background Art

[0003] Using a terminal device to collect voice signals from a distance (also known as collecting or picking up far-field voice signals) is extremely challenging. Long distance or far field refers to the distance between the microphone used to collect or pick up sound and the sound source. Generally speaking, in the field of acoustics, a distance of more than 10 meters between the microphone and the sound source is considered long distance, and a distance of 20 or 30 meters is considered ultra-long distance.

[0004] At present, far-field speech signals are mainly picked up by microphone arrays, and the picked up far-field speech signals are enhanced according to acoustic models to pick up speech signals in the target direction. Summary of the Invention

[0005] The present application provides a speech signal processing method and related equipment, which can make the collected speech signals (especially far-field speech signals) sound clearer and fuller.

[0006] In a first aspect, a speech signal processing method is provided, which is applied to an electronic device, and the method includes: obtaining a first speech signal to be processed; performing noise reduction processing on the first speech signal based on a first diffusion model to obtain a second speech signal; and performing repair processing on the first speech signal and the second speech signal based on a second diffusion model to obtain a target speech signal, wherein the second diffusion model is a conditional diffusion model, and the repair processing includes: using the first speech signal as a condition for the second diffusion model to repair the high-frequency components of the second speech signal.

[0007] This solution not only reduces noise, making the target speech signal clearer, but also restores high-frequency components, making the target speech signal fuller. Furthermore, compared to solutions that only reduce noise on far-field speech signals, this solution can achieve a fuller sound.

[0008] Furthermore, since the attenuated high-frequency components may be submerged in noise, the noise reduction process may interfere with the attenuated high-frequency components of the first speech signal, potentially affecting the restoration effect. During the restoration process, the pre-noise-reduction speech signal is also used as a condition in the conditional diffusion model and serves as reference information for restoring the high-frequency components. This reduces the potential interference of the noise reduction process on the restoration process and improves the restoration effect.

[0009] Furthermore, by breaking down the processing of the first voice signal into two steps, each using two diffusion models, this reduces algorithm complexity and saves computing power while achieving the same effect as using a single model to process the far-field voice signal. In other words, using a single model to process the first voice signal requires a more complex algorithm to achieve the same effect as both models, placing higher demands on the power consumption and computing power of the terminal device.

[0010] In a possible embodiment, the first voice signal is a far-field voice signal, or the distance between a microphone of the electronic device and a sound source of the first voice signal is greater than or equal to a preset distance threshold.

[0011] The embodiments of the present application are particularly effective in capturing far-field voice signals and performing voice processing.

[0012] In a possible embodiment, noise reduction processing is performed on a first speech signal based on a first diffusion model to obtain a second speech signal, including: calculating the gradient of the first speech signal, where the gradient is used to characterize the probability distribution of the second speech signal, and the probability distribution of the second speech signal corresponds to the distribution of time-frequency points in a spectrogram of the second speech signal; and sampling the gradient of the first speech signal to obtain the second speech signal.

[0013] It can be understood that the above consideration of the probability distribution of the second speech signal is to consider the integrity of the distribution of the time-frequency points of the second speech signal, and the integrity includes the relationship between the time-frequency points.

[0014] It can also be understood that in the case of not considering the overall nature and analyzing each time-frequency point one by one, even if the total error (for example, the first error value) between the time-frequency points of the signal after noise reduction processing and the time-frequency points corresponding to the ideal clean signal (the signal after all noise is removed from the first speech signal in an ideal situation) is very small, it is very likely that a small number of time-frequency points will have large errors with the corresponding time-frequency points of the ideal clean signal, while most time-frequency points will have small errors. Therefore, the small number of time-frequency points with large errors will, in principle, cause a certain degree of damage to the noise-reduced speech, poor spectral continuity, and poor listening experience.

[0015] Compared with the solution that does not consider the overall distribution of time-frequency points, this application, on the basis of considering the overall distribution of time-frequency points, if the total error is controlled to the same small value (such as the first error value), the errors of almost all points will be more balanced and smaller, thereby reducing the damage of noise reduction processing to the speech signal in principle.

[0016] In one possible embodiment, calculating the gradient of a first speech signal includes: inputting the first speech signal into a first neural network to obtain a denoised gradient to be sampled of the first speech signal, where the denoised gradient to be sampled is used to characterize the probability distribution of a second speech signal; sampling the gradient of the first speech signal includes: predicting a denoised sampled signal based on the denoised gradient to be sampled; calculating the gradient of the first speech signal also includes: inputting the denoised sampled signal into the first neural network to obtain a denoised gradient to be corrected of the denoised sampled signal, where the denoised gradient to be corrected is used to characterize the probability distribution of the second speech signal; sampling the gradient of the first speech signal also includes: correcting the denoised sampled signal based on the denoised gradient to be corrected to obtain a denoised corrected signal; and generating a second speech signal based on the denoised corrected signal.

[0017] The above scheme further calculates the denoised gradient to be corrected based on the denoised sampled signal obtained based on the prediction of the denoised gradient to be sampled, and corrects the denoised sampled signal based on the denoised gradient to be corrected. Compared with the diffusion model that only calculates the gradient once, or compared with the diffusion model that only predicts without correction, the denoising effect is better.

[0018] In other words, sampling based on the gradient includes two sub-steps: prediction and correction. For example, the prediction step uses ancestral sampling, while the correction step uses Langevin kinetic sampling or annealed Langevin kinetic sampling. The correction step is used to correct the prediction results of the prediction step. The prediction and sampling steps in this application can be referred to the examples herein and are described here for a unified explanation.

[0019] In a possible embodiment, the denoised sampling signal is predicted based on the denoised gradient to be sampled, including: calculating the denoised drift coefficient based on the first speech signal based on the stochastic differential equation of the first diffusion model; calculating the denoised inverse drift coefficient based on the denoised drift coefficient and the denoised gradient to be sampled; and predicting the denoised sampling signal based on the denoised inverse drift coefficient, the first speech signal and a fourth Gaussian noise value, wherein the fourth Gaussian noise value is generated based on a third random seed.

[0020] It can be understood that the inverse drift coefficient can be used to describe the path that a noisy speech signal follows during sampling, toward a clean speech signal. Sampling based on the inverse drift coefficient to obtain a gradient can achieve noise reduction.

[0021] In a possible embodiment, the first neural network is obtained based on first input data and first target data, wherein the first input data includes a sample noisy speech signal and a sample noise reduction sampling signal, the sample noisy speech signal is generated by the first sample noise signal and the first sample speech signal, the sample noise reduction sampling signal is generated based on the sample noisy speech signal, and the first target data is used to characterize the probability distribution of the first sample speech signal.

[0022] The above scheme trains the neural network model by using target data to represent the probability distribution of the sample speech signal contained in the sample noisy speech signal, so that the gradient output by the trained neural network model can represent the probability distribution of the second speech signal.

[0023] In a possible embodiment, the sample denoising sampling signal is a fifth Gaussian noise value, which is generated based on the sample noisy speech signal, the sample speech signal and a sixth Gaussian noise value that obeys a standard normal distribution. The first target data is generated based on the standard deviation of the fifth Gaussian noise value and the sixth Gaussian noise value, and the sixth Gaussian noise value is generated based on a fourth random seed.

[0024] It can be understood that since the random numbers generated by the random seed at different times are different, the speech signal processing method provided in the present application performs noise reduction processing on the same segment of speech signal to be processed at different times, and the second speech signals output respectively are different or not completely the same, so that the target speech signals finally output respectively are different or not completely the same.

[0025] In a possible embodiment, the first speech signal and the second speech signal are repaired based on a second diffusion model to obtain a target speech signal, including: calculating the gradient of the second speech signal based on the first speech signal and the second speech signal, the gradient is used to characterize the probability distribution of the target speech signal, and the probability distribution of the target speech signal corresponds to the distribution of time-frequency points in the spectrogram of the target speech signal; sampling the gradient of the second speech signal based on a first Gaussian noise value to obtain the target speech signal, and the first Gaussian noise value is generated based on a first random seed.

[0026] In the above scheme, since Gaussian noise is superimposed according to the gradient in the sampling step, if the Gaussian noise conforms to the gradient or the probability distribution of the target speech signal, the attenuated high-frequency components will be enhanced, thereby repairing the high-frequency components.

[0027] It is understandable that traditional discriminative neural networks cannot repair high-frequency components by superimposing Gaussian noise based on probability distribution like generative neural networks.

[0028] In a possible embodiment, the gradient of the second voice signal is calculated based on the first voice signal and the second voice signal, including: inputting the first voice signal and the second voice signal into the second neural network to obtain the repaired gradient to be sampled of the second voice signal, and the repaired gradient to be sampled is used to characterize the probability distribution of the target voice signal; sampling the gradient of the second voice signal based on the first Gaussian noise value, including: predicting the repaired sampling signal based on the repaired gradient to be sampled and the first Gaussian noise value; calculating the gradient of the second voice signal based on the first voice signal and the second voice signal, and also including: inputting the repaired sampling signal into the second neural network to obtain the repaired gradient to be corrected of the repaired sampling signal, and the repaired gradient to be corrected is used to characterize the probability distribution of the target voice signal; sampling the gradient of the second voice signal, and also including: correcting the repaired sampling signal based on the repaired gradient to be corrected to obtain a repaired correction signal; generating the target voice signal based on the repaired correction signal.

[0029] The above solution further calculates the repaired gradient to be corrected based on the repaired sampled signal obtained by predicting the repaired gradient to be sampled, and corrects the repaired sampled signal based on the repaired gradient to be corrected. Compared with the diffusion model that only calculates the gradient once, or compared with the diffusion model that only predicts without correction, it is more effective in repairing high-frequency components.

[0030] In a possible embodiment, a repaired sampling signal is predicted based on the repaired gradient to be sampled and the first Gaussian noise value, including: calculating a repair drift coefficient based on the second speech signal based on a stochastic differential equation of a second diffusion model; calculating a repair inverse drift coefficient based on the repair drift coefficient and the repaired gradient to be sampled; and predicting a repaired sampling signal based on the repair inverse drift coefficient, the second speech signal and the first Gaussian noise value.

[0031] It can be understood that the inverse drift coefficient can be used to describe the path along which a noisy speech signal approaches a clean speech signal during the sampling process. Specifically, it reduces the time-frequency points that do not conform to the probability distribution of the target speech signal and, using the first Gaussian noise, generates a path along which time-frequency points conform to the probability distribution of the target speech signal.

[0032] In a possible embodiment, the second neural network is trained based on second input data and second target data, wherein the second input data includes a sample noisy speech signal, a sample attenuated speech signal, and a sample repaired sampling signal, the sample noisy speech signal is generated by the second sample noise signal and the second sample speech signal, the sample attenuated speech signal is obtained by convolving the sample noisy speech signal with a preset distance room impulse response, the sample repaired sampling signal is generated based on the sample attenuated speech signal and the second sample speech signal, the second target data is used to characterize the probability distribution of the second sample speech signal, and the preset distance room impulse response is used to simulate the process of sound being emitted from the sound source and propagating to the microphone when the distance between the sound source and the microphone of the electronic device is greater than or equal to a preset distance threshold.

[0033] The above scheme trains the neural network model by using the target data to represent the probability distribution of the sample speech signal used to generate the sample attenuated speech signal, so that the gradient output by the trained neural network model can represent the probability distribution of the target speech signal.

[0034] In one possible embodiment, the sample repair sampling signal is a second Gaussian noise value, which is generated based on the sample attenuated speech signal, the second sample speech signal and a third Gaussian noise value that obeys a standard normal distribution, and the second target data is generated based on the standard deviation of the second Gaussian noise value and the third Gaussian noise value, and the third Gaussian noise value is generated based on a second random seed.

[0035] It is understandable that since the random numbers generated by the random seed at different times are different, the speech signal processing method provided in this application performs repair processing on the same segment of speech signal to be processed at different times, and the target speech signals outputted respectively are different or not completely the same.

[0036] In a possible embodiment, a first speech signal is input into a first neural network to obtain a denoised gradient to be sampled of the first speech signal, including: inputting the first speech signal into the first neural network to obtain the first denoised gradient to be sampled of the first speech signal; calculating a denoised drift coefficient according to the first speech signal based on a stochastic differential equation of a first diffusion model, including: calculating a first denoised drift coefficient according to the first speech signal based on a stochastic differential equation of the first diffusion model; calculating a denoised inverse drift coefficient according to the denoised drift coefficient and the denoised gradient to be sampled, including: calculating a first denoised inverse drift coefficient according to the first denoised drift coefficient and the first denoised gradient to be sampled; calculating a denoised inverse drift coefficient according to the denoised inverse drift coefficient. The method comprises: predicting a first denoised sampling signal according to the first denoising inverse drift coefficient, the first speech signal and the first fourth Gaussian noise value; inputting the denoised sampling signal into the first neural network to obtain a denoising gradient to be corrected of the denoised sampling signal, comprising: inputting the first denoised sampling signal into the first neural network to obtain a first denoising gradient to be corrected of the first denoised sampling signal; and correcting the denoised sampling signal according to the denoising gradient to be corrected to obtain a denoised corrected signal, comprising: correcting the first denoised sampling signal according to the first denoising gradient to be corrected to obtain a first denoising corrected signal.

[0037] Alternatively, the first speech signal is input into the first neural network to obtain the denoising gradient to be sampled of the first speech signal, including: inputting the n-1th denoising correction signal and the first speech signal into the first neural network to obtain the nth denoising gradient to be sampled of the first speech signal; calculating the denoising drift coefficient according to the first speech signal based on the stochastic differential equation of the first diffusion model, including: calculating the nth denoising drift coefficient according to the n-1th denoising correction signal and the first speech signal based on the stochastic differential equation of the first diffusion model; calculating the denoising inverse drift coefficient according to the denoising drift coefficient and the denoising gradient to be sampled, including: calculating the nth denoising inverse drift coefficient according to the nth denoising drift coefficient and the nth denoising gradient to be sampled; predicting the denoising inverse drift coefficient, the first speech signal and the fourth Gaussian noise value. to the noise reduction sampling signal, including: predicting the nth noise reduction sampling signal according to the nth noise reduction inverse drift coefficient, the n-1th noise reduction correction signal and the nth fourth Gaussian noise value; inputting the noise reduction sampling signal into the first neural network to obtain the noise reduction gradient to be corrected of the noise reduction sampling signal, including: inputting the nth noise reduction sampling signal and the first speech signal into the first neural network to obtain the nth noise reduction gradient to be corrected of the nth noise reduction sampling signal; correcting the noise reduction sampling signal according to the noise reduction gradient to be corrected to obtain the noise reduction correction signal, including: correcting the nth noise reduction sampling signal according to the nth noise reduction gradient to be corrected to obtain the nth noise reduction correction signal, 2≤n≤N, n and N are both positive integers, and when n=N, the Nth noise reduction correction signal is used as the second speech signal.

[0038] The above scheme iterates the steps of gradient calculation, prediction and correction N times to achieve better noise reduction effect.

[0039] In a possible embodiment, inputting the first speech signal and the second speech signal into a second neural network to obtain a repaired gradient to be sampled of the second speech signal includes: inputting the first speech signal and the second speech signal into the second neural network to obtain a first repaired gradient to be sampled of the second speech signal; calculating a repair drift coefficient according to the second speech signal based on a stochastic differential equation of a second diffusion model, including: calculating a first repair drift coefficient according to the second speech signal based on a stochastic differential equation of the second diffusion model; calculating a repair inverse drift coefficient according to the repair drift coefficient and the repaired gradient to be sampled, including: calculating a first repair inverse drift coefficient according to the first repair drift coefficient and the first repaired gradient to be sampled; A repaired sampling signal is predicted based on the repaired inverse drift coefficient, the second speech signal and the first Gaussian noise value, including: predicting the first repaired sampling signal based on the first repaired inverse drift coefficient, the second speech signal and the first Gaussian noise value; inputting the repaired sampling signal into the second neural network to obtain a repaired gradient to be corrected of the repaired sampling signal, including: inputting the first repaired sampling signal into the second neural network to obtain a first repaired gradient to be corrected of the first repaired sampling signal; correcting the repaired sampling signal based on the repaired gradient to be corrected to obtain a repaired corrected signal, including: correcting the first repaired sampling signal based on the first repaired gradient to be corrected to obtain a first repaired corrected signal.

[0040] Alternatively, the first speech signal and the second speech signal are input into the second neural network to obtain the repaired gradient to be sampled of the second speech signal, including: inputting the first speech signal, the second speech signal and the m-1th repair correction signal into the second neural network to obtain the mth repaired gradient to be sampled of the second speech signal; calculating the repair drift coefficient according to the second speech signal based on the stochastic differential equation of the second diffusion model, including: calculating the mth repair drift coefficient according to the second speech signal and the m-1th repair correction signal based on the stochastic differential equation of the second diffusion model; calculating the repair inverse drift coefficient according to the repair drift coefficient and the repair gradient to be sampled, including: calculating the mth repair inverse drift coefficient according to the mth repair drift coefficient and the mth repair gradient to be sampled; calculating the repair inverse drift coefficient according to the repair inverse drift coefficient, the second speech signal and the first Gaussian noise. value, predicting a repaired sampling signal, including: predicting an mth repaired sampling signal according to the mth repair inverse drift coefficient, the m-1th repaired correction signal and the first Gaussian noise value; inputting the repaired sampling signal into the second neural network to obtain a repaired gradient to be corrected of the repaired sampling signal, including: inputting the first speech signal, the second speech signal and the mth repaired sampling signal into the second neural network to obtain an mth repaired gradient to be corrected of the mth repaired sampling signal; correcting the repaired sampling signal according to the repaired gradient to be corrected to obtain a repaired correction signal, including: correcting the mth repaired sampling signal according to the mth repaired gradient to be corrected to obtain an mth repaired correction signal, 2≤m≤M, m and M are both positive integers, and when m=M, the Mth repaired correction signal is used as the target speech signal.

[0041] The above scheme iterates the steps of gradient calculation, prediction and correction M times to achieve better effect of repairing high-frequency components.

[0042] In a second aspect, the present application provides an electronic device comprising one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code comprises computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation of the first aspect.

[0043] In a third aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method described in the first aspect and any possible implementation method of the first aspect.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0045] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0046] It is understandable that the electronic device provided in the second aspect, the chip system provided in the third aspect, the computer storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to perform the methods provided in this application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG1 is a schematic diagram of an example of a scenario to which an embodiment of the present application is applicable;

[0048] FIG2A is a time domain diagram and a time-frequency diagram of the speech signal 1 collected at 5 m in the scenario shown in FIG1 ;

[0049] FIG2B is a time domain diagram and a time-frequency diagram of the speech signal 2 collected at 30 m in the scenario shown in FIG1 ;

[0050] FIG3 is a schematic diagram of a speech signal processing method 100 provided in an embodiment of the present application;

[0051] FIG4 is a schematic diagram of a speech conversion method 200 provided in an embodiment of the present application;

[0052] FIG5 is a schematic diagram of a voice conversion method 300 provided in an embodiment of the present application;

[0053] FIG6 is a schematic diagram of a training process 400 of a deep neural network model provided in an embodiment of the present application;

[0054] FIG7 is a schematic diagram of a training process 500 of a deep neural network model provided in an embodiment of the present application;

[0055] FIG8 is a schematic diagram of the hardware structure of an electronic device 1000 provided in an embodiment of the present application;

[0056] FIG9 is a block diagram of a software system of an electronic device 1000 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0058] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are first introduced.

[0059] 1. Long distance or far field can be understood as the distance between the microphone used to collect or pick up sound and the sound source. Generally speaking, in the field of acoustics, a distance of more than 10 meters between the microphone and the sound source is considered long distance, and a distance of 20 or 30 meters between the microphone and the sound source is considered ultra-long distance. In the embodiments of this application, long distance can be understood as the distance between the microphone and the sound source is greater than or equal to a preset distance threshold.

[0060] A long-distance voice signal or far-field voice signal is a voice signal collected or picked up at a distance or in a far-field scenario. A far-field voice signal can be understood as a voice signal recorded when the microphone of an electronic device is at a distance greater than or equal to a preset distance threshold from the sound source.

[0061] 2. Room impulse response (RIR). This function simulates the process of sound propagating from a source to the microphones. This process may include noise and the sound from the noisy source. The distance between the sound source and the microphones affects the attenuation of high-frequency components. RIR can also cause reverberation due to factors such as wall reflections.

[0062] In some embodiments, the embodiments of the present application are applicable to scenarios where sound signals are collected at a long distance. The RIR data set used in the training of the second neural network in the present application is mainly a long-distance RIR data set, which may include, for example, long-distance RIRs recorded in an extra-large room, long-distance RIRs in an extra-large room simulated using a simulation algorithm, or long-distance RIRs generated in other ways, which are not limited in the present application. For example, the extra-large room here is generally a room with a straight-line distance greater than or equal to a preset distance threshold (such as 20m or 30m or 10m or 15m or 25m, etc.).

[0063] The long-range RIR in this application may also be referred to as a preset distance RIR, which is equal to the preset distance threshold. The preset distance RIR simulates the process of sound propagating from a sound source to the microphone of an electronic device when the distance between the sound source and the microphone is greater than or equal to the preset distance threshold.

[0064] It should be noted that while the RIR datasets in this application are all related to very large rooms, the long-range or far-field nature of this application is not coupled to large or very large rooms. In other words, the speech signal processing method provided in this application is applicable not only to enclosed spaces such as very large rooms, large lecture halls, and banquet halls, but also to non-enclosed spaces such as playgrounds and open-air venues surrounded by all sides. Furthermore, these enclosed and non-enclosed spaces all encompass a linear distance of more than 20 meters or more than 30 meters.

[0065] It is understandable that, because the second neural network's training process focuses on the probability distribution of sample speech signals, rather than on noise datasets or RIR datasets that interfere with the sample clean speech signal (or sample speech signal), the second neural network is unaffected by the type of noise and has better generalization performance. Therefore, even when using an interference dataset related to a closed space, the trained second neural network can still obtain the target speech signal for the speech signal to be processed in a non-enclosed space.

[0066] 3. Time domain diagram, spectrogram and high frequency components.

[0067] The horizontal axis of the time domain graph is time, and the vertical axis is amplitude. By performing feature transformation on the time domain graph, the corresponding spectrogram can be obtained.

[0068] A spectrogram (also known as a time-frequency graph) has time on the horizontal axis and frequency on the vertical axis. A spectrogram is composed of multiple time-frequency points. The brightness or depth of a point represents the energy level. For example, brighter or lighter points indicate greater energy; conversely, lower energy.

[0069] For example, the frequency range of the spectrogram is 0~8000KHz, 5000KHz~8000KHz is high frequency, and 0~5000KHz is medium and low frequency.

[0070] High-frequency components are the higher-frequency components of a speech signal, specifically represented by high-frequency time-frequency points in the spectrogram. For example, if the frequency range of a speech signal's spectrogram is 0 to 8000 kHz, the time-frequency points between 5000 and 8000 kHz are considered high-frequency components.

[0071] 4. Diffusion model: A type of generative model based on random processes that can be used to generate various types of data, such as images, text, and audio. Unlike traditional generative models, diffusion models do not rely on known labels or target data. Instead, they process random noise to gradually generate high-quality data from it. Therefore, diffusion models can be considered a type of unsupervised generative model.

[0072] Among diffusion models, the most common is the unconditional diffusion model, whose main task is to generate random samples of original data. However, for some specific applications, such as image restoration, image synthesis, and text-to-image generation, we may need more precise control over the generated samples, which requires the introduction of the conditional diffusion model.

[0073] An important feature of the conditional diffusion model is that it can control the generated samples by adding additional conditional information, such as specific parts of the image, text descriptions, etc. For example, in the task of text-to-image generation, we can control what elements the generated image should contain or what style it should present by adding text descriptions.

[0074] In terms of specific implementation, conditional diffusion models typically encode additional conditional information as part of the model during the training phase, and then control the generated results by decoding this conditional information during the generation phase. An important advantage of this approach is that it can provide more refined control over the generated results while maintaining the generative capabilities of the diffusion model.

[0075] The diffusion model involved in this application is a conditional diffusion model.

[0076] 5. Drift coefficient and reverse drift coefficient.

[0077] The drift coefficient and the diffusion coefficient are used together to characterize the difference between the random signal being noisy and the audio signal being processed during the process of adding noise to the clean signal.

[0078] The drift coefficient is the mean of the random signal sampled at the nth time in the N sampling process; the diffusion coefficient is the variance of the random signal sampled at the nth time in the N sampling process.

[0079] De-noising the audio signal by sampling. The inverse drift coefficient and the inverse diffusion coefficient are used together to characterize the difference between the denoised random signal and the clean signal during the denoising process of the audio signal.

[0080] The inverse drift coefficient is the mean of the random signal sampled at the nth time in the N sampling process; the inverse diffusion coefficient is the variance of the random signal sampled at the nth time in the N sampling process.

[0081] Here, 1≤n≤N and n and N are both integers.

[0082] 6. Gaussian noise: This is a common type of random noise with a Gaussian distribution, meaning that the noise intensity and frequency are uniformly distributed. In signal and image processing, Gaussian noise is often used as a tool to simulate various noises. The frequency response of Gaussian noise can be calculated using a Gaussian function. In the frequency domain, the spectral density function of Gaussian noise exhibits a Gaussian curve.

[0083] 7. Random seed is a computer science term for a random number generated with a true random number (seed) as the initial condition. Generally, computer random numbers are pseudo-random numbers, starting with a true random number (seed) and then iterating through a specific algorithm to generate random numbers.

[0084] It is understandable that in the present application, the random numbers (eg, Gaussian noise values) generated based on the same random seed at different times are different.

[0085] Currently, most approaches rely on microphone arrays to pick up far-field speech signals, then enhance them based on acoustic models to capture speech signals at the target location. These methods are driven by linear models and have limited performance. Errors in the localization module can distort the target speech signal, and these methods require multiple microphones, which is costly.

[0086] FIG1 is a schematic diagram of an example of a scenario to which an embodiment of the present application is applicable.

[0087] As shown in Figure 1, the dashed box can represent the enclosed space or non-enclosed space described above. A sound source is emitting sound. For example, microphone A (e.g., included in terminal device 1) and microphone B (e.g., included in terminal device 2) are used to collect speech signals at 5m and 30m from the sound source, respectively. The same sound signals are collected from the same sound source at the same time and at both normal and long distances.

[0088] It should be noted that only 5m and 30m are used as examples here to illustrate the problems that need to be solved in long-distance recording compared to normal distance recording, and it is not limited to specific values.

[0089] For example, students are taking a class in a large lecture hall (or other enclosed space). The teacher is lecturing at the podium. Student 1, sitting in the second row (approximately 5 meters from the podium), is recording the conversation using terminal device 1, while student 2, sitting in the last row (approximately 30 meters from the podium), is recording the conversation using terminal device 2. Generally speaking, the teacher's voice is easily audible in the audio recorded by student 1. However, in the audio recorded by student 2, the teacher's voice is very soft and has a very low signal-to-noise ratio. Even if the volume is increased to amplify the teacher's voice, the noise will also be amplified, and the teacher's voice will still be drowned out by the noise. Furthermore, even if some sentences can be barely heard, the sound will not be full-bodied enough.

[0090] For another example, students are listening to a speech on the playground (or other non-enclosed space). Student 1, who is about 5 meters away from the speaker, is recording using terminal device 1, and student 2, who is about 30 meters away, is recording using terminal device 2. Generally speaking, the speaker's voice can be easily heard in the audio recorded by student 1; however, in the audio recorded by student 2, the speaker's voice is very soft and has a very low signal-to-noise ratio. Even if the volume is increased to amplify the speaker's voice, the noise will also be amplified, and the speaker's voice will still be drowned out by the noise. Moreover, even if some sentences can be barely heard, the sound is not full enough.

[0091] FIG2A is a time domain diagram and a time-frequency diagram of the speech signal 1 collected at a distance of 5 m in the scenario shown in FIG1 .

[0092] FIG2B is a time domain diagram and a time-frequency diagram of the speech signal 2 collected at a distance of 30 m in the scenario shown in FIG1 .

[0093] It should be noted that, since the volume of the voice signal 2 collected at 30 m is very low, the volume needs to be amplified to barely hear the sound from the sound source, so FIG2B is processed after the volume is increased.

[0094] As shown in Figure 2A, the upper portion is time domain graph 1, and the lower portion is spectrogram 1, obtained by feature conversion based on time domain graph 1. As shown in spectrogram 1, the harmonic structure of speech signal 1 in the mid- and low-frequency regions is relatively obvious, and the time-frequency point distribution in the high-frequency region is also relatively obvious. Overall, there are a small number of irregularly distributed time-frequency points that are noise, but the energy of the time-frequency points of the mid- and low-frequency harmonic components and the high-frequency components is significantly greater than that of the noise time-frequency points, which means that the noise time-frequency points do not overwhelm the time-frequency points of the speech signal. From a listening perspective, the sound is very full and clear, and the noise is negligible.

[0095] As shown in Figure 2B, the upper portion is Time Domain Graph 2, and the lower portion is Spectrogram 2, obtained by feature transformation based on Time Domain Graph 2. Time Domain Graph 2 clearly shows a very low signal-to-noise ratio. As shown in Spectrogram 2, overall, the energy difference between the noise time-frequency points and the non-noise time-frequency points in Speech Signal 2 is not significant, indicating a very low signal-to-noise ratio. The harmonic structure of the mid- and low-frequency frequencies is vaguely visible in the first half of the speech signal, but becomes slightly more pronounced in the second half. The high-frequency components are almost completely drowned out by the noise time-frequency points. From a listening perspective, the noise is much louder than the non-noise. Even with very careful listening, the speech cannot be clearly heard in the first half. With a little more effort, the speech can be roughly understood in the second half, but the sound is noticeably less natural.

[0096] It is understandable that for the collection of far-field speech or long-distance speech, as the distance increases, on the one hand, the signal-to-noise ratio of the speech signal collected by the microphone gradually decreases, which will make the speech signal less and less clear in hearing; on the other hand, the high-frequency components of the speech signal will attenuate more, which will make the speech signal less and less full in hearing.

[0097] Therefore, how to improve the voice quality of the collected far-field voice signals has become an urgent problem to be solved.

[0098] In view of this, the present application provides a speech signal processing method, which first reduces the noise of the speech signal to be processed based on a diffusion model, and then repairs the high-frequency components based on the diffusion model, so as to finally obtain the target speech signal.

[0099] This solution not only reduces noise, making the target speech signal clearer, but also restores high-frequency components, making the target speech signal fuller. Furthermore, compared to solutions that only reduce noise on far-field speech signals, this solution can achieve a fuller sound.

[0100] Furthermore, since the attenuated high-frequency components may be submerged in noise, the noise reduction process may interfere with the attenuated high-frequency components of the first speech signal, potentially affecting the restoration effect. During the restoration process, the pre-noise-reduction speech signal is also used as a condition in the conditional diffusion model and serves as reference information for restoring the high-frequency components. This reduces the potential interference of the noise reduction process on the restoration process and improves the restoration effect.

[0101] Furthermore, by breaking down the processing of far-field voice signals into two steps, each using two diffusion models, this reduces algorithm complexity and saves computing power while achieving the same effect as using a single model. This means that using a single model to process far-field voice signals requires a more complex algorithm to achieve the same effect as both models, placing higher demands on the power consumption and computing power of the terminal device.

[0102] In some embodiments, the restoration process includes calculating the gradient of the denoised speech signal and sampling the gradient based on Gaussian noise. Since Gaussian noise is superimposed on the gradient during the sampling step, if the Gaussian noise conforms to the gradient or the probability distribution of the target speech signal, the attenuated high-frequency components are enhanced, thereby restoring the high-frequency components.

[0103] Traditional discriminative neural networks cannot, like generative neural networks, repair high-frequency components by superimposing Gaussian noise based on probability distribution.

[0104] FIG3 is a schematic diagram of a speech signal processing method 100 provided in an embodiment of the present application.

[0105] S101: Acquire a first speech signal to be processed.

[0106] For example, the microphone module 110 of an electronic device is used to remotely collect a speech signal s(k) to be processed. The speech signal to be processed, recorded by the microphone of the electronic device, is a time-domain speech signal. The feature conversion module 120 can convert the time-domain speech signal s(k) into a first speech signal x(f, t) in the time-frequency domain, and subsequent processing is performed based on the first speech signal x(f, t). Here, f represents frequency, and t represents time frame.

[0107] S102 : Perform noise reduction processing on the first speech signal using the noise reduction module 130 based on the first diffusion model to obtain a second speech signal y(f, t).

[0108] As a possible example of S102, Example 1 includes:

[0109] Step A-1: ​​Calculate the gradient of the first speech signal.

[0110] The gradient is used to characterize the probability distribution of the second speech signal, and the probability distribution of the second speech signal corresponds to the distribution of time-frequency points in the spectrogram of the second speech signal.

[0111] Step A-2: Sampling the gradient of the first speech signal to obtain a second speech signal.

[0112] It can be understood that the above consideration of the probability distribution of the second speech signal is to consider the integrity of the distribution of the time-frequency points of the second speech signal, and the integrity includes the relationship between the time-frequency points.

[0113] It is also understandable that, if the method of analyzing each time-frequency point one by one without considering the overall performance is used, even if the total error (e.g., the first error value) between the time-frequency points of the speech-enhanced signal and the time-frequency points corresponding to the ideal clean signal (the signal after all noise is removed from the first speech signal in the ideal situation) is very small, it is very likely that the following situation will occur: a small number of time-frequency points will have large errors (e.g., significantly larger than the first error value) relative to the corresponding time-frequency points of the ideal clean signal, while the errors of most time-frequency points will be very small (e.g., significantly smaller than the first error value). Therefore, the small number of time-frequency points with large errors will, in principle, cause a certain degree of damage to the denoised speech, resulting in poor spectral continuity and a poor listening experience.

[0114] Compared with the scheme that does not consider the overall distribution of time-frequency points, the present application takes into account the overall distribution of time-frequency points. If the total error is controlled to the same small value (such as the first error value), the errors of almost all time-frequency points will be more balanced and smaller, thereby reducing the damage of noise reduction processing to the speech signal in principle.

[0115] The total error in this application can be understood as the loss function value used by the neural network during the training process for calculating the denoised gradient to be sampled or the denoised gradient to be corrected.

[0116] As a possible specific implementation of Example 1, two gradient calculations are performed, and a correction process is performed after the sampling process. Example 1-1 includes:

[0117] Step C-1: Input the first speech signal into the first neural network to obtain the denoised sampled gradient of the first speech signal. The denoised sampled gradient is used to represent the probability distribution of the second speech signal.

[0118] Step C-2: predicting the denoised sampled signal based on the denoised sampled gradient.

[0119] Step C-3: input the denoised sample signal into the first neural network to obtain the denoised gradient to be corrected of the denoised sample signal, where the denoised gradient to be corrected is used to characterize the probability distribution of the second speech signal.

[0120] Step C-4: correcting the noise reduction sampled signal according to the noise reduction gradient to be corrected to obtain a noise reduction correction signal, and using the noise reduction correction signal to generate a second speech signal.

[0121] Among them, step C-1 and step C-3 are a specific example of step A-1; step C-2 and step C-3 are a specific example of step A-2.

[0122] The above scheme further calculates the denoised gradient to be corrected based on the denoised sampled signal obtained based on the prediction of the denoised gradient to be sampled, and corrects the denoised sampled signal based on the denoised gradient to be corrected. Compared with the diffusion model that only calculates the gradient once, or compared with the diffusion model that only predicts without correction, the denoising effect is better.

[0123] As a possible specific implementation of step C-2, the noise reduction sampling signal is predicted by the drift coefficient and the inverse drift coefficient. Example 1-2 includes:

[0124] Step E-1: Calculate a noise reduction drift coefficient according to the first speech signal based on a stochastic differential equation of the first diffusion model.

[0125] Step E-2: Calculate the noise reduction inverse drift coefficient according to the noise reduction drift coefficient and the noise reduction gradient to be sampled.

[0126] Step E-3: predicting a noise reduction sampling signal based on the noise reduction inverse drift coefficient, the first speech signal, and the fourth Gaussian noise value.

[0127] The fourth Gaussian noise value is generated based on the third random seed.

[0128] It can be understood that the inverse drift coefficient can be used to describe the path that a noisy speech signal follows during sampling, toward a clean speech signal. Sampling based on the inverse drift coefficient to obtain a gradient can achieve noise reduction.

[0129] Specifically, a more specific implementation will be introduced below in combination with Example 1-1 and Example 1-2 based on FIG4 .

[0130] Specifically, for the first neural network in Example 1-1, the training process will be introduced below based on Figure 6.

[0131] S103: Perform restoration processing on the first speech signal and the second speech signal based on the second diffusion model to obtain a target speech signal.

[0132] The second diffusion model is a conditional diffusion model, and the restoration process includes: using the first speech signal as a condition of the second diffusion model to restore the high-frequency components of the second speech signal.

[0133] For example, the restoration module 140 based on the second diffusion model is used to restore the first speech signal x(f, t) and the second speech signal y(f, t) to obtain the restored speech signal r(f, t). The restored speech signal r(f, t) is input into the feature inverse transformation module 150 to obtain the target speech signal

[0134] As a possible example of S102, Example 2 includes:

[0135] Step B-1, based on the first speech signal and the second speech signal, calculate the gradient of the second speech signal, the gradient is used to characterize the probability distribution of the target speech signal, and the probability distribution of the target speech signal corresponds to the distribution of time-frequency points in the spectrogram of the target speech signal.

[0136] Step B-2: Sampling the gradient of the second speech signal based on the first Gaussian noise value to obtain a target speech signal, where the first Gaussian noise value is generated based on the first random seed.

[0137] In the above scheme, since Gaussian noise is superimposed according to the gradient in the sampling step, if the Gaussian noise conforms to the gradient or the probability distribution of the target speech signal, the attenuated high-frequency components will be enhanced, thereby repairing the high-frequency components.

[0138] It is understandable that traditional discriminative neural networks cannot repair high-frequency components by superimposing Gaussian noise based on probability distribution like generative neural networks.

[0139] As a possible specific implementation of Example 2, two gradient calculations are performed, and a correction process is performed after the sampling process. Example 2-1 includes:

[0140] Step D-1: Input the first speech signal and the second speech signal into the second neural network to obtain the repaired gradient to be sampled of the second speech signal, and the repaired gradient to be sampled is used to characterize the probability distribution of the target speech signal.

[0141] Step D-2: predicting the repaired sampled signal based on the repaired sampled gradient and the first Gaussian noise value.

[0142] Step D-3: input the repaired sample signal into the second neural network to obtain the repaired gradient to be corrected of the repaired sample signal, and the repaired gradient to be corrected is used to characterize the probability distribution of the target speech signal.

[0143] Step D-4, sampling the gradient of the second speech signal, also includes: correcting the repaired sampled signal according to the repaired gradient to be corrected to obtain a repaired corrected signal, and the repaired corrected signal is used to generate the target speech signal.

[0144] Step D-1 and step D-3 are specific examples of step B-1, and step D-2 and step D-3 are specific examples of step B-2.

[0145] The above solution further calculates the repaired gradient to be corrected based on the repaired sampled signal obtained by predicting the repaired gradient to be sampled, and corrects the repaired sampled signal based on the repaired gradient to be corrected. Compared with the diffusion model that only calculates the gradient once, or compared with the diffusion model that only predicts without correction, it is more effective in repairing high-frequency components.

[0146] As a possible specific implementation of step D-2, the sampled signal is repaired by predicting the drift coefficient and the inverse drift coefficient. Example 2-2 includes:

[0147] Step F-1: Calculate the restoration drift coefficient according to the second speech signal based on the stochastic differential equation of the second diffusion model.

[0148] Step F-2: Calculate the repaired inverse drift coefficient based on the repaired drift coefficient and the repaired gradient to be sampled.

[0149] Step F-3: predicting and obtaining a repaired sampled signal based on the repaired inverse drift coefficient, the second speech signal, and the first Gaussian noise value.

[0150] It can be understood that the inverse drift coefficient can be used to describe the path along which a noisy speech signal approaches a clean speech signal during the sampling process. Specifically, it reduces the time-frequency points that do not conform to the probability distribution of the target speech signal and, using the first Gaussian noise, generates a path along which time-frequency points conform to the probability distribution of the target speech signal.

[0151] Specifically, a more specific implementation method will be introduced below in combination with Example 2-1 and Example 2-2 based on Figure 5.

[0152] Specifically, for the second neural network in Example 2-1, the training process will be introduced below based on Figure 7.

[0153] The embodiments of this application not only reduce noise, making the target speech signal clearer, but also restore high-frequency components, making the target speech signal more audible. Furthermore, compared to solutions that only reduce noise on far-field speech signals, this can make the target speech signal more audible.

[0154] As a possible further example of Example 1-1 and Example 1-2, Example 3, the steps of Example 1-1 and Example 1-2 need to be iterated N times to achieve a better noise reduction effect. Example 3 includes:

[0155] When the sampling number n=1, in step G-1, the first speech signal is input into the first neural network to obtain the first denoised sampled gradient of the first speech signal.

[0156] Step G-2: Calculate a first noise reduction drift coefficient according to the first speech signal based on the stochastic differential equation of the first diffusion model.

[0157] Step G-3: Calculate the first denoising inverse drift coefficient according to the first denoising drift coefficient and the first denoising gradient to be sampled.

[0158] Step G-4: predicting and obtaining a first denoised sampling signal based on the first denoised inverse drift coefficient, the first speech signal, and the first fourth Gaussian noise value.

[0159] Step G-5: input the first denoised sample signal into the first neural network to obtain the first denoised gradient to be corrected of the first denoised sample signal.

[0160] Step G-6: Correct the first noise reduction sampling signal according to the first noise reduction gradient to be corrected to obtain a first noise reduction corrected signal.

[0161] Alternatively, when the number of sampling times is n, and 2≤n≤N, and n and N are both positive integers, in step H-1, the n-1th noise reduction correction signal and the first speech signal are input into the first neural network to obtain the nth noise reduction gradient to be sampled of the first speech signal.

[0162] Step H-2: Calculate the nth noise reduction drift coefficient based on the stochastic differential equation of the first diffusion model and the n-1th noise reduction correction signal and the first speech signal.

[0163] Step H-3: Calculate the nth denoising inverse drift coefficient according to the nth denoising drift coefficient and the nth denoising gradient to be sampled.

[0164] Step H-4: predicting and obtaining the nth noise reduction sampling signal according to the nth noise reduction inverse drift coefficient, the n-1th noise reduction correction signal, and the nth fourth Gaussian noise value.

[0165] Step H-5: input the nth denoised sample signal and the first speech signal into the first neural network to obtain the nth denoised gradient to be corrected of the nth denoised sample signal.

[0166] Step H-6: Correct the nth denoised sampled signal according to the nth denoised gradient to be corrected to obtain the nth denoised corrected signal.

[0167] FIG4 is a schematic diagram of a speech conversion method 200 provided in an embodiment of the present application. FIG4 can be used as a possible further example of Example 3. For example, the noise reduction module 130 based on the first diffusion model includes the initialization module 201 to the judgment module 207 as shown in FIG4 .

[0168] Initialization module 201, initialize sampling times n = 1 (indicates the current first sampling), initialize sampling time s = 0.999, initialize sampling signal The initialized parameters are output to the drift coefficient calculation module 202 .

[0169] Wherein, x(f,t) is the first speech signal mentioned above.

[0170] in, Represents the denoised sampled signal.

[0171] Wherein, when the sampling number n satisfies 2≤n≤N, when the steps from the drift coefficient calculation module 202 to the judgment module 207 are executed for the nth time, is the n-1th noise reduction correction signal, which can be understood as the noise reduction correction signal obtained by the n-1th sampling.

[0172] Drift coefficient calculation module 202, the input of this module is x(f,t), s, The function of this module is to calculate the drift coefficient fc according to the stochastic differential equation of the diffusion model, such as The drift coefficient calculated by this module is output to the inverse drift coefficient calculation module 204 to calculate the inverse drift coefficient.

[0173] Here, the stochastic differential equation may be a forward stochastic differential equation of a diffusion model.

[0174] The step performed by the drift coefficient calculation module 202 may be understood as a possible example of step G-2 or step H-2.

[0175] Gradient calculation module 203, the input of this module is x(f,t), s, The function of this module is to use the trained first neural network model to calculate the denoised sampled gradient.

[0176] The step performed by the gradient calculation module 203 may be understood as a possible example of step G-1 or step H-1.

[0177] Specifically, the training process of the first neural network model will be described in detail below in conjunction with FIG6 .

[0178] The gradient calculated by this module is output to the inverse drift coefficient calculation module 204 to calculate the inverse drift coefficient.

[0179] The inverse drift coefficient calculation module 204, the input of this module is the drift coefficient fc and the gradient grad, using the formula fcr = -fc + g 2 (t)grad is used to calculate the inverse drift coefficient fcr, which is output to the prediction module 205 for prediction. g(t) is a predefined diffusion coefficient. In this application, g(t) can be a constant selected based on experience, such as 1.2.

[0180] The step performed by the inverse drift coefficient calculation module 204 may be understood as a possible example of step G-3 or step H-3.

[0181] The prediction module 205, which may also be called the sampling prediction module 205, uses the formula To predict the new sampled signal, z1 is Gaussian noise with mean zero and variance 1 (i.e., the fourth Gaussian noise mentioned above). It should be noted that generating Gaussian noise requires setting a random seed. For evidence collection, the random seed is set when the terminal device is reset.

[0182] Among them, the formula Said that Updated to A new sampling signal is obtained, that is, the noise reduction sampling signal obtained by the n-th sampling, also called the n-th noise reduction sampling signal.

[0183] Exemplarily, the sampling methods include but are not limited to ancestral sampling, Langevin kinetic sampling and annealing Langevin kinetic sampling (a combination of one or two), but no matter which sampling method is used, it is necessary to generate random numbers based on random seeds (that is, Gaussian noise with a mean of 0 and a variance of 1 in this application), and then sample based on the random numbers.

[0184] The step performed by the prediction module 205 can be understood as a possible example of step G-4 or step H-4.

[0185] After predicting the denoised sample signal, the prediction module 205 outputs the prediction result to the gradient calculation module 203 .

[0186] Gradient calculation module 203, the input of this module is s,x(f,t). is the nth sampling signal after update in the prediction module 205 The function of this module is to use the trained neural network model to calculate the denoising gradient to be corrected. The gradient calculated by this module is output to the correction module 206 for correction.

[0187] The step performed by the gradient calculation module 203 may be understood as a possible example of step G-5 or step H-5.

[0188] Correction module 206 uses the annealed Langevin dynamic sampling method to predict the new sampled signal according to the denoised gradient to be corrected. Perform correction to obtain the nth correction signal The annealing Langevin kinetic sampling method can be described in the related art. Alternatively, the Langevin kinetic sampling method can be used, and the Langevin kinetic sampling method can be described in the related art.

[0189] When 1≤n<N, it serves as the n+1th input of the drift coefficient calculation module 202, and when n=N, it serves as the final output of this method.

[0190] The step performed by the correction module 206 may be understood as a possible example of step G-6 or step H-6.

[0191] The judgment module 207 judges whether n is equal to N (the preset total number of sampling steps, for example, assuming it is 10). If n=N, the correction signal Output as the second speech signal, otherwise update the sampling number and sampling time, n←n+1 (that is, update n with n+1), s=max(0.03,(n-1)×(0.999-0.03) / N)), and cyclically execute the drift coefficient calculation module 202.

[0192] The beneficial effects of method 200 can refer to the beneficial effects of the corresponding solution in method 100.

[0193] As a possible further example of Example 2-1 and Example 2-2, Example 4, the steps of Example 2-1 and Example 2-2 need to be iterated M times to achieve a better effect of repairing high-frequency components. Example 4 includes:

[0194] When the sampling number m=1, in step I-1, the first speech signal and the second speech signal are input into the second neural network to obtain the first repaired to-be-sampled gradient of the second speech signal.

[0195] Step I-2: Calculate a first repair drift coefficient according to the second speech signal based on the stochastic differential equation of the second diffusion model.

[0196] Step I-3: Calculate the first repaired inverse drift coefficient according to the first repaired drift coefficient and the first repaired gradient to be sampled.

[0197] Step I-4: predicting and obtaining a first repaired sampling signal based on the first repaired inverse drift coefficient, the second speech signal, and the first Gaussian noise value.

[0198] Step I-5: input the first repaired sampling signal into the second neural network to obtain the first repaired gradient to be corrected of the first repaired sampling signal.

[0199] Step I-6: Correct the first repaired sampling signal according to the first repaired gradient to be corrected to obtain a first repaired correction signal.

[0200] Alternatively, when the number of times is m and 2≤m≤M, and m and M are both positive integers, in step J-1, the first speech signal, the second speech signal and the m-1th repaired and corrected signal are input into the second neural network to obtain the mth repaired gradient to be sampled of the second speech signal.

[0201] Step J-2: Calculate the mth restoration drift coefficient according to the second speech signal and the m-1th restoration correction signal based on the stochastic differential equation of the second diffusion model.

[0202] Step J-3: Calculate the mth repaired inverse drift coefficient according to the mth repaired drift coefficient and the mth repaired gradient to be sampled.

[0203] Step J-4: predicting and obtaining the mth repaired sampling signal according to the mth repaired inverse drift coefficient, the m-1th repaired correction signal, and the first Gaussian noise value.

[0204] Step J-5: Input the first speech signal, the second speech signal, and the mth repaired sample signal into the second neural network to obtain the mth repaired gradient to be corrected of the mth repaired sample signal.

[0205] Step J-6: Correct the mth repaired sampling signal according to the mth repaired gradient to be corrected to obtain the mth repaired correction signal.

[0206] FIG5 is a schematic diagram of a speech conversion method 300 provided in an embodiment of the present application. FIG5 can be used as a possible further example of Example 4. For example, the repair module 140 based on the second diffusion model includes the initialization module 301 to the judgment module 307 as shown in FIG5.

[0207] Initialization module 301, initialize sampling times m = 1 (indicates the current first sampling), initialize sampling time s = 0.999, initialize sampling signal The initialized parameters are output to the drift coefficient calculation module 302 .

[0208] Wherein, y(f,t) is the second speech signal mentioned above.

[0209] in, Represents the repaired sampling signal.

[0210] Wherein, when the sampling number m satisfies 2≤m≤M, when the steps from the drift coefficient calculation module 302 to the judgment module 307 are executed for the mth time, is the (m-1)th repair and correction signal, which can be understood as the repair and correction signal obtained by the (m-1)th sampling.

[0211] Drift coefficient calculation module 302, the input of this module is y(f,t), s and The function of this module is to calculate the drift coefficient fc according to the stochastic differential equation of the diffusion model, such as The drift coefficient calculated by this module is output to the inverse drift coefficient calculation module 304 to calculate the inverse drift coefficient.

[0212] Here, the stochastic differential equation may be a forward stochastic differential equation of a diffusion model.

[0213] The step performed by the drift coefficient calculation module 302 may be understood as a possible example of step I-2 or step J-2.

[0214] Gradient calculation module 303, the input of this module is y(f,t), s, and x(f,t). The function of this module is to use the trained first neural network model to calculate and repair the gradient to be sampled.

[0215] The step performed by the gradient calculation module 303 may be understood as a possible example of step I-1 or step J-1.

[0216] Specifically, the training process of the first neural network model will be described in detail below in conjunction with FIG6 .

[0217] The gradient calculated by this module is output to the inverse drift coefficient calculation module 304 to calculate the inverse drift coefficient.

[0218] The inverse drift coefficient calculation module 304 takes the drift coefficient fc and the gradient grad as input, calculates the inverse drift coefficient fcr using the formula fcr=-fc+grad, and outputs it to the prediction module 305 for prediction.

[0219] The step performed by the inverse drift coefficient calculation module 304 may be understood as a possible example of step I-3 or step J-3.

[0220] The prediction module 305, which may also be referred to as the sampling prediction module 305, uses the formula To predict the new sampled signal, z2 is Gaussian noise with mean zero and variance 1 (i.e., the first Gaussian noise mentioned above). It should be noted that generating Gaussian noise requires setting a random seed. For evidence collection, the random seed is set when the terminal device is reset.

[0221] Among them, the formula Said that Updated to A new sampling signal is obtained, that is, the repaired sampling signal obtained by the n-th sampling, also called the m-th repaired sampling signal.

[0222] Exemplarily, the sampling methods include but are not limited to ancestral sampling, Langevin kinetic sampling and annealing Langevin kinetic sampling (a combination of one or two), but no matter which sampling method is used, it is necessary to generate random numbers based on random seeds (that is, Gaussian noise with a mean of 0 and a variance of 1 in this application), and then sample based on the random numbers.

[0223] The step performed by the prediction module 305 can be understood as a possible example of step I-4 or step J-4.

[0224] After predicting and repairing the sampled signal, the prediction module 305 outputs the prediction result to the gradient calculation module 303 .

[0225] Gradient calculation module 303, the input of this module is s, y(f,t) and x(f,t). is the updated mth sampling signal in the prediction module 305 The function of this module is to use the trained neural network model to calculate and repair the gradient to be corrected. The gradient calculated by this module is output to the correction module 306 for correction.

[0226] The step performed by the gradient calculation module 303 can be understood as a possible example of step I-5 or step J-5.

[0227] Correction module 306 uses the annealing Langevin dynamic sampling method to repair the predicted new sampling signal according to the gradient to be corrected Perform correction and obtain the mth correction signal The annealing Langevin kinetic sampling method can be described in the related art. Alternatively, the Langevin kinetic sampling method can be used, and the Langevin kinetic sampling method can be described in the related art.

[0228] When 1≤m<M, it is used as the (m+1)th input of the drift coefficient calculation module 302 , and when m=M, it is used as the final output of this method.

[0229] The step performed by the correction module 306 may be understood as a possible example of step I-6 or step J-6.

[0230] The judgment module 307 judges whether m is equal to M (the preset total number of sampling steps, for example, assuming it is 10). If m=M, the correction signal It is output as the repaired speech signal r(f, t), otherwise the sampling times and sampling time are updated, m←m+1 (i.e., m is updated with m+1), s=max(0.03,(m-1)×(0.999-0.03) / M)), and the drift coefficient calculation module 302 is executed cyclically.

[0231] The beneficial effects of method 300 can refer to the beneficial effects of the corresponding solution in method 100.

[0232] As a further example of Example 1-1, Example 5, the first neural network is trained based on first input data and first target data, wherein the first input data includes a sample noisy speech signal and a sample noise reduction sampling signal, the sample noisy speech signal is generated by a first sample noise signal and a first sample speech signal, the sample noise reduction sampling signal is generated based on the sample noisy speech signal, and the first target data is used to characterize the probability distribution of the first sample speech signal.

[0233] The above scheme trains the neural network model by using target data to represent the probability distribution of the sample speech signal contained in the sample noisy speech signal, so that the gradient output by the trained neural network model can represent the probability distribution of the second speech signal.

[0234] Among them, the sample denoising sampling signal is the fifth Gaussian noise value, the fifth Gaussian noise value is generated based on the sample noisy speech signal, the sample speech signal and the sixth Gaussian noise value that obeys the standard normal distribution, the first target data is generated based on the standard deviation of the fifth Gaussian noise value and the sixth Gaussian noise value, and the sixth Gaussian noise value is generated based on the fourth random seed.

[0235] It can be understood that since the random numbers generated by the random seed at different times are different, the speech signal processing method provided in the present application performs noise reduction processing on the same segment of speech signal to be processed at different times, and the second speech signals output respectively are different or not completely the same, so that the target speech signals finally output respectively are different or not completely the same.

[0236] Figure 6 shows a schematic diagram of a training process 400 of a deep neural network model provided in an embodiment of the present application. The training process shown in Figure 6 can be understood as a specific example of the training process of the first neural network in method 100, that is, a specific example of Example 5.

[0237] The training process diagram of the deep neural network model includes the following modules.

[0238] The adding module 401 is used to linearly add the sample noise and the sample clean speech signal Ya(k) according to a certain signal-to-noise ratio to obtain the sample noisy speech signal Ys(k), which is a time domain signal.

[0239] The sample noise comes from a sample noise set, and the sample clean speech signal comes from a clean speech signal set. The sample noise set can be simulated or actually recorded. The sample clean speech data set can be pre-recorded.

[0240] The feature transformation module 402 performs feature transformation on the sample noisy speech signal Ys(k) to obtain the sample noisy speech signal Yx(f, t) in the time-frequency domain after feature transformation; it also performs feature transformation on the sample clean speech signal Ya(k) to obtain the sample clean speech signal YA(f, t) in the time-frequency domain after feature transformation.

[0241] For example, the feature transformation is short-time Fourier transform, amplitude transform, etc.

[0242] Sampling time generation module 403, this module does not need input information. The function of this module is to randomly generate a l min to l max The decimal between is used as the sampling time, where l min is the minimum value, such as 0.003, l max The sampling time generated by this module is output to the sample generation module 405.

[0243] The Gaussian noise generating module 404 is configured to generate a Gaussian noise value z3 (ie, the sixth Gaussian noise value mentioned above) with a mean of zero and a variance of 1.

[0244] The sample generation module 405 is used to generate a sample denoised sampling signal Yp(f,t) according to Yp(f,t)=μ(f,t,l)+σ(l)z3, where μ(f,t,l) is the mean and σ(l) is the standard deviation.

[0245] Among them, μ(f,t,l)=lx(f,t)+(1-l)A(f,t),

[0246] The first neural network 406 inputs the sample noisy speech signal Yx(f, t), the sample sampling time l and the sample noise reduction sampling signal Yp(f, t) into the first neural network 406, and outputs the noise reduction sample gradient, i.e. Ygrad1 = DNN s (Yp(f,t),l,Yx(f,t)).

[0247] The first neural network here can be composed of a variety of deep neural network models, such as noise conditional scoring network (NCSN)++ network, convolutional neural network (CNN), convolutional recurrent neural network (CRNN), U-net and other network models.

[0248] Loss function calculation module 407. The calculation formula of the loss function is Will As the target data, the loss function is calculated with the output prediction gradient.

[0249] The beneficial effects of the training process 400 can be seen from the beneficial effects of the corresponding solution in the method 100.

[0250] As a further example of Example 2-1, Example 6, the second neural network is trained based on second input data and second target data, wherein the second input data includes a sample noisy speech signal, a sample attenuated speech signal, and a sample repaired sampling signal, the sample noisy speech signal is generated by the second sample noise signal and the second sample speech signal, the sample attenuated speech signal is obtained by convolving the sample noisy speech signal with a preset distance room impulse response, the sample repaired sampling signal is generated based on the sample attenuated speech signal and the second sample speech signal, the second target data is used to characterize the probability distribution of the second sample speech signal, the first distance room impulse response is used to simulate the process of the sound source being emitted by the sound source and propagating to the microphone when the distance between the sound source and the microphone of the electronic device is a preset distance, and the preset distance is a preset distance threshold.

[0251] The above scheme trains the neural network model by using the target data to represent the probability distribution of the sample speech signal used to generate the sample attenuated speech signal, so that the gradient output by the trained neural network model can represent the probability distribution of the target speech signal.

[0252] Among them, the sample repair sampling signal is a second Gaussian noise value, the second Gaussian noise value is generated based on the sample attenuated speech signal, the sample speech signal and a third Gaussian noise value that obeys the standard normal distribution, the second target data is generated based on the standard deviation of the second Gaussian noise value and the third Gaussian noise value, and the third Gaussian noise value is generated based on the second random seed.

[0253] It is understandable that since the random numbers generated by the random seed at different times are different, the speech signal processing method provided in this application performs repair processing on the same segment of speech signal to be processed at different times, and the target speech signals outputted respectively are different or not completely the same.

[0254] Figure 7 shows a schematic diagram of a training process 500 of a deep neural network model provided in an embodiment of the present application. The training process shown in Figure 7 can be understood as a specific example of the training process of the second neural network in method 100, that is, a specific example of Example 6.

[0255] The training process diagram of the deep neural network model includes the following modules.

[0256] The adding module 501 is used to linearly add the sample noise and the sample clean speech signal Ya(k) according to a certain signal-to-noise ratio to obtain the sample noisy speech signal Ys(k), which is a time domain signal.

[0257] The sample noise comes from a sample noise set, and the sample clean speech signal comes from a clean speech signal set. The sample noise set can be simulated or actually recorded. The sample clean speech data set can be pre-recorded.

[0258] The convolution module 502 is configured to perform convolution calculation on the sample noisy speech signal Ys(k) and the long-range room impulse response (ie, long-range RIR) to obtain a sample attenuated speech signal Yd(k) with attenuated high-frequency components.

[0259] The long-distance RIRs come from the long-distance RIR dataset, which includes long-distance RIRs recorded in very large rooms and long-distance RIRs in very large rooms simulated using simulation algorithms.

[0260] The feature transformation module 503 is used to perform feature transformation on the sample attenuated speech signal Yd(k) to obtain the sample attenuated speech signal Yx(f, t) in the time-frequency domain after the feature transformation; and also perform feature transformation on the sample clean speech signal Ya(k) to obtain the sample clean speech signal YA(f, t) in the time-frequency domain after the feature transformation.

[0261] The noise reduction module 504 based on the first diffusion model is used to perform noise reduction processing on the sample noisy speech signal Yx(f,t) and output a sample noise-reduced speech signal Yy(f,t).

[0262] Sampling time generation module 505, this module does not need input information. The function of this module is to randomly generate a l min to l max The decimal between is used as the sampling time, where l min is the minimum value, such as 0.003, l max The sampling time generated by this module is output to the sample generation module 507.

[0263] The Gaussian noise generating module 506 is configured to generate a Gaussian noise value z4 (ie, the third Gaussian noise value mentioned above) with a mean of zero and a variance of 1.

[0264] The sample generation module 507 is used to generate a sample repair sampling signal Yq(f,t) according to Yq(f,t)=μ(f,t,l)+σ(l)z4, where μ(f,t,l) is the mean and σ(l) is the standard deviation.

[0265] Among them, μ(f,t,l)=lx(f,t)+(1-l)A(f,t),

[0266] The second neural network 508 inputs the sample attenuated speech signal Yx(f, t), the sample sampling time l and the sample repaired sampling signal Yq(f, t) into the second neural network 508, and outputs the repaired sample gradient, i.e. Ygrad2=DNN r (Yq(f,t),l,Yx(f,t),Yy(f,t)).

[0267] The second neural network here can be composed of a variety of deep neural network models, such as NCSN++ network, CNN, CRNN, U-Net and other network models.

[0268] Loss function calculation module 509. The calculation formula of the loss function is Will As the target data, the loss function is calculated with the output prediction gradient.

[0269] The beneficial effects of the training process 500 can be seen from the beneficial effects of the corresponding solution in the method 100.

[0270] Please refer to FIG8 , which shows a schematic diagram of the hardware structure of an electronic device 1000 provided in an embodiment of the present application.

[0271] The electronic device 1000 can be a headset, a mobile phone, a smart screen, a tablet computer, a wearable electronic device, an in-vehicle electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, etc. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device 1000.

[0272] 8 , the electronic device 1000 may include a processor 1010 , an audio module 1020 , a microphone 1020A, and optionally, a speaker 1020B.

[0273] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0274] The processor 1010 may include one or more processing units, for example: the processor 1010 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc.

[0275] The controller may be the nerve center and command center of the electronic device 1000. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0276] Processor 1010 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 1010 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 1010. If processor 1010 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 1010 latency, and thus improves system efficiency.

[0277] In this application, the processor 1010 is used to first perform noise reduction on the speech signal to be processed based on the diffusion model, and then repair the high-frequency components based on the diffusion model to finally obtain the target speech signal. On the one hand, noise reduction can make the target speech signal clearer in terms of auditory perception, and on the other hand, repairing the high-frequency components can make the target speech signal sound fuller in terms of auditory perception. Moreover, compared with the solution of only performing noise reduction processing on the far-field speech signal, the target speech signal can be made to sound fuller in terms of auditory perception.

[0278] In some embodiments, the processor 1010 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0279] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0280] The electronic device 1000 can implement audio functions such as music playback, recording, voice calls, etc. through the audio module 1020, the speaker 1020B, the microphone 1020A, and the application processor (not shown in the figure).

[0281] In a voice call or recording scenario, microphone 1020A is used to record far-field voice signals.

[0282] Optionally, during a voice call, the speaker 1020B is used to play the voice of the party talking to the user; in a recording scenario, if the user wants to listen to the recording content, the speaker 1020B plays the recording content.

[0283] Next, the software system of the electronic device 1000 will be described.

[0284] For example, the electronic device 1000 may be a mobile phone. The software system of the electronic device 1000 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software system of the electronic device 1000.

[0285] FIG9 shows a block diagram of a software system of an electronic device 1000 provided in an embodiment of the present application. Referring to FIG9 , the layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, namely, from top to bottom, the application layer (application layer), the application framework layer (framework layer), the hardware abstraction layer (HAL), the driver layer, and the hardware layer.

[0286] The application layer may include a series of application packages, such as a dialing application, a gallery application, etc. (not shown in the figure). In the embodiment of the present application, the application package may include applications such as recording and voice calls, all of which require the use of a microphone to record audio. The terminal device may perform the voice signal processing of the present application on the audio recorded by the microphone, including noise reduction processing and high-frequency component repair processing.

[0287] Alternatively, the application layer may also include other applications that require the voice signal recorded by the microphone to be processed according to the present application, which is not limited in this application.

[0288] The framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. In an embodiment of the present application, the framework layer includes a microphone service interface and a noise reduction and high-frequency component repair service interface. Among them, the noise reduction and high-frequency component repair service interface can provide an API and programming framework for applications that obtain noise reduction and high-frequency component repair services. The microphone service can be used to provide an API and programming framework for applications that call the microphone.

[0289] The Hardware Abstraction Layer (HAL) is an interface layer between the operating system kernel and upper-layer software, providing a virtual hardware platform for the operating system. In the embodiment of the present application, the hardware abstraction layer may include a microphone hardware abstraction layer and a noise reduction and high-frequency component repair algorithm. The microphone hardware abstraction layer may provide virtual hardware for microphone 1, microphone 2, or more microphone devices. The noise reduction and high-frequency component repair algorithm may include operating code and data for implementing the voice signal processing method provided in the embodiment of the present application.

[0290] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, including microphone device drivers and digital signal processor drivers. The microphone device driver drives the microphone sensor to collect sound signals and the audio signal processor to pre-process the sound signals to generate digital audio signals. The digital signal processor driver drives the digital signal processor to process the digital audio signals.

[0291] The hardware layer includes a sensor and an audio signal processor. Among them, the sensor includes microphone 1 and microphone 2. The microphone included in the sensor corresponds one-to-one to the virtual microphone included in the microphone hardware abstraction layer. The audio signal processor can be used to convert the sound signal collected by the microphone into an audio digital signal. The digital signal processor can be used to process the audio digital signal. It should be noted that the software structure diagram of the electronic device shown in Figure 9 provided in this application is only used as an example and does not limit the specific module division in different layers of the Android operating system. For details, please refer to the introduction of the Android operating system software structure in conventional technology.

[0292] The following describes the method in the embodiment of the present application in detail in combination with the above hardware structure and system structure:

[0293] In response to enabling recording or voice call applications, such applications may call the noise reduction and high-frequency component repair service interface to obtain the noise reduction and high-frequency component repair service application programming interface and programming framework.

[0294] On the one hand, the noise reduction and high-frequency component restoration service can call the microphone service in the framework layer to collect sound signals from the environment. Specifically, the microphone service can call the microphone 1 in the microphone hardware abstraction layer to send a command to the microphone 1 sensor in the hardware layer to collect sound signals. The microphone hardware abstraction layer sends this command to the microphone device driver in the driver layer. Based on this command, the microphone device driver can activate microphone 1, thereby acquiring sound signals from the environment and generating digital audio signals through the audio signal processor.

[0295] The Noise Reduction and High-Frequency Repair service initializes the Noise Reduction and High-Frequency Repair algorithm. This algorithm obtains the digital audio signal generated by the audio signal processor through the microphone hardware abstraction layer. Then, based on the speech signal processing method stored in the Noise Reduction and High-Frequency Repair algorithm, it processes the acquired digital audio signal using the digital signal processor to produce a digital audio signal with noise reduction and high-frequency repair.

[0296] Specifically, how to process the digital audio signal to obtain the digital audio signal after noise reduction and restoration of high-frequency components can be referred to the method flow charts shown in Figures 3 to 7 above.

[0297] Finally, the noise reduction and high frequency component restoration algorithm can transmit the digital audio signal after noise reduction and high frequency component restoration to the noise reduction and high frequency component restoration service, and then transmit it back to the application layer.

[0298] The present invention provides a chip system comprising one or more processors configured to retrieve and execute instructions stored in a memory, thereby executing the method of the present invention. The chip system may be composed of a chip or may include a chip and other discrete devices.

[0299] Among them, the chip system may include an input circuit or interface for sending information or data, and an output circuit or interface for receiving information or data.

[0300] The present application also provides a computer program product, which, when executed by a processor, implements the method described in any method embodiment of the present application.

[0301] The computer program product can be stored in a memory and finally converted into an executable target file that can be executed by a processor through preprocessing, compilation, assembly and linking.

[0302] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements the method described in any method embodiment of the present application. The computer program can be a high-level language program or an executable target program.

[0303] The computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0304] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and equipment and the technical effects produced can refer to the corresponding processes and technical effects in the aforementioned method embodiments, and will not be repeated here.

[0305] In the several embodiments provided in this application, the disclosed systems, devices and methods can be implemented in other ways. For example, some features of the method embodiments described above can be ignored or not executed. The device embodiments described above are merely schematic, and the division of units is only a logical function division. There may be other division methods in actual implementation, and multiple units or components may be combined or integrated into another system. In addition, the coupling between the units or the coupling between the components may be direct coupling or indirect coupling, and the above coupling includes electrical, mechanical or other forms of connection.

[0306] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0307] It should be understood that the term "plurality" used herein refers to two or more. The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0308] The terms (or numbers) "first", "second", ... etc. that appear in the embodiments of the present application are only used for descriptive purposes, that is, they are only used to distinguish different objects, such as different "coordinates", etc., and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first", "second", ... etc. may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "at least one (item)" refers to one or more. "Multiple" means two or more. "At least one of the following (item)" or similar expressions refers to any combination of these items, including any combination of a single (item) or plural (items).

[0309] In short, the above description is only a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included in the scope of protection of this application.

Claims

1. A speech signal processing method, applied to electronic equipment, characterized in that: The method comprises: Acquire a first speech signal to be processed; performing noise reduction processing on the first speech signal based on a first diffusion model to obtain a second speech signal; The first speech signal and the second speech signal are repaired based on a second diffusion model to obtain a target speech signal, wherein the second diffusion model is a conditional diffusion model, and the repair processing includes: using the first speech signal as a condition for the second diffusion model to repair the high-frequency components of the second speech signal.

2. The method according to claim 1, wherein The first voice signal is a far-field voice signal, or the distance between the microphone of the electronic device and the sound source of the first voice signal is greater than or equal to a preset distance threshold.

3. The method according to claim 1, wherein The repairing process of the first speech signal and the second speech signal based on the second diffusion model to obtain a target speech signal includes: Calculating a gradient of the second speech signal based on the first speech signal and the second speech signal, wherein the gradient is used to characterize a probability distribution of the target speech signal, where the probability distribution of the target speech signal corresponds to a distribution of time-frequency points in a spectrogram of the target speech signal; The target speech signal is obtained by sampling the gradient of the second speech signal based on a first Gaussian noise value, where the first Gaussian noise value is generated based on a first random seed.

4. The method according to claim 3, wherein Calculating the gradient of the second speech signal based on the first speech signal and the second speech signal includes: inputting the first speech signal and the second speech signal into a second neural network to obtain a repaired to-be-sampled gradient of the second speech signal, wherein the repaired to-be-sampled gradient is used to characterize the probability distribution of the target speech signal; The sampling process of the gradient of the second speech signal based on the first Gaussian noise value includes: predicting and repairing the sampled signal according to the repaired gradient to be sampled and the first Gaussian noise value; Calculating the gradient of the second speech signal based on the first speech signal and the second speech signal further includes: inputting the repaired sampled signal into the second neural network to obtain a repaired gradient to be corrected of the repaired sampled signal, wherein the repaired gradient to be corrected is used to characterize the probability distribution of the target speech signal; The sampling process of the gradient of the second speech signal further includes: correcting the repaired sampled signal according to the repaired gradient to be corrected to obtain a repaired corrected signal, and the repaired corrected signal is used to generate the target speech signal.

5. The method according to claim 4, wherein The predicting and repairing the sampled signal according to the repaired gradient to be sampled and the first Gaussian noise value includes: Calculating a repair drift coefficient according to the second speech signal based on a stochastic differential equation of the second diffusion model; Calculating a repair inverse drift coefficient according to the repair drift coefficient and the repaired gradient to be sampled; The repaired sampling signal is predicted based on the repaired inverse drift coefficient, the second speech signal and the first Gaussian noise value.

6. The method according to claim 4, wherein The second neural network is trained based on second input data and second target data, wherein the second input data includes a sample noisy speech signal, a sample attenuated speech signal and a sample repaired sampling signal, the sample noisy speech signal is generated by a second sample noise signal and a second sample speech signal, the sample attenuated speech signal is obtained by convolving the sample noisy speech signal with a preset distance room impulse response, the sample repaired sampling signal is generated based on the sample attenuated speech signal and the second sample speech signal, the second target data is used to characterize the probability distribution of the second sample speech signal, and the preset distance room impulse response is used to simulate the process of sound being emitted by the sound source and propagating to the microphone of the electronic device when the distance between the sound source and the microphone is greater than or equal to a preset distance threshold.

7. The method according to claim 6, wherein The sample repair sampling signal is a second Gaussian noise value, which is generated based on the sample attenuated speech signal, the second sample speech signal and a third Gaussian noise value that obeys a standard normal distribution. The second target data is generated based on the standard deviation of the second Gaussian noise value and the third Gaussian noise value, and the third Gaussian noise value is generated based on a second random seed.

8. The method according to claim 1, wherein The performing noise reduction processing on the first speech signal based on the first diffusion model to obtain a second speech signal includes: Calculating a gradient of the first speech signal, where the gradient is used to characterize a probability distribution of the second speech signal, where the probability distribution of the second speech signal corresponds to a distribution of time-frequency points in a spectrogram of the second speech signal; Sampling the gradient of the first speech signal to obtain the second speech signal.

9. The method according to claim 8, wherein Calculating the gradient of the first speech signal includes: inputting the first speech signal into a first neural network to obtain a denoised sampled gradient of the first speech signal, wherein the denoised sampled gradient is used to characterize a probability distribution of the second speech signal; The sampling process of the gradient of the first speech signal includes: predicting a noise reduction sampling signal according to the noise reduction gradient to be sampled; The step of calculating the gradient of the first speech signal further includes: inputting the denoised sampled signal into the first neural network to obtain a denoised gradient to be corrected for the denoised sampled signal, wherein the denoised gradient to be corrected is used to characterize the probability distribution of the second speech signal; The sampling processing of the gradient of the first speech signal also includes: correcting the noise reduction sampled signal according to the noise reduction gradient to be corrected to obtain a noise reduction correction signal, and the noise reduction correction signal is used to generate the second speech signal.

10. The method according to claim 9, wherein The step of predicting the denoised sampled signal according to the denoised sampled gradient includes: Calculating a noise reduction drift coefficient according to the first speech signal based on a stochastic differential equation of the first diffusion model; Calculating a noise reduction inverse drift coefficient according to the noise reduction drift coefficient and the noise reduction gradient to be sampled; The denoised sampling signal is predicted based on the denoising inverse drift coefficient, the first speech signal and a fourth Gaussian noise value, wherein the fourth Gaussian noise value is generated based on a third random seed.

11. The method according to claim 9, wherein The first neural network is trained based on first input data and first target data, wherein the first input data includes a sample noisy speech signal and a sample noise reduction sampling signal, the sample noisy speech signal is generated by a first sample noise signal and a first sample speech signal, the sample noise reduction sampling signal is generated based on the sample noisy speech signal, and the first target data is used to characterize the probability distribution of the first sample speech signal.

12. The method according to claim 11, wherein The sample denoising sampling signal is a fifth Gaussian noise value, which is generated based on the sample noisy speech signal, the sample speech signal and a sixth Gaussian noise value that obeys a standard normal distribution. The first target data is generated based on the standard deviation of the fifth Gaussian noise value and the sixth Gaussian noise value, and the sixth Gaussian noise value is generated based on a fourth random seed.

13. An electronic device, characterized in that: The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 12.

14. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Speech enhancement method, device and equipment

    CN112712818A

  • Speech conversion model training method and device, speech conversion method and device and related equipment

    CN114758663A

  • Audio signal processing method and device, electronic equipment and storage medium

    CN116741191A

  • Voice-driven posture action generation method and device based on diffusion model

    CN117292704A

  • Audio processing method and device, storage medium and electronic equipment

    CN117612548A