Voice signal processing method and related equipment
By obtaining the probability distribution of the speech signal and using random seed generation gradients for sampling, the speech quality degradation caused by reverb in closed space is solved, and effective dereverb and speech signal recovery in a wider range is achieved.
Patent Information
- Application Number
- CN202410078687.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-25
AI Technical Summary
When recording or voice calls are performed in a closed space, the reverb voice signal will reduce the recording quality and the accuracy of speech recognition. The prior art dereverb methods do not work well under long reverb times or damage the voice signal.
By obtaining the probability distribution of the speech signal, random numbers are generated using random seeds for sampling, combining clean speech signals and noise signals to generate gradients, dereverberation processing is performed, and the spectral continuity and harmonic structure of the speech signal are restored.
It improves the spectrum continuity and listening quality of voice signals, reduces voice signal damage, and improves the success rate and user experience of voice recognition.
Smart Images

Figure CN120375844A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to a method for processing voice signals and related devices. Background Art
[0002] In scenarios such as using a terminal device for recording, voice calls, or voice recognition (such as voice input, or a user interacting with a smart assistant through voice, etc.), the terminal device needs to collect voice through a microphone. When the terminal device collects sound in a closed space (such as a room, classroom, etc.), it will inevitably collect reverberant voice signals. The existence of reverberation greatly reduces the listening quality of the voice during recording or voice calls, as well as the accuracy of voice recognition. Summary of the Invention
[0003] This application provides a method for processing voice signals and related devices, which can improve the user experience.
[0004] In a first aspect, a method for processing a voice signal is provided, which is applied to an electronic device. The method includes: obtaining a first voice signal to be processed; preprocessing the first voice signal to obtain a probability distribution of the first voice signal, where the probability distribution of the first voice signal corresponds to the distribution of time-frequency points in the spectrogram of the first voice signal; sampling the first voice signal according to the probability distribution of the first voice signal to obtain a reverberation-removed first voice signal, and using the reverberation-removed first voice signal as a target voice signal.
[0005] In the above solution, the reverberation of the first voice signal is removed by using the probability distribution of the voice signal to obtain a reverberation-removed target voice signal, which ensures the continuity of the voice spectrum, improves the listening quality of the voice during recording or voice calls, and the success rate of voice recognition, thereby improving the user experience.
[0006] Compared with the discriminative reverberation-removal method, the reverberation-removed voice signal according to the voice signal processing method provided in the embodiments of this application can ensure the spectral continuity of the reverberation-removed voice signal and reduce the damage to the voice signal; compared with the reverberation-removal method that depends on linear assumptions, it can process a wider range of reverberant voice signals. For example, even for voice signals with a longer reverberation time, it can ensure excellent reverberation-removal effects.
[0007] It can be understood that the spectrogram represents the energy magnitude of the speech signal at different time-frequency points. The probability distribution of the first speech signal can also be understood as the distribution of time-frequency points on the spectrogram and the energy distribution of the speech signal. Among them, the time-frequency points with higher energy (corresponding to the clean speech part) are more concentrated, and the probability of extracting response sample points during the sampling process is greater. The time-frequency points with lower energy (corresponding to the reverberation component) are more dispersed, and the probability of extracting response sample points during the sampling process is smaller. Thus, it is possible to approximate the overall distribution with a relatively small number of sample points to achieve dereverberation. Therefore, compared with the spectrogram of the first speech signal, the spectrogram of the target speech signal removes the time-frequency points of the reverberation component and can also restore the harmonic structure of the clean speech signal.
[0008] It can also be understood that in the related art, the probability distribution of the speech signal is not considered, but the speech spectrum is treated as discrete points, which will inevitably lead to poor continuity of the speech spectrum and further affect the listening quality of the noise-reduced speech. In this application, the time-frequency points of the speech signal are sampled as continuous points, ensuring the continuity of the speech spectrum.
[0009] It can also be understood that in the spectrogram of the dereverberated target speech signal, the harmonic structure is clearly visible. Compared with the spectrogram of the first speech signal, the harmonics of the reverberation component superimposed on the harmonics of the clean signal are removed, making the harmonic result clearer, indicating that the dereverberation effect is very good. Compared with the spectrogram after dereverberation by the discriminant dereverberation method, the speech signal is not damaged and the spectrum continuity is high. In terms of listening perception, the sound is more natural and full, and closer to human voice. Especially for the scenario of speech recognition, it can improve the recognition rate of speech by the electronic device and enhance the user experience. Especially for the scenarios of recording or voice calls, it can improve the quality of the voice and thus enhance the listening perception and the user experience.
[0010] In a possible embodiment, sampling processing is performed on the first speech signal according to the probability distribution of the first speech signal, including: generating a first random number according to a preset random seed; performing sampling processing according to the first random number and the probability distribution of the first speech signal.
[0011] Among them, the preset random seed is determined through a large number of experiments, which can make the listening perception of the dereverberated signal better.
[0012] It can be understood that since the random numbers generated by the random seed are different at different times, the method provided in this application processes the first speech signal at different times and outputs different or not completely the same dereverberated target speech signals respectively.
[0013] Exemplarily, sampling according to a probability distribution includes two sub-steps: prediction and correction. Among them, the prediction step uses ancestral sampling; the correction step uses Langevin dynamics sampling or annealed Langevin dynamics sampling, and the correction step is used to correct the prediction result of the prediction step.
[0014] In a possible embodiment, preprocessing the first speech signal to obtain the probability distribution of the first speech signal includes: predicting a second speech signal according to the first speech signal, where the second speech signal is an estimated clean speech signal; predicting the nth first Gaussian noise value according to the first speech signal; generating the nth gradient according to the second speech signal, the nth first Gaussian noise value, and the time step of the nth sampling; sampling processing according to the probability distribution of the first speech signal includes: in the case of n = 1, sampling the first speech signal using the first gradient to obtain the first sampling signal of the first speech signal, where the first gradient is used to characterize the probability distribution of the first speech signal; or, in the case of n ≥ 2, sampling the (n - 1)th sampling signal of the first speech signal using the nth gradient to obtain the nth sampling signal of the first speech signal, where the nth gradient is used to characterize the probability distribution of the (n - 1)th sampling signal, and n is a positive integer.
[0015] In the above solution, the predicted clean speech signal and the noise signal are combined to generate a gradient and sample. Compared with only using the predicted clean speech signal as the dereverberated signal, the spectral continuity is better and the damage to the speech signal is less; compared with only generating a gradient based on the predicted noise signal and sampling on the basis of the predicted clean signal, the number of sampling steps can be significantly reduced, saving computing power and time.
[0016] In a possible embodiment, sampling the first speech signal using the first gradient to obtain the first sampling signal of the first speech signal includes: generating the first second Gaussian noise value according to a preset first random seed; predicting the first sampling signal of the first speech signal according to the first second Gaussian noise value, the first gradient, and the first speech signal.
[0017] In a possible embodiment, sampling the (n - 1)th sampling signal of the first speech signal using the nth gradient to obtain the nth sampling signal of the first speech signal includes: generating the nth second Gaussian noise value according to a preset first random seed; predicting the nth sampling signal of the first speech signal according to the nth second Gaussian noise value, the nth gradient, and the speech signal obtained by the (n - 1)th sampling.
[0018] In a possible embodiment, predicting the nth first Gaussian noise value according to the first speech signal includes: in the case of n = 1, obtaining the first first Gaussian noise value according to the first speech signal, the time step of the first sampling, and the third speech signal, where the third speech signal is generated according to the first speech signal, the third Gaussian noise value, and the time step of the first sampling, and the third Gaussian noise value is generated according to a preset second random seed; or, in the case of n ≥ 2, inputting the first speech signal, the time step of the nth sampling, and the (n - 1)th sampling signal into a noise prediction module to obtain the nth first Gaussian noise value.
[0019] In a possible embodiment, the noise prediction module is a first neural network model trained based on first input data and first target data. The first input data includes a reverberant sample speech signal, a sample sampling time, and a reverberant sample sampling signal, where the reverberant sample sampling signal is generated according to the sample speech signal, a fourth Gaussian noise value generated based on a preset third random seed, and the sample sampling time. The first target data includes a sample clean speech and a sample noise obtained by subtracting the sample sampling signal. The sample clean speech is used to generate the reverberant sample speech signal.
[0020] In a possible embodiment, predicting the second speech signal according to the first speech signal includes: inputting the first speech signal into a reverberation prediction module to obtain the second speech signal, where the reverberation prediction module is a second neural network model trained based on second input data and second target data. The second input data includes a reverberant sample speech signal, and the second target data includes a clean speech signal corresponding to the reverberant sample speech signal.
[0021] In a possible embodiment, the sample speech signal is a time-frequency domain speech signal obtained based on feature transformation. The feature transformation includes short-time Fourier transform and amplitude transformation.
[0022] It can be understood that the role of the amplitude transformation is to make the features of the signal more prominent and improve the success rate of neural network learning.
[0023] In a second aspect, the present application provides an electronic device, which includes one or more processors and one or more memories; wherein, the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code. The computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation manner in the first aspect.
[0024] In a third aspect, an embodiment of this application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the processors are configured to call computer instructions to cause the electronic device to execute the method described in the first aspect and any possible implementation manner of the first aspect.
[0025] In a fourth aspect, this application provides a computer-readable storage medium including instructions, which, when running on an electronic device, cause the electronic device to execute the method described in the first aspect and any possible implementation manner of the first aspect.
[0026] In a fifth aspect, this application provides a computer program product including instructions, which, when running on an electronic device, cause the electronic device to execute the method described in the first aspect and any possible implementation manner of the first aspect.
[0027] It can be understood that the electronic device provided in the second aspect, the chip system provided in the third aspect, the computer storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute the method provided in this application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, and will not be elaborated here. Description of the Drawings
[0028] Figure 1A A schematic diagram showing an example of a scenario applicable to an embodiment of this application;
[0029] Figure 1B A schematic diagram showing another example of a scenario applicable to an embodiment of this application;
[0030] Figure 2A A time-domain diagram and a spectrogram of the reverberant speech signal #1 provided by an embodiment of this application are shown;
[0031] Figure 2B A time-domain diagram and a spectrogram of the speech signal #2 after reverberation removal by a second related technology are shown;
[0032] Figure 3 A time-domain diagram and a spectrogram of the speech signal #3 after reverberation removal by the speech signal processing method provided by this application are shown;
[0033] Figure 4 A schematic diagram showing the processing method 100 of the speech signal provided by an embodiment of this application is shown;
[0034] Figure 5 A schematic diagram showing a specific embodiment of the processing method 100 of the speech signal provided by an embodiment of this application is shown;
[0035] Figure 6A schematic diagram showing the training process of the first neural network model provided by an embodiment of the present application;
[0036] Figure 7 A schematic diagram showing the training process of the second neural network model provided by an embodiment of the present application;
[0037] Figure 8 A schematic diagram of the hardware structure of an electronic device 1000 provided by an embodiment of the present application;
[0038] Figure 9 A block diagram of the software system of an electronic device 1000 provided by an embodiment of the present application. Detailed implementation manners
[0039] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0040] To better understand the implementation manners of the present application, the technical terms related to the implementation manners of the present application will be introduced first.
[0041] 1. Reverberation: Reverberation is an important concept in the audio field. It refers to the echo of sound caused by phenomena such as reflection, refraction, and diffraction during the sound propagation process. Simply put, reverberation is the sound delay effect produced by the reflection and attenuation of sound waves in space. The generation of reverberation is determined by the reflection and attenuation of sound in space. When a sound source emits sound waves, surrounding objects will reflect the sound waves, and the sound waves reach the listener's ears after multiple reflections. These reflected waves are superimposed with the direct wave after changes in time and intensity, producing an effect of lingering sound.
[0042] The method for processing the voice signal of the present application is particularly applicable to processing voice signals with reverberation. In the present application, a voice signal with reverberation is understood as a voice signal with a reverberation effect in terms of auditory perception. Correspondingly, a clean voice signal is a voice signal without reverberation.
[0043] In the present application, when a voice signal with reverberation is subtracted from the corresponding clean voice signal, the obtained data can be called the reverberation component.
[0044] 2. Dereverberation refers to the process of removing the influence of reverberation from the sound after it is picked up by a microphone.
[0045] 3. Diffusion Model: A type of generative model based on stochastic processes, which can be used to generate various types of data such as images, text, audio, etc. Different from traditional generative models, diffusion models do not rely on known labels or target data, but rather process random noise and gradually generate high-quality data from the noise. Therefore, diffusion models can be regarded as an unsupervised generative model.
[0046] In diffusion models, the most common is the unconditional diffusion model, whose main task is to generate random samples of the original data. However, for some specific applications, such as image inpainting, image synthesis, text-to-image generation, etc., we may need more refined control over the generated samples, which requires the introduction of conditional diffusion models.
[0047] An important feature of conditional diffusion models is that they can control the generated samples by adding additional conditional information, such as specific parts of an image, text descriptions, etc. For example, in the task of text-to-image generation, we can control which elements the generated image should contain or what style it should present by adding text descriptions.
[0048] In the specific implementation, conditional diffusion models usually encode the additional conditional information as part of the model during the training stage of the diffusion model, and then decode this conditional information during the generation stage to control the generation result. An important advantage of this method is that it can provide more refined control over the generation result while maintaining the generation ability of the diffusion model.
[0049] The diffusion model involved in this application is a conditional diffusion model. During the training stage, the reverberant speech signal is used as the conditional input to train the diffusion model.
[0050] 4. Random seed is a computer professional term, a kind of random number with a true random number (seed) as the initial condition for the object of random numbers. Generally, the random numbers generated by computers are pseudo-random numbers, which use a true random number (seed) as the initial condition and then use a certain algorithm to continuously iterate to generate random numbers.
[0051] It can be understood that in this application, the random numbers generated according to the same random seed at different times are different.
[0052] 5. The probability distribution of the speech signal can be understood as the distribution of time-frequency points in the spectrogram of the speech signal.
[0053] It can be understood that the electronic device records sound through a microphone to obtain a time-domain speech signal. The time-domain speech signal is subjected to feature transformation to obtain a time-frequency domain speech signal. The time-frequency domain speech signal can be represented by a spectrogram (or, in this application, also referred to as a probability distribution diagram or a feature map or a time-frequency map).
[0054] On a spectrogram, the horizontal axis represents time, the vertical axis represents frequency, and the darkness (or brightness) of the color represents the energy intensity of the speech signal at that moment and frequency. That is to say, the spectrogram represents the energy magnitude of the speech signal at different time-frequency points. By observing the spectrogram, the characteristics of the speech signal at different times and frequencies can be intuitively understood (therefore, the spectrogram is also called a feature map), so as to analyze and process the speech signal.
[0055] Specifically, it will be described below in combination with Figure 2A 、 Figure 2B and Figure 3 for an example introduction of the spectrogram.
[0056] 6. Spectral continuity (or, in this application, it can also be called spectral consistency): The spectral continuity of speech refers to the continuity of the speech signal in the frequency domain, that is, the spectrum of the speech signal changes continuously over time.
[0057] When the speech signal is damaged, its spectral distribution may change, thus affecting the spectral continuity of the speech signal. For example, when the speech signal is interfered by noise, echo, reverberation, etc., its spectral distribution may become uneven, thus affecting the spectral continuity of the speech signal. Therefore, in speech signal processing and analysis, it is necessary to consider whether the speech signal is damaged and the degree of damage, so as to select appropriate methods and technologies to process and analyze the speech signal and improve the quality and accuracy of the speech signal.
[0058] 7. Gaussian noise: It is a common random noise, which has the characteristics of Gaussian distribution, that is, the intensity and frequency of the noise are evenly distributed. In signal processing and image processing, Gaussian noise is often used as a tool to simulate various noises. The frequency response of Gaussian noise can be calculated by the Gaussian function. In the frequency domain, the spectral density function of Gaussian noise presents a Gaussian curve.
[0059] 8. Impulse response: In signal processing, the impulse response, or impulse response function (IR function, IRF) generally refers to the output (response) of the system when the input is a unit impulse function. More generally, the impulse response is the reaction of any dynamic system to certain external changes. In both cases, the impulse response describes the reaction of the system as a function of time (or possibly as a function of other independent variables that parameterize the dynamic behavior of the system).
[0060] The acoustic impulse response (IR), also known as the impulse response, refers to the time-domain (i.e., time-amplitude) response characteristics of a system under test to an impulse excitation signal. The impulse response contains rich information about sound propagation, including necessary information such as arrival time, direct sound frequency components, discrete reflected sounds, concrete attenuation characteristics, signal-to-noise ratio, and speech intelligibility, as well as the overall frequency response.
[0061] 9. Convolution: It is an integral operation that can be used to describe the relationship between the input and output of a linear time-invariant system. For example, in signal processing, for an input signal passing through a linear system, its output can be obtained through the convolution operation of the input and a function (impulse response function) characterizing the system characteristics. The physical meaning of convolution lies in that it can describe the response of the system to the input signal.
[0062] 10. Time step is a concept used in numerical simulations, which represents the discretization of the simulation system in time. In numerical simulations, the time range of the problem is divided into continuous small time intervals, and each time interval is called a time step. Within each time step, the state of the system is updated and evolved according to physical laws and calculation methods. The size of the time step can be adjusted as needed, usually determined according to the stability and accuracy requirements of the simulation system. A smaller time step can improve the accuracy of the simulation but will increase the computational complexity and time consumption.
[0063] In scenarios such as using a terminal device for recording, voice calls, or speech recognition (such as voice input, or voice interaction between the user and a smart assistant), the terminal device needs to collect voice through a microphone. When the terminal device collects sound in an enclosed space (such as a room, classroom, bathroom, etc.), it will inevitably collect reverberant voice signals. The existence of reverberation greatly reduces the listening quality of the voice during recording or voice calls, as well as the accuracy of speech recognition.
[0064] Figure 1A The figure shows a schematic diagram of an example of a scenario applicable to the embodiments of the present application.
[0065] As Figure 1A shown, when the user is washing up in the bathroom and wants to send a message to Xiaoming, it is inconvenient to type with hands, so the user selects voice input. The user says "I'm washing up", and the voice signal collected by the electronic device is a reverberant voice signal. If speech recognition is performed based on this reverberant voice signal, the text output in the input box of the chat interface may be "I'm washing up number".
[0066] Figure 1B The figure shows a schematic diagram of another example of a scenario applicable to the embodiments of the present application.
[0067] As Figure 1B shown, after an electronic device collects a reverberant speech signal, if the speech signal processing method of the present application is used to obtain a dereverberated speech signal, and then speech recognition is performed based on the dereverberated speech signal, the text output in the input box of the chat interface is "I'm washing up", which can improve the accuracy of speech recognition.
[0068] It should be noted that Figure 1A and Figure 1B taking the bathroom as an example for illustration, it can also be replaced with other scenarios that generate reverberation; Figure 1A and Figure 1B taking voice input as an example for illustration, it can also be replaced with other applications that need to collect the user's voice.
[0069] Figure 2A Fig. shows the time-domain diagram and spectrogram of the reverberant speech signal #1 provided by the embodiment of the present application.
[0070] As Figure 2A shown, it includes two parts. The upper part is the time-domain diagram #1 of the reverberant speech signal #1, and the lower part is the spectrogram #1.
[0071] Among them, the horizontal axis of the time-domain diagram #1 is time, and the vertical axis is amplitude. By performing feature transformation on the reverberant speech signal #1 in this time domain, the corresponding reverberant speech signal #1 in the time-frequency domain is obtained, as shown in the spectrogram #1.
[0072] Among them, the horizontal axis of the spectrogram #1 is time, and the vertical axis is frequency (for example, including different frequency bands, ranging from 0 Hz to 8 KHz). Each point in the spectrogram is a time-frequency point, and the superposition of several time-frequency points can be considered as a harmonic (for example, the line connected by multiple points is a harmonic). The amplitude (or energy) of the time-frequency point is represented by the brightness of the time-frequency point. The brighter the time-frequency point, the greater the amplitude, and the darker the time-frequency point, the smaller the amplitude; in particular, the black area (for example Figure 2A the area pointed by the white arrow in) indicates no sound.
[0073] The probability distribution of the speech signal mentioned above refers to the distribution of each time-frequency point in the spectrogram. Taking the spectrogram #1 as an example, the probability distribution of the reverberant speech signal #1 refers to the distribution of each time-frequency point in the spectrogram #1.
[0074] There is no obvious boundary in time for the time-frequency points in the spectrogram #1, indicating that there are still time-frequency points of the reverberation component when the user does not make a sound.
[0075] It should be noted that the spectrograms provided in the present application (the original figure is a color figure) are all illustrated in grayscale.
[0076] In the first related technology, reverberation removal is performed on reverberant speech based on a statistical signal processing model. It is assumed that the reverberant signal and the direct signal have a linear relationship. Then, a reverberation prediction filter is used to predict the reverberation, and the predicted reverberation is subtracted from the reverberant speech signal to obtain the reverberation-removed speech signal. For this processing model, when the reverberation time is less than a certain threshold, the effect of reverberation removal is relatively good. When the reverberation time is very long, the assumption of a linear relationship does not hold. This results in limited performance of reverberation removal by this processing model when the reverberation time is relatively long.
[0077] In the second related technology, reverberation removal is performed on speech based on a discriminative deep neural network model. Such methods act on the reverberant speech signal through a mask or mapping function represented by a discriminative deep neural network model to obtain the reverberation-removed speech signal. The disadvantage of this type of method is that it will cause damage to the speech signal.
[0078] Figure 2B The time-domain graph and spectrogram of speech signal #2 after reverberation removal by the second related technology are shown.
[0079] And Figure 2A similarly, Figure 2B it includes two parts. The upper part is the time-domain graph #2, and the lower part is the spectrogram #2.
[0080] As Figure 2B shown, in a larger white box (taking this box as an example), almost all time-frequency points are black, indicating that there are almost no time-frequency points, or rather, there is almost no speech signal in this time-frequency region. And at the corresponding position in Figure 2A there is clearly a speech signal. This shows that the speech signal obtained after reverberation removal by the second related technology is damaged.
[0081] As Figure 2B shown, in a smaller white box (taking this box as an example), there are gaps between harmonics, indicating poor spectral continuity.
[0082] It can be understood that the damage to the speech signal and the poor spectral continuity will result in a very dry and insufficiently full sound in terms of auditory perception.
[0083] The present application provides a method for processing a speech signal. According to the probability distribution of a first reverberant speech signal, the reverberant speech signal is sampled to obtain a first reverberation-removed speech signal.
[0084] It is understandable that a spectrogram represents the energy magnitude of a speech signal at different time-frequency points. The probability distribution of the first speech signal can also be understood as the distribution of time-frequency points on the spectrogram and the energy distribution of the speech signal. Among them, the time-frequency points with higher energy (corresponding to the clean speech part) are more concentrated, and the probability of extracting response sample points during the sampling process is greater. The time-frequency points with lower energy (corresponding to the reverberation component) are more dispersed, and the probability of extracting response sample points during the sampling process is smaller. Thus, it is possible to approximate the overall distribution with a relatively small number of sample points to achieve dereverberation. Therefore, compared with the spectrogram of the first speech signal, the spectrogram of the target speech signal removes the time-frequency points of the reverberation component and can also restore the harmonic structure of the clean speech signal.
[0085] Figure 3 The time-domain graph and spectrogram of speech signal #3 after dereverberation processing by the speech signal processing method provided by this application are shown.
[0086] Similar to Figure 2A Similar, Figure 3 It includes two parts. The upper part is the time-domain graph #3, and the lower part is the spectrogram #3.
[0087] As shown in the spectrogram #3, the boundary between human voices at different time periods is particularly obvious (there are obvious black strip areas between the harmonics in different time periods in the time domain). It is understandable that in Figure 2A there are harmonics at the corresponding positions of the black strip areas, which are the harmonics generated by reverberation. For example, when a microphone records a 3s audio in a closed space, only the sound source emits sound in the 1st second and the 3rd second, and no sound is emitted in the 2nd second; while in the reverberant audio recorded by the microphone, there may be an echo of the sound emitted by the sound source in the 1st second in the 2nd second, and there are also time-frequency points in the 2nd second in the spectrogram corresponding to this audio.
[0088] As shown in the spectrogram #3, in a larger white box (taking this box as an example), the harmonic structure is clearly visible. Compared with the corresponding position in the spectrogram #1, the overlap between the harmonics is removed, making the harmonic result clearer, indicating that the dereverberation effect is very good. Compared with the corresponding position in the spectrogram #2, the speech signal is not damaged and the reduction degree of the speech signal is high. Especially for the scenario of speech recognition, it can improve the recognition rate (or recognition success rate) of the electronic device for speech and enhance the user experience. Especially for the scenarios of recording or voice calls, it can improve the quality of the voice and thus enhance the listening experience and enhance the user experience.
[0089] As shown in the spectrogram #3, in a smaller white box (taking this box as an example), compared with the corresponding position in the spectrogram #2, the continuity between the harmonics is stronger.
[0090] It is understandable that the voice signal is not damaged and has high spectral continuity. In terms of the sense of hearing, the sound is more natural and full, and closer to the human voice.
[0091] In addition, as shown in spectrogram #3, especially the harmonics in the middle and high frequencies are well restored.
[0092] Figure 4 The schematic diagram of the processing method 100 of the voice signal provided by the embodiment of the present application is shown.
[0093] S101, obtain the first voice signal to be processed.
[0094] For example, use an electronic device to collect the first voice signal to be processed in a closed space (such as a bathroom, a classroom). Among them, the first voice signal recorded by the microphone of the electronic device is a voice signal in the time domain.
[0095] Exemplarily, through feature transformation, the first voice signal in the time domain can be converted into the first voice signal in the time-frequency domain, and then the first voice signal in the time-frequency domain is subjected to subsequent processing.
[0096] S102, preprocess the first voice signal to obtain the probability distribution of the first voice signal.
[0097] Among them, the probability distribution of the first voice signal corresponds to the distribution of time-frequency points in the spectrogram of the first voice signal.
[0098] Exemplarily, an estimated clean signal and a first Gaussian noise value (which can be understood as the reverberation component in the first voice signal, and will be specifically introduced in combination with the second neural network below) are predicted according to the first voice signal, and the probability distribution of the first voice signal is determined according to the estimated clean signal and the first Gaussian noise value. Optionally, a gradient is determined according to the estimated clean signal and the first Gaussian noise value, and the gradient is used to characterize the probability distribution of the first voice signal.
[0099] S103, according to the probability distribution of the first voice signal, perform sampling processing on the first voice signal to obtain the first voice signal after reverberation removal, and use the first voice signal after reverberation removal as the target voice signal.
[0100] Exemplarily, a first random number is generated according to a preset random seed; sampling processing is performed according to the first random number and the probability distribution of the first voice signal.
[0101] Among them, the preset random seed is determined through a large number of experiments, and can make the signal after reverberation removal have a better sense of hearing.
[0102] It can be understood that since the random numbers generated by the random seed are different at different times, the method 100 provided in this application processes the first speech signal at different times, and the output target speech signals after reverberation removal are different or not completely the same.
[0103] Alternatively, the target speech signal is determined according to the first speech signal after reverberation removal. For example, if the first speech signal after reverberation removal is a speech signal in the time-frequency domain, the first speech signal after reverberation removal is subjected to feature conversion to obtain a speech signal in the time domain as the target speech signal.
[0104] The above solution can perform reverberation removal processing on the first speech signal with reverberation to obtain the target speech signal after reverberation removal, improve the listening quality of speech during recording or voice calls, and the success rate during speech recognition, thereby enhancing the user experience.
[0105] Compared with the discriminant reverberation removal method, the speech signal after reverberation removal according to the speech signal processing method provided in the embodiments of the present application can ensure the spectral continuity of the reverberation-removed speech signal and reduce the damage to the speech signal; compared with the reverberation removal method relying on linear assumptions, it can process a wider range of reverberant speech signals. For example, even for speech signals with a longer reverberation time, excellent reverberation removal effects can be ensured.
[0106] Further examples are given below for the above steps.
[0107] Example 1, the target speech signal is obtained through N times of sampling processing. Example 1 gives the specific implementation manners of S102 and S103.
[0108] Among them, S102 can specifically be: Step 1-a, predicting a second speech signal according to the first speech signal, where the second speech signal is an estimated clean speech signal; Step 1-b, predicting the nth first Gaussian noise value according to the first speech signal; Step 1-c, generating the nth gradient according to the second speech signal, the nth first Gaussian noise value, and the time step of the nth sampling.
[0109] Among them, S103 can specifically include 2 cases.
[0110] Step 2-a, in the case of the first sampling, when n = 1, using the first gradient to sample the first speech signal to obtain the first sampling signal of the first speech signal, where the first gradient is used to characterize the probability distribution of the first speech signal.
[0111] Alternatively, in step 2-b, for samplings other than the first sampling, when 2 ≤ n ≤ N, the nth gradient is used to sample the (n - 1)th sampled signal of the first speech signal to obtain the nth sampled signal of the first speech signal, where the nth gradient is used to characterize the probability distribution of the (n - 1)th sampled signal, N ≥ 2 and both n and N are integers. Here, n is the current sampling number (or step number), and N is the preset total sampling number (or step number).
[0112] It can be understood that the predicted clean speech signal and the noise signal are combined to generate a gradient and sample. Compared with using only the predicted clean speech signal as the dereverberated signal, the spectral continuity is better and the damage to the speech signal is less; compared with generating a gradient only based on the predicted noise signal, sampling based on the predicted clean signal can significantly reduce the sampling steps, saving computing power and time.
[0113] Example 2: In Example 1, a second Gaussian noise value is sampled and generated as a random number, which is used for each corresponding sampling. In other words, each sampling is based on a second Gaussian noise value. For example, in step 2-a, for the first sampling, when n = 1, the first second Gaussian noise value is generated according to the preset first random seed; according to the first second Gaussian noise value, the first gradient, and the first speech signal, the first sampled signal of the first speech signal is predicted. For example, in step 2-b, for the nth sampling, the nth second Gaussian noise value is generated according to the preset first random seed; according to the nth second Gaussian noise value, the nth gradient, and the speech signal obtained from the (n - 1)th sampling, the speech signal obtained from the nth sampling is predicted.
[0114] Example 3: A specific implementation manner of step 1-b is given.
[0115] In step 1-b-a, for the first sampling, when n = 1, the first first Gaussian noise value is obtained according to the first speech signal, the time step of the first sampling, and the third speech signal. Wherein, the third speech signal is generated according to the first speech signal, the third Gaussian noise value, and the time step of the first sampling, and the third Gaussian noise value is generated according to the preset second random seed.
[0116] Alternatively, in step 1-b-b, for samplings other than the first sampling, when 2 ≤ n ≤ N, the first speech signal, the time step of the nth sampling, and the (n - 1)th sampled signal are input into the noise prediction module to obtain the nth first Gaussian noise value.
[0117] Example 4. In Example 3, the noise prediction module is a second neural network model trained based on the first input data and the first target data. The first input data includes a reverberant sample speech signal, a sample sampling time, and a reverberant sample sampling signal. The reverberant sample sampling signal is generated based on the sample speech signal, a fourth Gaussian noise value generated based on a preset third random seed, and the sample sampling time. The first target data includes the sample clean speech and the sample noise obtained by subtracting the sample sampling signal. The sample clean speech is used to generate the reverberant sample speech signal.
[0118] Example 5. Step 1-a may specifically be to input the first speech signal into the reverberation prediction module to obtain a second speech signal. The reverberation prediction module is a first neural network model trained based on the second input data and the second target data. The second input data includes the reverberant sample speech signal, and the second target data includes the clean speech signal corresponding to the reverberant sample speech signal.
[0119] Optionally, the sample speech signal is a time-frequency domain speech signal obtained based on feature transformation. The feature transformation includes short-time Fourier transform and amplitude transformation.
[0120] It can be understood that the role of the amplitude transformation is to make the features of the signal more prominent and improve the success rate of neural network learning.
[0121] Figure 5 FIG. shows a schematic diagram of a specific embodiment of the speech signal processing method 100 provided by the embodiments of the present application.
[0122] This specific embodiment includes two stages: a working stage and a training stage. The process of the working stage is as Figure 5 shown and is respectively executed by the signal acquisition module 101 to the feature inverse transformation module 108. The training stage will be introduced below in combination with Figure 6 and Figure 7
[0123] The signal acquisition module 101 is used to acquire the reverberant speech signal x(t) in the time domain, where t is time. Input x(t) into the feature transformation module 102.
[0124] The feature transformation module 102. The function of this module is to perform feature transformation on the input, such as short-time Fourier transform and amplitude transformation. The input of this module is the reverberant speech signal x(t) in the time domain. The output of this module is the reverberant signal X(f,t) in the time-frequency domain, where f in X(f,t) represents frequency and t represents time. This module also outputs X(f,t) to the noise prediction module 104 and the reverberation prediction module 103.
[0125] It is understandable that the signal acquisition module 101 and the feature transformation module 102 are used for an example of executing S101.
[0126] A reverberation prediction module 103, which is used to predict the estimated clean signal q of the reverberant signal. Alternatively, the estimated clean signal q in this application is the second speech signal mentioned above, also known as the clean signal or the estimated signal or the estimated clean signal. Wherein, the input of this module is X(f,t), the output is q, and it is given to the gradient calculation module 105.
[0127] Exemplarily, the reverberation prediction module 103 can be a reverberation prediction filter based on a linear hypothesis, or a trained deep neural network model (the first neural network model mentioned above).
[0128] A noise prediction module 104, which predicts the Gaussian noise value p at the current sampling step through a trained deep neural network model (the second neural network model mentioned above). Wherein, the input of this module is: X(f,t), t, the output is p, and it is given to the gradient calculation module 105. Wherein, n is the current sampling step. If it is executed for the first time, take n = 1. Each time it is executed, n is incremented by 1. When n>1, is the speech signal obtained from the (current nth) sampling at the (n - 1)th sampling, and is obtained from the sampling module 106. When n = 1, g is the third Gaussian noise value with a mean of zero and a variance of 1 generated according to the random seed, and σ is the standard deviation of the third Gaussian noise value.
[0129] A gradient calculation module 105, which is used to calculate the gradient (grad) according to the estimated clean signal q output by the reverberation prediction module 103 and the predicted noise p output by the noise prediction module 104. The formula is where s represents the time step, and the calculation method of the time step is s = max(1 - n / N, 0.003), where N is the preset total number of sampling steps, n is the current sampling step. If it is executed for the first time, take n = 1. Each time it is executed, n is incremented by 1. σ is the standard deviation of the predicted Gaussian noise p (that is, the first Gaussian noise value mentioned above).
[0130] It is understandable that the reverberation prediction module 103 to the gradient calculation module 105 are used for an example of executing S102. For example, it can be a more specific embodiment of steps 1-a to 1-c in Example 1. The reverberation prediction module 103 is used for a more specific embodiment in Example 5. The noise prediction module 104 is used for a more specific embodiment in Example 3 or Example 4.
[0131] The sampling module 106 is used to perform sampling based on the gradient calculated by the gradient calculation module 105 to obtain a newly sampled speech signal. Specifically, reference can be made to steps 2-a and 2-b above.
[0132] The sampling method includes, but is not limited to, ancestral sampling, Langevin dynamics sampling, and annealed Langevin dynamics sampling (a combination of one or two). However, regardless of which sampling method is used, a first random number needs to be generated based on a random seed, and then sampling is performed based on the first random number.
[0133] Exemplarily, sampling based on the gradient calculated by the gradient calculation module 105 includes two sub-steps: prediction and correction. Among them, the prediction step uses ancestral sampling; the correction step uses Langevin dynamics sampling or annealed Langevin dynamics sampling, and the correction step is used to correct the prediction result of the prediction step.
[0134] For example, in the prediction step, the prediction formula is where is the speech signal obtained by the nth sampling, grad is the gradient, and z is the second Gaussian noise value that follows a normal distribution with a mean of zero and a variance of 1. It should be noted that setting a random seed is required to generate the second Gaussian noise value. For evidentiary purposes, the random seed is set when the terminal device is reset. -f(x t ,X(f,t))+g 2 (t)grad is the inverse drift coefficient, and g(t) is the inverse diffusion coefficient. Among them, in f(x t ,X(f,t)), the f in f(*) is the drift coefficient, and g(t) is the inverse diffusion coefficient. The drift coefficient and the diffusion coefficient are determined when designing the forward stochastic differential equation. The drift coefficient where s is the sampling time step, s is initialized to 1, and for the nth sampling, s = max(1 - n / N, 0.003). The diffusion coefficient g(t) = 2 t .
[0135] For example, in the correction step, the annealed Langevin dynamics sampling method is used to correct the prediction result.
[0136] It can be understood that the sampling module 106 is used to execute an example of S103. For example, it can be a more specific embodiment of steps 2-a and 2-b in Example 1 or Example 2.
[0137] The judgment module 107 is used to judge whether the number of sampling steps has reached the preset total number of sampling steps N. If not, the sampled speech signal is input to the noise prediction module 104 for processing; otherwise, the sampled speech signal is input to the feature inverse transformation module 108.
[0138] The feature inverse transformation module 108 is used to perform feature inverse transformation on the sampled speech signal, such as inverse short-time Fourier transform, to obtain the dereverberated speech signal in the time domain.
[0139] It can be understood that the above-mentioned noise prediction module 104 to judgment module 107 constitute a conditional diffusion model, and the condition of this conditional diffusion model is the reverberant speech signal.
[0140] Figure 6 The figure shows a schematic diagram of the training process of the first neural network model provided by the embodiment of the present application.
[0141] The reverberation prediction model (the first neural network model) needs to be trained before the Figure 5 embodiment shown is executed, and it includes the following several modules.
[0142] The acoustic impulse response module 201 is used to obtain the acoustic impulse response from the acoustic impulse response data set. This data set can be generated by simulation or actually recorded.
[0143] The clean speech signal module 202 is used to obtain the clean speech signal from the speech data set, and this speech data set can be pre-recorded.
[0144] The convolution module 203 is used to convolve the acoustic impulse response and the clean speech signal to obtain the reverberant speech signal, and this signal is a signal in the time domain.
[0145] The feature transformation module 204: performs feature transformation on the reverberant signal to obtain the sampled reverberant speech signal in the time-frequency domain after feature transformation. For example, this feature transformation is short-time Fourier transform, amplitude transformation, etc.
[0146] The first neural network 205: inputs the sampled speech signal into the deep neural network model. There are various choices for the deep neural network model here, such as convolutional neural networks (CNN), long short-term memory (LSTM), U-Net and other network models.
[0147] The loss function calculation module 206: calculates the loss function using the output signal of the network model and the clean speech signal. For example, the loss function can choose mean square loss function, absolute value loss function, signal-to-noise ratio loss function, and signal-to-interference ratio loss function, etc.
[0148] Figure 7 The figure shows a schematic diagram of the training process of the second neural network model provided by the embodiment of the present application.
[0149] The training process block diagram of the noise prediction model (the second neural network model) includes the following modules.
[0150] For the acoustic impulse response module 301 to the feature transformation module 304, reference can be made to the descriptions of the acoustic impulse response module 201 to the feature transformation module 204.
[0151] The sampling time generation module 305, which requires no input information. The function of this module is to randomly generate a decimal number between t min and t max as the sampling time, where t min is the minimum value, such as 0.003, and t max is the maximum value, such as 1. The sampling time generated by this module is output to the sample generation module.
[0152] The Gaussian noise generation module 306 is used to generate the fourth Gaussian noise value h with a mean of zero and a variance of 1.
[0153] The sample generation module 307 is used to generate samples based on the sampling time according to x = μ + σh, where μ is the mean of h and σ is the standard deviation of h (where the mean and standard deviation both depend on the sampling time). Various calculation methods can be adopted.
[0154] The sample generation module 307 can control the maximum and minimum values (or amplitude) of the variance curve and the inflection point of the variance curve. On the one hand, it can bring performance improvement (that is, the spectrogram is restored better and the speech quality is improved), and on the other hand, it can bring compression of the sampling steps (reduce the sampling steps).
[0155] The second neural network 308 inputs the reverberant sample speech signal (that is, the signal after feature transformation), the sample sampling time (that is, the sampling time generated by the sampling time generation module 305), and the reverberant sample sampling signal (that is, the sample generated above) into the noise prediction model. The noise prediction model can be composed of various deep neural networks, such as U-Net network, CNN network, convolutional recurrent neural network (CRN) network, etc.
[0156] The loss function calculation module 309 calculates the loss function using the predicted noise output by the noise prediction model, the clean speech signal after feature transformation, and the generated samples.
[0157] Specifically, the clean speech signal and the generated samples are subtracted to obtain the noise signal as the target data, and the loss function is calculated with the predicted noise output.
[0158] For example, the loss function here can be calculated according to criteria such as mean square error, signal-to-noise ratio, and signal loss ratio.
[0159] It can be understood that taking the implementation of the conditional diffusion model by training a neural network model as an example, the reverberant sample speech signal, the current sampling time t, and the sample x generated at the current sampling time t are input into the second neural network model to predict x t+1 relative to the noise superimposed on x (or it can also be understood as the reverberation component), and the signal obtained by subtracting x t from the corresponding clean speech signal is used as the target data for loss function calculation. t Thus, after the second neural network is trained, when a reverberant speech signal is input, the output Gaussian noise p can also be understood as the predicted reverberation component in the reverberant speech signal.
[0160] Please refer to
[0161] which shows a schematic hardware structure diagram of an electronic device 1000 provided by an embodiment of the present application. Refer to Figure 8 Figure 8 which shows a schematic hardware structure diagram of an electronic device 1000 provided by an embodiment of the present application. See Figure 8 The electronic device 1000 may include a processor 110, an audio module 120, a microphone 120A, and optionally, a speaker 120B.
[0162] It can be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0163] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc.
[0164] Among them, the controller can be the nerve center and command center of the electronic device 1000. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0165] A memory can also be set in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0166] In this application, the processor 110 is configured to obtain a first reverberant voice signal; preprocess the first voice signal to obtain the probability distribution of the first voice signal; and perform sampling processing on the first voice signal according to the probability distribution of the first voice signal to obtain a target voice signal after reverberation removal. It can perform reverberation removal processing on the first reverberant voice signal to obtain a target voice signal after reverberation removal, improve the listening quality of the voice during a voice call, and the success rate of speech recognition, thereby enhancing the user experience. Compared with the discriminant reverberation removal method, it can ensure the spectral continuity of the reverberation-removed voice signal and reduce the damage to the voice signal; compared with the reverberation removal method that relies on linear assumptions, it can process a wider range of reverberant voice signals.
[0167] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0168] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0169] The electronic device 1000 can implement audio functions such as music playback, recording, and voice calls through the audio module 120, the speaker 120B, the microphone 120A, and an application processor (not shown in the figure), etc.
[0170] In a voice call, recording, or voice recognition scenario, the microphone 120A is used to record reverberant voice signals.
[0171] Optionally, during a voice call, the speaker 120B is used to play the voice of the party on the other side of the call with the user; in a recording scenario, if the user wants to audition the recorded content, the speaker 120B plays the recorded content; in a voice recognition scenario, after the electronic device recognizes the voice signal of the user recorded by the microphone 120A, it transmits the feedback voice signal to the user through the speaker 120B.
[0172] Next, the software system of the electronic device 1000 will be described.
[0173] Exemplarily, the electronic device 1000 may be a mobile phone. The software system of the electronic device 1000 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily describe the software system of the electronic device 1000.
[0174] Figure 9 Shows a block diagram of a software system of an electronic device 1000 provided by an embodiment of the present application. Refer to Figure 9 , the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom are the application layer (application layer), the application framework layer (framework layer), the hardware abstraction layer (hardware abstraction layer, HAL), the driver layer, and the hardware layer.
[0175] The application layer may include a series of application packages, such as a dialing application, a gallery application, etc. (not shown in the figure). In the embodiments of the present application, the application package may include applications such as recording, voice call, voice assistant, etc. Such applications all need to use a microphone to record audio and need to perform reverberation removal processing on the audio recorded by the microphone. Among them, the voice assistant is used to recognize the user's voice and interact with the user.
[0176] Alternatively, the application layer may also include other applications that need to perform reverberation removal processing on the voice signals recorded by the microphone, which is not limited in the present application.
[0177] The framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. In the embodiments of the present application, the framework layer includes a microphone service interface and a reverberation removal service interface. Among them, the reverberation removal service interface can provide APIs and programming frameworks for applications that obtain the reverberation removal service. The microphone service can be used to provide APIs and programming frameworks for applications that call the microphone.
[0178] The hardware abstraction layer (HAL) is an interface layer located between the operating system kernel and the upper-layer software, providing a virtual hardware platform for the operating system. In the embodiments of the present application, the hardware abstraction layer may include a microphone hardware abstraction layer and a reverberation removal algorithm. The microphone hardware abstraction layer can provide virtual hardware for microphone 1, microphone 2, or more microphone devices. The reverberation removal algorithm may include the running code and data for implementing the voice signal processing method provided in the embodiments of the present application.
[0179] The driver layer is the layer between the hardware and the software. The driver layer includes drivers for various hardware. The driver layer may include a microphone device driver, a digital signal processor driver, etc. The microphone device driver is used to drive the microphone sensor to collect sound signals and drive the audio signal processor to preprocess the sound signals to obtain audio digital signals. The digital signal processor driver is used to drive the digital signal processor to process the audio digital signals.
[0180] The hardware layer includes sensors and an audio signal processor. Among them, the sensors include microphone 1 and microphone 2. The microphones included in the sensors correspond one-to-one with the virtual microphones included in the microphone hardware abstraction layer. The audio signal processor can be used to convert the sound signals collected by the microphones into audio digital signals. The digital signal processor can be used to process the audio digital signals. It should be noted that what the present application provides Figure 9The schematic diagram of the software structure of the electronic device shown is only an example and does not limit the specific module division in different layers of the Android operating system. For details, reference can be made to the introduction to the software structure of the Android operating system in the conventional technology.
[0181] Next, in combination with the above-mentioned hardware structure and system structure, the method in the embodiments of the present application will be specifically described:
[0182] In response to enabling applications such as recording, voice call, or voice recognition, these applications can call the dereverberation service interface to obtain the application programming interface and programming framework provided by the dereverberation service.
[0183] On the one hand, the dereverberation service can call the microphone service in the framework layer to collect the sound signal in the environment through the microphone service. Among them, the microphone service can send an instruction to collect the sound signal to the microphone 1 in the microphone hardware abstraction layer by calling the microphone 1, and the microphone hardware abstraction layer sends the instruction to the microphone device driver in the driver layer. The microphone device driver can start the microphone 1 according to the above instruction, so as to obtain the sound signal in the environment and generate a digital audio signal through the audio signal processor.
[0184] On the other hand, the dereverberation service can initialize the dereverberation algorithm. The dereverberation algorithm can obtain the digital audio signal generated by the audio signal processor through the microphone hardware abstraction layer. Then, according to the voice signal processing method stored in the dereverberation algorithm, the dereverberation algorithm can use the digital signal processor to process the obtained digital audio signal to obtain the dereverberated digital audio signal.
[0185] Specifically, how to process the digital audio signal to obtain the dereverberated digital audio signal can be referred to the method flow chart shown above. Figures 4 to 7 shown.
[0186] Finally, the dereverberation algorithm can transmit the dereverberated digital audio signal back to the dereverberation service, and then back to the application layer.
[0187] The embodiments of the present application provide a chip system, which includes one or more processors for calling and running the instructions stored in the memory from the memory, so that the methods in the above embodiments of the present application are executed. The chip system can be composed of chips or can include chips and other discrete devices.
[0188] Among them, the chip system can include an input circuit or interface for sending information or data, and an output circuit or interface for receiving information or data.
[0189] The present application also provides a computer program product which, when executed by a processor, implements the method described in any of the method embodiments of the present application.
[0190] This computer program product can be stored in a memory and is finally converted into an executable target file that can be executed by a processor through processes such as preprocessing, compilation, assembly, and linking.
[0191] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements the method described in any of the method embodiments of the present application. This computer program can be a high-level language program or an executable target program.
[0192] This computer-readable storage medium can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0193] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the technical effects generated by the above-described devices and equipment can refer to the corresponding processes and technical effects in the foregoing method embodiments, and will not be elaborated herein.
[0194] In several embodiments provided in this application, the disclosed systems, devices, and methods can be implemented in other ways. For example, some features of the above-described method embodiments can be ignored or not executed. The device embodiments described above are merely illustrative. The division of units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system. In addition, the coupling between units or the coupling between each component can be direct coupling or indirect coupling. The above coupling includes electrical, mechanical, or other forms of connection.
[0195] It should be understood that in various embodiments of this application, the magnitude of the sequence numbers of each process does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0196] It should be understood that the "multiple" mentioned in this application refers to two or more. The term " / and" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0197] Terms (or numbers) such as "first", "second",... that appear in the embodiments of this application are only for descriptive purposes, that is, only to distinguish different objects, such as different "coordinates", etc., and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second",... can explicitly or implicitly include one or more features. In the description of the embodiments of this application, "at least one (item)" means one or more. The meaning of "multiple" is two or more. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of a single (item) or multiple (items).
[0198] In summary, the above description is only a preferred embodiment of the technical solution of this application, and is not used to limit the protection scope of this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A voice signal processing method, applied to an electronic device, characterized in that, The method includes: Obtaining a first speech signal to be processed; Performing preprocessing on the first speech signal to obtain a probability distribution of the first speech signal, where the probability distribution of the first speech signal corresponds to the distribution of time-frequency points in the spectrogram of the first speech signal; Performing sampling processing on the first speech signal according to the probability distribution of the first speech signal to obtain the first speech signal after reverberation removal, and using the first speech signal after reverberation removal as the target speech signal.
2. The method according to claim 1, wherein The performing preprocessing on the first speech signal to obtain the probability distribution of the first speech signal includes: predicting a second speech signal according to the first speech signal, where the second speech signal is an estimated clean speech signal; predicting the nth first Gaussian noise value according to the first speech signal; generating the nth gradient according to the second speech signal, the nth first Gaussian noise value, and the time step of the nth sampling; The performing sampling processing according to the probability distribution of the first speech signal includes: in the case of n = 1, sampling the first speech signal by using the first gradient to obtain the first sampling signal of the first speech signal, where the first gradient is used to characterize the probability distribution of the first speech signal; or, in the case of n ≥ 2, sampling the (n - 1)th sampling signal of the first speech signal by using the nth gradient to obtain the nth sampling signal of the first speech signal, where the nth gradient is used to characterize the probability distribution of the (n - 1)th sampling signal, and n is a positive integer.
3. The method according to claim 2, characterized in that, The sampling the first speech signal by using the first gradient to obtain the first sampling signal of the first speech signal includes: Generating the first second Gaussian noise value according to a preset first random seed; Predicting the first sampling signal of the first speech signal according to the first second Gaussian noise value, the first gradient, and the first speech signal.
4. The method according to claim 3, characterized in that, The sampling the (n - 1)th sampling signal of the first speech signal by using the nth gradient to obtain the nth sampling signal of the first speech signal includes: Generating the nth second Gaussian noise value according to the preset first random seed; Predicting the speech signal obtained by the nth sampling according to the nth second Gaussian noise value, the nth gradient, and the speech signal obtained by the (n - 1)th sampling.
5. The method according to claim 2, wherein The predicting the nth first Gaussian noise value according to the first speech signal includes: In the case of n = 1, obtaining the first first Gaussian noise value according to the first speech signal, the time step of the first sampling, and a third speech signal, where the third speech signal is generated according to the first speech signal, a third Gaussian noise value, and the time step of the first sampling, and the third Gaussian noise value is generated according to a preset second random seed; or, When n≥2, input the first voice signal, the time step of the nth sampling, and the (n - 1)th sampled signal into a noise prediction module to obtain the nth first Gaussian noise value.
6. The method according to claim 5, characterized in that, The noise prediction module is a first neural network model trained based on first input data and first target data. The first input data includes a reverberant sample voice signal, a sample sampling time, and a reverberant sample sampled signal. The reverberant sample sampled signal is generated according to the sample voice signal, a fourth Gaussian noise value generated based on a preset third random seed, and the sample sampling time. The first target data includes a sample clean voice and a sample noise obtained by subtracting the sample sampled signal. The sample clean voice is used to generate the reverberant sample voice signal.
7. The method according to claim 2, characterized in that, Predicting a second voice signal according to the first voice signal includes: Input the first voice signal into a reverberation prediction module to obtain the second voice signal. Among them, the reverberation prediction module is a second neural network model trained based on second input data and second target data. The second input data includes a reverberant sample voice signal, and the second target data includes a clean voice signal corresponding to the reverberant sample voice signal.
8. The method according to claim 7, wherein The sample voice signal is a time-frequency domain voice signal obtained based on feature transformation. The feature transformation includes short-time Fourier transform and amplitude transformation.
9. An electronic device, characterized in that, The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors. The memory is used to store computer program code. The computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
10. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes one or more processors. The one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions. When the instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.