Voice activity detection method and related equipment

The time frame signal gradient and diffusion model prediction of the audio signal are calculated through a neural network to generate a voice activity sequence, which solves the impact of signal-to-noise ratio changes on voice activity detection in the existing technology and achieves stable and accurate detection under various signal-to-noise ratio conditions.

CN120612967APending Publication Date: 2025-09-09HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410235157.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing voice activity detection methods are unstable in conditions with high or low signal-to-noise ratios. In particular, in extremely low signal-to-noise ratios, it is difficult to accurately distinguish between speech and noise, which affects the processing effect of subsequent audio algorithms.

Method used

A neural network-based method is used to characterize the probability distribution of speech signals by calculating the time frame signal gradient of audio signals, and a diffusion model is used for prediction and correction to generate speech activity sequences that are independent of the influence of noise signals.

Benefits of technology

It can stably and accurately detect voice activity under various signal-to-noise ratio conditions, improving the accuracy of subsequent audio algorithm processing results and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612967A_ABST
    Figure CN120612967A_ABST
Patent Text Reader

Abstract

The invention provides a voice activity detection method and related equipment. The method comprises the following steps: acquiring an audio signal to be processed, and determining a voice activity sequence of the audio signal based on a diffusion model; wherein the processing of the audio signal based on the diffusion model comprises the following steps: obtaining a time frame signal in a sequence form according to audio features, calculating a gradient of the time frame signal based on a neural network model, predicting a sampling signal according to the gradient, and obtaining a voice activity sequence for indicating whether a voice signal exists in each time frame based on the sampling signal. Compared with a related detection method, the method focuses on probability distribution of voice signals, is not affected by the signal-to-noise ratio, and can work normally under the extremely low signal-to-noise ratio; noise signals are not concerned, influence of noise types is avoided, and generalization performance is better. Therefore, an accurate voice activity detection result can be stably obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice detection, and in particular to a voice activity detection method and related equipment. Background Art

[0002] Voice activity detection (VAD) is used to detect speech in an audio signal (including whether speech is present, or which time periods contain speech and which do not), thereby distinguishing speech from non-speech segments. VAD is a prerequisite for many audio algorithms, such as speech enhancement, speech recognition, and voice wake-up algorithms, and plays a vital role in audio applications.

[0003] Currently, one implementation method is to use the zero-crossing rate, energy, and other features of the speech signal to perform VAD. This method can achieve good results when the signal-to-noise ratio is high, but the effect is poor when the signal-to-noise ratio is low, and may even fail to work properly. In addition, when using energy features for VAD, if the energy of the speech signal is unstable, the effect will also deteriorate. Another implementation method is to use a discriminant neural network to perform VAD. This method uses a training data set that includes clean speech data and noise data. The discriminant neural network learns the mapping between the audio signal and the speech activity detection results based on the training data set. It can achieve good results in noise conditions similar to the noise data in the training set, but the effect is poor in noise conditions not included in the training set, that is, the generalization is poor. Summary of the Invention

[0004] The present application provides a voice activity detection method and related devices, which can stably obtain accurate voice activity detection results.

[0005] In a first aspect, a voice activity detection method is provided, which is applied to an electronic device, and the method includes: obtaining an audio signal to be processed, the audio features of the audio signal including a time-frequency spectrum; obtaining a time frame signal based on the audio signal, the time frame signal being a sequence corresponding to the time frames of the time-frequency spectrum; calculating a gradient of the time frame signal based on the audio signal and the time frame signal based on a neural network model, the gradient of the time frame signal being used to characterize the probability distribution of a voice signal included in the audio signal; and predicting a sampling signal of the time frame signal based on the gradient of the time frame signal.

[0006] The voice activity detection method provided in the present application has a very stable VAD effect and accurate VAD results under high or low signal-to-noise ratio conditions; therefore, the voice activity detection method provided in the present application can stably obtain accurate voice activity detection results.

[0007] Therefore, thanks to the voice activity detection method provided in this application that stably provides accurate VAD results, subsequent audio algorithms (such as speech recognition, voice wake-up, etc.) can stably obtain more accurate processing results, greatly improving the user experience.

[0008] In other words, the voice activity detection method provided by this application uses a neural network to obtain the distribution of voice signals in an audio signal, outputting a gradient representing the probability distribution of the voice signal, and then generating a voice activity sequence based on the predicted sampled signal based on this gradient. Therefore, the voice activity detection method provided by this application does not focus on the noise signal in the audio signal, nor does it distinguish between the noise signal and the voice signal. Instead, it focuses on the probability distribution of the voice signal. As a result, the voice activity detection method provided by this application is not affected by the signal-to-noise ratio or the type of noise. Furthermore, the voice activity detection method provided by this application can accurately detect voice activity even in extremely low signal-to-noise ratio conditions.

[0009] It can be understood that by inputting a sequence of time frame signals and the audio signal to be processed into a neural network model, a probability distribution of the speech signal can be output. The sampled signal obtained by gradient prediction of the time frame signal (or sampling the time frame signal) can be understood as being able to represent the probability distribution of the speech signal. The probability distribution of the speech signal represented by the sampled signal corresponds to fewer time-frequency points than the probability distribution of the speech signal represented by the gradient of the time frame signal. A smaller number of sample points is used to approximate the overall distribution, thereby achieving noise reduction.

[0010] It can also be understood that compared with the traditional VAD method, the voice activity detection method provided by the present application is not affected by the signal-to-noise ratio and is more stable.

[0011] It can also be understood that compared to the artificial intelligence VAD method, the voice activity detection method provided in this application does not learn the noise signal, nor does it learn the difference between the noise signal and the voice signal. Therefore, it is not affected by the noise type and has better generalization performance.

[0012] In a possible embodiment, a speech activity sequence of an audio signal to be processed is obtained based on a sampling signal, including: calculating the gradient of the sampling signal based on a neural network model, where the gradient of the sampling signal is used to characterize the probability distribution of the speech signal; correcting the sampling signal based on the gradient of the sampling signal to obtain a corrected signal; and obtaining a speech activity sequence based on the corrected signal.

[0013] It can be understood that calculating the gradient based on the predicted signal and then performing correction can improve the accuracy of the prediction, thereby improving the accuracy of the final speech activity sequence.

[0014] In a possible embodiment, the neural network model is trained based on a training data set and a loss function value, the training data set includes a sample sequence and a sample audio signal, the sample audio signal includes a sample speech signal, the sample sequence is generated based on the sample audio signal, and the loss function value is used to characterize the probability distribution of the sample speech signal.

[0015] That is to say, by setting the loss function, the neural network learns the probability distribution of the sample signal in the sample audio during the training process, so that on the inference side of the neural network, the probability distribution of the speech signal in the audio signal can be obtained based on the input audio signal and the sequence signal.

[0016] In a possible embodiment, the training data set also includes sample audio features of the sample audio signal, the sample audio features include a sample time-frequency spectrum, and the gradient of the time frame signal is calculated based on the neural network model, including: inputting the time frame signal, the audio signal and the audio features into the neural network model to obtain the gradient of the time frame signal, wherein the sample sequence is a first Gaussian noise value, the first Gaussian noise value is generated based on the sample speech signal, the sample audio signal and a second Gaussian noise value that obeys a standard normal distribution, the loss function value is generated based on the variance of the first Gaussian noise value and the second Gaussian noise value, and the second Gaussian noise value is generated based on the first random seed.

[0017] In the above scheme, the time-frequency spectrum is also used as training data during the training process, so that the neural network can more accurately obtain the probability distribution of the speech signal, thereby improving the accuracy of the neural network finally trained.

[0018] The first random seed is a preset value determined through a large number of experiments, which can achieve better voice activity detection results.

[0019] It is understandable that, because the random number generated by the random seed at different times is different, the method provided in the present application processes the audio signal at different times, and the voice activity detection sequences outputted at the respective times are different or not completely the same.

[0020] In a possible embodiment, the mean value of the first Gaussian noise value is generated according to a sample voice activity sequence and a sample audio signal, where the sample voice activity sequence is a result of voice activity detection generated based on the sample voice.

[0021] It can be understood that using a sample speech activity sequence to generate a sample sequence can make the sample sequence generated more accurate, thereby reducing the number of training steps of the neural network and accelerating the convergence of the neural network.

[0022] In a possible embodiment, the variance of the first Gaussian noise value is generated according to the diffusion intensity, and the diffusion intensity is determined based on a stochastic differential equation.

[0023] It can be understood that the diffusion strength can adjust the randomness of the first Gaussian noise value, and by setting the diffusion strength value, the neural network can be adjusted to achieve better training effects.

[0024] In a possible embodiment, the sample audio features also include: sample time domain features, and the audio features also include: time domain features; wherein the sample time domain features include sample zero-crossing rate features, and the time domain features include zero-crossing rate features; and / or, the sample time domain features include sample energy features, and the time domain features include energy features.

[0025] It is understandable that although these features have little effect on VAD under low signal-to-noise ratio conditions, when the signal-to-noise ratio is high, VAD based on these features is very effective and can improve the training efficiency of the neural network.

[0026] That is to say, the voice activity detection method provided in the present application is not only applicable to situations with low signal-to-noise ratios, but also to situations with high signal-to-noise ratios. During the model training process, when the signal-to-noise ratio is low, the zero-crossing rate and energy have a weak effect on the training of the neural network, and mainly rely on the sample sequence and sample time-frequency spectrum for training; while when the signal-to-noise ratio is high, the zero-crossing rate and energy have a strong effect on the training of the neural network, and can assist the generated sample sequence and sample time-frequency spectrum in training the neural network, thereby improving the efficiency of neural network training. Correspondingly, during the model inference process, when the signal-to-noise ratio is low, the zero-crossing rate and energy have a weak effect on VAD, and mainly rely on the time frame signal and time-frequency spectrum for inference; while when the signal-to-noise ratio is high, the effect of VAD based on zero-crossing rate and energy is very good, and can effectively assist other information in performing VAD on the audio signal, thereby improving the efficiency and accuracy of VAD.

[0027] In a possible embodiment, a time frame signal is obtained based on an audio signal, including: obtaining a first time frame signal based on the audio signal; calculating the gradient of the time frame signal based on a neural network model, including: calculating the gradient of the first time frame signal based on the neural network model, the gradient of the first time frame signal is used to characterize the probability distribution of the speech signal; predicting a sampling signal of the time frame signal based on the gradient of the time frame signal, including: predicting the first sampling signal based on the gradient of the first time frame signal.

[0028] Exemplarily, one or more audio features of the audio signal are used as the first time frame signal.

[0029] In a possible embodiment, a time frame signal is obtained according to an audio signal, including: obtaining an nth time frame signal according to an n-1th sampling signal, wherein the n-1th sampling signal is obtained according to the audio signal, 2≤n≤N, and N and n are both integers; calculating the gradient of the time frame signal based on a neural network model, including: calculating the gradient of the nth time frame signal based on the neural network model, the gradient of the nth time frame signal is used to characterize the probability distribution of the speech signal; predicting the sampling signal of the time frame signal according to the gradient of the time frame signal, including: predicting the nth sampling signal according to the gradient of the nth time frame signal.

[0030] That is to say, the above steps of calculating gradients and predictions need to be iterated N times in order to obtain more accurate results.

[0031] In a possible embodiment, obtaining a voice activity sequence of the audio signal to be processed according to the sampling signal includes: when n=N, obtaining a voice activity sequence according to the Nth sampling signal.

[0032] That is to say, the sampled signal obtained in the Nth iteration can be used as a speech activity sequence when the diffusion model does not include a correction process.

[0033] In a possible embodiment, obtaining an n-th time frame signal based on the n-1th sampling signal includes: calculating the gradient of the n-1th sampling signal, where the gradient of the n-1th sampling signal is used to characterize the probability distribution of the speech signal; correcting the n-1th sampling signal based on the gradient of the n-1th sampling signal to obtain the n-1th corrected signal, and using the n-1th corrected signal as the n-1th time frame signal.

[0034] That is, when the diffusion model includes a correction process, the correction step also needs to be performed N times, and the correction signal is used as the time frame signal of the next iteration.

[0035] In a possible embodiment, a voice activity sequence is obtained based on the Nth sampling signal, including: calculating the gradient of the Nth sampling signal, the gradient of the Nth sampling signal is used to characterize the probability distribution of the voice signal; correcting the Nth sampling signal based on the gradient of the Nth sampling signal to obtain the Nth corrected signal, and using the Nth corrected signal as the voice activity sequence.

[0036] That is, the correction signal obtained in the Nth iteration process is used as the speech activity sequence.

[0037] In a possible embodiment, a sampling signal of the time frame signal is predicted according to the gradient of the time frame signal, including: obtaining a drift coefficient according to the time frame signal based on a stochastic differential equation of a conditional diffusion model, wherein the condition of the conditional diffusion model is an audio signal; obtaining an inverse drift coefficient according to the drift coefficient and the gradient of the time frame signal; predicting the sampling signal according to the time frame signal, the inverse drift coefficient and a third Gaussian noise value, wherein the third Gaussian noise value is generated based on a second random seed.

[0038] In a possible embodiment, the first sampling signal is predicted according to the gradient of the first time frame signal, including: obtaining a first drift coefficient according to the first time frame signal based on a stochastic differential equation of a conditional diffusion model, wherein the condition of the conditional diffusion model is an audio signal; obtaining a first inverse drift coefficient according to the first drift coefficient and the gradient of the first time frame signal; predicting the first sampling signal according to the first time frame signal, the first inverse drift coefficient and the first third Gaussian noise value, wherein the first third Gaussian noise value is generated based on the first second random seed.

[0039] In a possible embodiment, the nth sampling signal is predicted according to the gradient of the nth time frame signal, including: obtaining the nth drift coefficient according to the nth time frame signal and the first time frame signal based on a stochastic differential equation of a conditional diffusion model, wherein the condition of the conditional diffusion model is the audio signal; obtaining the nth inverse drift coefficient according to the nth drift coefficient and the gradient of the nth time frame signal; predicting the nth sampling signal according to the nth time frame signal, the nth inverse drift coefficient and the nth third Gaussian noise value, wherein the nth third Gaussian noise value is generated based on the nth second random seed.

[0040] In a second aspect, the present application provides an electronic device comprising one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code comprises computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation of the first aspect.

[0041] In a third aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method described in the first aspect and any possible implementation method of the first aspect.

[0042] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0043] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0044] It is understandable that the electronic device provided in the second aspect, the chip system provided in the third aspect, the computer storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to perform the methods provided in this application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1A A schematic diagram showing an example of a scenario to which the embodiments of the present application are applicable;

[0046] Figure 1B A schematic diagram showing another example of a scenario to which the embodiments of the present application are applicable;

[0047] Figure 2A A schematic diagram showing an example of the result of performing VAD on audio signal #1 provided in an embodiment of the present application;

[0048] Figure 2B A schematic diagram showing an example of the result of performing VAD on audio signal #2 using the voice activity detection method provided by an embodiment of the present application;

[0049] Figure 3 2 is a schematic diagram illustrating a voice activity detection method 200 provided in an embodiment of the present application;

[0050] Figure 4 1 is a schematic diagram showing a voice activity detection method 300 provided in an embodiment of the present application;

[0051] Figure 5 4 is a schematic diagram illustrating a voice activity detection method 400 provided in an embodiment of the present application;

[0052] Figure 6 1 is a schematic diagram illustrating a voice activity detection method 500 provided in an embodiment of the present application;

[0053] Figure 7 1 shows a hardware structure diagram of an electronic device 1000 provided in an embodiment of the present application;

[0054] Figure 8A block diagram of a software system of an electronic device 1000 provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0055] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0056] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are first introduced.

[0057] 1. An audio signal may contain human speech and / or non-human speech. The portion of the audio signal containing human speech is called a speech signal, and the portion not containing human speech is called a noise signal.

[0058] The signal to interference plus noise ratio (SNR) refers to the ratio of the strength of the received useful signal (such as the voice signal in this application) to the strength of the received interference signal (including noise and interference, such as the noise signal in this application); it can be simply understood as the "signal-to-noise ratio".

[0059] The lower the signal-to-noise ratio, the greater the interference of the noise signal on the speech signal.

[0060] Ultra-low signal-to-noise ratio or extremely low signal-to-noise ratio can be understood as the noise signal significantly interfering with the speech signal, for example, signal-to-noise ratio <-5dB, or signal-to-noise ratio <-10dB, or other values ​​indicating an ultra-low or extremely low signal-to-noise ratio.

[0061] 2. VAD is used to detect speech in an audio signal (including whether speech is present, or which time periods contain speech and which do not), thereby distinguishing speech from non-speech segments. VAD is a prerequisite for many audio algorithms, such as speech enhancement, speech recognition, and voice wake-up algorithms, and plays a vital role in audio applications.

[0062] It can be understood that, assuming that an audio signal includes a speech segment and a non-speech segment, the VAD algorithm performs VAD processing on the audio signal, so that the subsequent audio algorithm can focus on processing the speech segment (optionally, without processing the non-speech segment), which can effectively reduce the power consumption of the electronic device and improve the processing efficiency of the audio algorithm.

[0063] 3. The time domain characteristics of audio signals include: zero-crossing rate, energy fluctuation rate, energy peak, etc.; the frequency domain characteristics of audio signals include: fundamental frequency, harmonics, spectrum centroid, spectrum band characteristics, etc.

[0064] The zero-crossing rate (ZCR) refers to the rate at which a signal changes sign, for example, from positive to negative, or vice versa. The short-term zero-crossing rate indicates the number of times the speech signal waveform crosses the horizontal axis (zero level) in a frame of speech. For continuous speech signals, a zero crossing occurs when the time domain waveform crosses the time axis; for discrete signals, a zero crossing occurs when adjacent sample values ​​change sign. Therefore, the zero-crossing rate is the number of times a sample changes sign.

[0065] In this application, the time-frequency domain features of the audio signal are time-frequency diagrams, also called spectrograms, where the horizontal axis is the time domain (for example, in time frames) and the vertical axis is the frequency domain.

[0066] In the present application, the audio features of the audio signal to be processed include time-frequency domain features, and may optionally include time-domain features and / or frequency-domain features. In the embodiments of the present application, the audio features of the audio signal to be processed are mainly described by taking as an example the audio features including time-frequency spectrum, zero-crossing rate, and energy, but the scope of protection of the audio features is not limited. For example, the audio features may also include fundamental frequency.

[0067] 4. The probability distribution of the speech signal in this application can be characterized by the gradient of the sampled signal or by the gradient of the time frame signal.

[0068] The time frame signal may be a sequence corresponding to the time frames of the time spectrum of the audio signal to be processed. That is, the time frame signal is in the form of a sequence and is arranged in sequence according to the order of the time frames of the time spectrum; that is, the sequence includes a plurality of sequentially arranged numbers, each of which corresponds to a time frame in sequence.

[0069] The sampled signal is a signal obtained by sampling (including prediction and optionally correction) the time frame signal. Accordingly, the sampled signal is also a signal in sequence form and also corresponds to the time frame of the above-mentioned time spectrum.

[0070] In other words, the sampling signal and the time frame signal can be understood as sequences related to the speech signal. The gradient of the sequence related to the speech signal is then used to characterize the probability distribution of the speech signal. The probability distribution of the speech signal corresponds to the distribution of time-frequency points on the speech signal's time-frequency graph.

[0071] For example, the gradient of the time frame signal can be understood as the distribution of the values ​​of the numbers included in the time frame signal sequence, or as the characteristics of the time frame signal (including time domain characteristics, frequency domain characteristics, or time-frequency domain characteristics, etc.).

[0072] For another example, the gradient of the sampling signal can be understood as the distribution of the values ​​of the numbers included in the sequence of the sampling signal, or as the characteristics of the sampling signal (including time domain characteristics, frequency domain characteristics, or time-frequency domain characteristics, etc.).

[0073] 5. Voice activity sequence: used to indicate whether there is a voice signal in each time frame of the time-spectrum; in other words, used to characterize or indicate which time frames in the audio signal to be processed have voice signals, and / or which frames do not have voice signals (or are noise signals).

[0074] For example, the value of the number in the sequence corresponding to the time frame with a voice signal is 1, and the value of the number in the sequence corresponding to the time frame without a voice signal is 0 (or other values ​​can also be used, such as the value corresponding to the voice signal is 0, and the value corresponding to the no voice signal is 1).

[0075] The speech activity sequence distribution can be understood as the probability distribution of numbers with a value of 1 and numbers with a value of 0 in the speech activity sequence, or can be understood as the probability distribution of time-frequency points indicating speech signals in audio signals in each time frame.

[0076] In other words, the speech activity sequence is also a sequence corresponding to the time frames of the time spectrum. It is understandable that because speech signals have short-term stationarity, for example, speech signals within the 10-30ms range can be considered stable, allowing audio data to be framed and VAD performed on the audio data by frame (e.g., time frame). For example, framing is performed using overlapping segments with 5-20ms as a frame. This means that the previous and next frames overlap. This overlapping portion is called frame shift, and the ratio of frame shift to frame length is generally 0-0.5.

[0077] 6. Diffusion model: A type of generative model based on random processes that can be used to generate various types of data, such as images, text, and audio. Unlike traditional generative models, diffusion models do not rely on known labels or target data. Instead, they process random noise to gradually generate high-quality data from it. Therefore, diffusion models can be considered a type of unsupervised generative model.

[0078] Among diffusion models, the most common is the unconditional diffusion model, whose main task is to generate random samples of original data. However, for some specific applications, such as image restoration, image synthesis, and text-to-image generation, we may need more precise control over the generated samples, which requires the introduction of the conditional diffusion model.

[0079] An important feature of the conditional diffusion model is that it can control the generated samples by adding additional conditional information, such as specific parts of the image, text descriptions, etc. For example, in the task of text-to-image generation, we can control what elements the generated image should contain or what style it should present by adding text descriptions.

[0080] In terms of specific implementation, conditional diffusion models typically encode additional conditional information as part of the model during the training phase, and then control the generated results by decoding this conditional information during the generation phase. An important advantage of this approach is that it can provide more refined control over the generated results while maintaining the generative capabilities of the diffusion model.

[0081] The diffusion model involved in this application is a conditional diffusion model. During the training phase, the audio signal is used as a conditional input to the diffusion model for training.

[0082] 7. Drift coefficient and reverse drift coefficient.

[0083] The drift coefficient and the diffusion coefficient are used together to characterize the difference between the random signal being noisy and the audio signal being processed during the process of adding noise to the clean signal.

[0084] The drift coefficient is the mean of the random signal sampled at the nth time in the N sampling process; the diffusion coefficient is the variance of the random signal sampled at the nth time in the N sampling process.

[0085] The audio signal to be processed is denoised by sampling. The inverse drift coefficient and the inverse diffusion coefficient are used together to characterize the difference between the denoised random signal and the clean signal during the denoising process of the audio signal to be processed.

[0086] The inverse drift coefficient is the mean of the random signal sampled at the nth time in the N sampling process; the inverse diffusion coefficient is the variance of the random signal sampled at the nth time in the N sampling process.

[0087] Here, 1≤n≤N and n and N are both integers.

[0088] 8. Gaussian noise: This is a common type of random noise with a Gaussian distribution, meaning its intensity and frequency are uniformly distributed. In signal and image processing, Gaussian noise is often used to simulate various noises. The frequency response of Gaussian noise can be calculated using a Gaussian function. In the frequency domain, the spectral density function of Gaussian noise exhibits a Gaussian curve.

[0089] 9. Random seed is a computer science term for a random number generated using a true random number (seed) as its initial condition. Typical computer random numbers are pseudo-random numbers, which use a true random number (seed) as their initial condition and then use a specific algorithm to iterate and generate random numbers.

[0090] It is understandable that in the present application, the random numbers (eg, Gaussian noise values) generated based on the same random seed at different times are different.

[0091] 10. Generalization: The generalization of a neural network refers to its ability to classify and predict unknown data during the learning process. Those skilled in the art will appreciate that generalization is a crucial concept in machine learning, as machine learning algorithms are typically expected to make accurate predictions on new, unseen data, not just perform well on known training data.

[0092] At present, the traditional VAD method is to use the zero-crossing rate, energy and other characteristics of the speech signal to perform VAD. This type of method can achieve good results when the signal-to-noise ratio is high, but the effect is poor when the signal-to-noise ratio is low, and may even fail to work properly. In addition, when using energy features for VAD, if the energy of the speech signal is unstable, the effect will also deteriorate. The method of using artificial intelligence to perform VAD is to use a discriminant neural network to perform VAD. This type of method uses a training data set that includes clean speech data and noise data. The discriminant neural network learns the mapping between the audio signal and the speech activity detection results based on the training data set. It can achieve good results in noise conditions similar to the noise data in the training set, but the effect is poor in noise conditions not included in the training set, that is, the generalization is poor.

[0093] That is to say, the current VAD method cannot stably achieve accurate VAD effects when the signal-to-noise ratio is too low or the noise type is complex. Therefore, the current VAD method is not robust.

[0094] Figure 1A A schematic diagram showing an example of a scenario to which an embodiment of the present application is applicable.

[0095] In a quiet environment, the user successfully wakes up the electronic device with voice. As shown in S101, the electronic device is in the off state. The user says, "Hello YOYO", and the time domain diagram of the audio signal #1 recorded by the microphone of the electronic device is as follows: Figure 1A As shown in the lower left corner, a very clear waveform is displayed. The electronic device performs VAD on the audio signal based on the VAD algorithm, determines the voice signal in the audio signal, and processes the voice signal based on the voice wake-up algorithm to wake up the electronic device and display the interface shown in S102, including the smart assistant panel.

[0096] It is understandable that in a quiet environment, the signal-to-noise ratio of the audio signal is relatively high, and both the traditional VAD method and the VAD method using artificial intelligence can achieve good results.

[0097] Figure 2A A schematic diagram showing an example of the result of performing VAD on the audio signal #1 provided in an embodiment of the present application.

[0098] like Figure 2A As shown, Figure 1A The pulse diagram added to the time domain diagram in the figure can be understood as a visualization of the VAD results. The time domain with pulses indicates speech, and the time domain without pulses indicates no speech.

[0099] Figure 1B A schematic diagram showing another example of a scenario to which the embodiments of the present application are applicable.

[0100] In a noisy environment (loud noise, extremely low signal-to-noise ratio), the user's voice fails to wake up the electronic device. Figure 1B As shown in S103 on the left side, the screen is off and the user says, "Hello YOYO". The time domain diagram of the audio signal #2 recorded by the microphone of the electronic device is as follows: Figure 1B As shown in the lower left corner, the speech waveform is almost completely submerged by the noise waveform, making it difficult to distinguish. For example, if an electronic device performs VAD on an audio signal using a traditional VAD algorithm, due to the low signal-to-noise ratio, the zero-crossing rate characteristics or energy characteristics of the speech signal and the noise signal are almost indistinguishable, and therefore it is likely that the speech and non-speech regions cannot be distinguished. Assuming that the judgment result is that audio signal #2 is entirely non-speech, the subsequent voice wake-up algorithm will not be performed, and the electronic device will not be awakened, and the interface described in S103 will still be displayed. For another example, if the electronic device performs VAD on audio signal #2 using an artificial intelligence VAD algorithm and inputs audio signal #2 into a neural network for processing, because the noise in the current noisy environment is not included in the neural network's training database, the output of the neural network is likely to be inaccurate. If the judgment result identifies the actual non-speech region as a speech region for subsequent voice wake-up algorithm processing, then the electronic device will not process the speech signal and it is likely that the electronic device will not be awakened, and the interface described in S103 will still be displayed.

[0101] To improve the robustness of the VAD method, the present application provides a voice activity detection method that obtains an audio signal to be processed and determines a voice activity sequence of the audio signal based on a diffusion model. The processing of the audio signal based on the diffusion model includes obtaining a time frame signal in a sequence based on audio features, calculating the gradient of the time frame signal based on a neural network model, the gradient being used to characterize the probability distribution of the voice signal included in the audio signal, predicting a sampled signal based on the gradient, and obtaining a voice activity sequence based on the sampled signal to indicate whether a voice signal exists within each time frame.

[0102] In other words, the voice activity detection method provided by this application uses a neural network to obtain the distribution of voice signals in an audio signal, outputting a gradient representing the probability distribution of the voice signal, and then generating a voice activity sequence based on the predicted sampled signal based on this gradient. Therefore, the voice activity detection method provided by this application does not focus on the noise signal in the audio signal, nor does it distinguish between the noise signal and the voice signal. Instead, it focuses on the probability distribution of the voice signal. As a result, the voice activity detection method provided by this application is not affected by the signal-to-noise ratio or the type of noise. Furthermore, the voice activity detection method provided by this application can accurately detect voice activity even in extremely low signal-to-noise ratio conditions.

[0103] It can be understood that the time frame signal in sequence form and the audio signal to be processed are input into the neural network model, and the probability distribution of the speech signal is output. The sampled signal obtained by gradient prediction of the time frame signal (or sampling the time frame signal) can be understood as the sampled signal can also be used to represent the probability distribution of the speech signal. The probability distribution of the speech signal represented by the sampled signal corresponds to fewer time-frequency points than the probability distribution of the speech signal represented by the gradient of the time frame signal. A smaller number of sample points is used to approximate the overall distribution to achieve noise reduction.

[0104] In some embodiments, the processing of the diffusion model further includes: calculating a gradient of the sampled signal, and correcting the sampled signal based on the gradient.

[0105] It can be understood that calculating the gradient based on the predicted signal and then performing correction can improve the accuracy of the prediction, thereby improving the accuracy of the final speech activity sequence.

[0106] In some embodiments, the training data set of the neural network includes a sample sequence and a sample audio signal, the loss function value is used to characterize the probability distribution of the sample speech signal, the sample audio signal includes a sample speech signal, and the sample sequence is generated based on the sample audio signal. Thus, the loss function value is used to update the parameters that need to be trained in the neural network so that the output of the trained neural network approaches the probability distribution of the sample speech signal in the input sample audio signal.

[0107] That is to say, by setting the loss function, the neural network learns the probability distribution of the sample signal in the sample audio during the training process, so that on the inference side of the neural network, the probability distribution of the speech signal in the audio signal can be obtained based on the input audio signal and the sequence signal.

[0108] Figure 2B A schematic diagram shows an example of the result of performing VAD on audio signal #2 using the voice activity detection method provided by an embodiment of the present application.

[0109] like Figure 2B As shown, Figure 1B The pulse diagram added to the time domain diagram in the figure can be understood as a visualization of the VAD results. The time domain with pulses indicates speech, and the time domain without pulses indicates no speech.

[0110] That is to say, under the condition of extremely low signal-to-noise ratio, the voice activity detection method provided by the present application can accurately obtain the result of voice activity detection.

[0111] It can also be understood that compared with the traditional VAD method, the voice activity detection method provided by the present application is not affected by the signal-to-noise ratio and is more stable.

[0112] It can also be understood that compared to the artificial intelligence VAD method, the voice activity detection method provided in this application does not learn the noise signal, nor does it learn the difference between the noise signal and the voice signal. Therefore, it is not affected by the noise type and has better generalization performance.

[0113] In some embodiments, the audio features may further include time domain features such as zero-crossing rate features and energy features, or may further include frequency domain features.

[0114] It is understandable that although these features have little effect on VAD under low signal-to-noise ratio conditions, when the signal-to-noise ratio is high, VAD based on these features is very effective and can improve the training efficiency of the neural network.

[0115] That is to say, the voice activity detection method provided in the present application is not only applicable to situations with low signal-to-noise ratios, but also to situations with high signal-to-noise ratios. During the model training process, when the signal-to-noise ratio is low, the zero-crossing rate and energy have a weak effect on the training of the neural network, and mainly rely on the sample sequence and sample time-frequency spectrum for training; while when the signal-to-noise ratio is high, the zero-crossing rate and energy have a strong effect on the training of the neural network, and can assist the generated sample sequence and sample time-frequency spectrum in training the neural network, thereby improving the efficiency of neural network training. Correspondingly, during the model inference process, when the signal-to-noise ratio is low, the zero-crossing rate and energy have a weak effect on VAD, and mainly rely on the time frame signal and time-frequency spectrum for inference; while when the signal-to-noise ratio is high, the effect of VAD based on zero-crossing rate and energy is very good, and can effectively assist other information in performing VAD on the audio signal, thereby improving the efficiency and accuracy of VAD.

[0116] Therefore, the voice activity detection method provided in the present application has a very stable VAD effect and accurate VAD results under high or low signal-to-noise ratio conditions; thus, the voice activity detection method provided in the present application can stably obtain accurate voice activity detection results.

[0117] Therefore, thanks to the voice activity detection method provided in this application that stably provides accurate VAD results, subsequent audio algorithms (such as speech recognition, voice wake-up, etc.) can stably obtain more accurate processing results, greatly improving the user experience.

[0118] Figure 3 A schematic diagram of a voice activity detection method 200 provided in an embodiment of the present application is shown.

[0119] First, obtain the audio signal x(k) to be processed. For example, the audio signal to be processed may be Figure 1A The signal recorded by the microphone of an electronic device.

[0120] Here, k corresponds to one sample. For example, if the sampling rate is 16k, there are 16k samples in 1 second.

[0121] Exemplarily, a microphone is used to record an audio signal in the time domain.

[0122] Next, based on the zero-crossing rate calculation module 201, the energy calculation module 202, and the time-frequency spectrum calculation module 203, the zero-crossing rate sequence c(l), the energy sequence h(l), and the time-frequency spectrum z(f,l) of the audio signal x(k) are calculated. Where f is the frequency index, l is the time frame index, and represents the lth time frame. Alternatively, l represents the lth window (windowing involved in the short-time Fourier transform), l is the index of the lth window, and f represents the index of the frequency corresponding to the lth window, l∈[1,L], f∈[1,F], and the lth window corresponds to F frequencies. Where l, f, L, and F are all positive integers.

[0123] Exemplarily, the time-frequency spectrum calculation module 203 calculates the time-frequency spectrum using short-time Fourier transform to obtain the time-frequency spectrum.

[0124] It is understandable that one or both of the zero-crossing rate calculation module 201 and the energy calculation module 202 may optionally perform the above steps.

[0125] Again, the zero-crossing rate sequence c(l), energy sequence h(l) and time-frequency spectrum z(f,l) are input into the diffusion model-based voice activity detection module 204, and voice activity detection is performed using the diffusion model-based voice activity detection algorithm, and finally the voice activity sequence y(l) is output.

[0126] For the diffusion model, please refer to the relevant description above.

[0127] It should be noted that the audio signal to be processed obtained in this application can be obtained offline or online. The online acquisition method can be, for example, in scenarios such as voice recognition and voice wake-up; the offline acquisition method can be, for example, an electronic device for an audio signal downloaded from the network or a pre-stored audio signal. This is explained here in a unified manner and will not be repeated below.

[0128] Combination of the above Figure 3 The overall idea of ​​the voice activity detection method provided by the embodiment of the present application is introduced below. Figures 4 to 6 This method is introduced in detail.

[0129] Figure 4 FIG2 is a schematic diagram of a voice activity detection method 300 provided in an embodiment of the present application. The diffusion model-based voice activity detection module 204 in the method 200 performs the method 300. In other words, the method 300 can be an implementation of the diffusion model-based voice activity detection module 204 obtaining a voice activity sequence.

[0130] S301: Acquire an audio signal to be processed.

[0131] The audio features of the audio signal include a time-frequency spectrum.

[0132] Optionally, the audio feature further includes: a time domain feature, wherein the time domain feature includes a zero-crossing rate feature and / or an energy feature.

[0133] S302: Obtain a time frame signal according to the audio signal.

[0134] The time frame signal is a sequence corresponding to the time frame of the time spectrum.

[0135] Exemplarily, when the audio features include time-frequency spectrum and time-domain features, the time frame signal is obtained based on one or more of the audio features. For example, the time frame signal can be determined based on the time-frequency spectrum and / or can be determined based on the time-domain features.

[0136] Optionally, the audio features may further include frequency domain features. Similarly, the time frame signal may be determined based on the frequency domain features.

[0137] It is understandable that other features in the audio features besides the time-frequency spectrum features can help improve the accuracy and efficiency of VAD. For details, please refer to the relevant description above and will not be repeated here.

[0138] S303 , calculating the gradient of the time frame signal based on the neural network model and the audio signal and the time frame signal.

[0139] The gradient of the time frame signal is used to characterize the probability distribution of the speech signal included in the audio signal. For example, the probability distribution of the speech signal corresponds to the distribution of time-frequency points in the time-frequency graph (or spectrogram) of the speech signal.

[0140] Exemplarily, the neural network model is trained based on a training data set and a loss function value, the training data set includes a sample sequence and a sample audio signal, the sample audio signal includes a sample speech signal, the sample sequence is generated based on the sample audio signal, and the loss function value is used to characterize the probability distribution of the sample speech signal.

[0141] In one possible implementation, the training data set also includes sample audio features of the sample audio signal, and the sample audio features include a sample time-frequency spectrum. As an example of S303, S303a inputs the time frame signal, the audio signal, and the audio features into the neural network model to obtain the gradient of the time frame signal, wherein the sample sequence is a first Gaussian noise value, which is generated based on the sample speech signal, the sample audio signal, and a second Gaussian noise value that follows a standard normal distribution, and the loss function value is generated based on the variance of the first Gaussian noise value and the second Gaussian noise value, which is generated based on the first random seed.

[0142] It is understandable that the random seed involved in this application is a preset value determined through a large number of experiments, which can achieve better voice activity detection results.

[0143] It is understandable that, because the random seed generates different random numbers at different times, the method provided in this application processes the audio signal at different times, and the voice activity detection sequences outputted are different or not completely the same.

[0144] As can be understood, inputting sample audio features, including the sample time-frequency spectrum, into the neural network adds training data, namely the time-frequency spectrum, to the sample sequence. This allows the neural network to more accurately determine the probability distribution of the sample speech signal, improving the accuracy of the ultimately trained neural network. Consequently, after inputting the time-frequency signal and audio features into the neural network, the neural network can more accurately output the probability distribution of the speech signal.

[0145] Optionally, in S303a, the sample audio features further include: sample time domain features, wherein the sample time domain features include sample zero-crossing rate features and / or sample energy features. The sample audio features on the training side correspond to the audio features on the inference side.

[0146] It is understandable that by adding these features to the training side to train the neural network, it is also possible to accelerate convergence and improve motion efficiency when the signal-to-noise ratio of the sample audio signal is high, so that the VAD efficiency can be improved when the signal-to-noise ratio on the neural network inference side is high.

[0147] Exemplarily, in S303a, the mean value of the first Gaussian noise value is generated according to the sample voice activity sequence and the sample audio signal, and the sample voice activity sequence is a result of voice activity detection generated based on the sample voice.

[0148] It can be understood that using a sample speech activity sequence to generate a sample sequence can make the sample sequence generated more accurate, thereby reducing the number of training steps of the neural network and accelerating the convergence of the neural network.

[0149] Exemplarily, in S303a, the variance of the first Gaussian noise value is generated according to the diffusion intensity, and the diffusion intensity is determined based on a stochastic differential equation.

[0150] It can be understood that the diffusion strength can adjust the randomness of the first Gaussian noise value, and by setting the diffusion strength value, the neural network can be adjusted to achieve better training effects.

[0151] S304: predicting a sampling signal of the time frame signal according to the gradient of the time frame signal.

[0152] As an example of S304, S304c, based on the stochastic differential equation of the conditional diffusion model, a drift coefficient is obtained according to the time frame signal, wherein the condition of the conditional diffusion model is the audio signal; based on the drift coefficient and the gradient of the time frame signal, an inverse drift coefficient is obtained; based on the time frame signal, the inverse drift coefficient and the third Gaussian noise value, the sampling signal is predicted, wherein the third Gaussian noise value is generated based on the second random seed.

[0153] S305 , obtaining a voice activity sequence of the audio signal according to the sampling signal. The voice activity sequence corresponds to a time frame of the time spectrum. The voice activity sequence is used to indicate whether a voice signal exists in each time frame of the time spectrum.

[0154] As an example of S305, S305a, the gradient of the sampling signal is calculated based on the neural network model, and the gradient of the sampling signal is used to characterize the probability distribution of the speech signal; according to the gradient of the sampling signal, the sampling signal is corrected to obtain a corrected signal; according to the corrected signal, a speech activity sequence is obtained.

[0155] It can be understood that calculating the gradient based on the predicted signal and then performing correction can improve the accuracy of the prediction, thereby improving the accuracy of the final speech activity sequence.

[0156] The above scheme achieves stable and accurate VAD results regardless of whether the signal-to-noise ratio is high or low. Therefore, the voice activity detection method provided by this application can consistently obtain accurate voice activity detection results. Consequently, subsequent audio algorithms (such as speech recognition and voice wake-up) can also consistently obtain more accurate processing results, greatly improving the user experience.

[0157] It is understandable that S302 to S304 may be executed N times in a loop to further improve the accuracy of the prediction.

[0158] Execution for the first time:

[0159] As an example of S302, S302a, a first time frame signal is obtained according to audio features.

[0160] As an example of S303, S303b, the gradient of the first time frame signal is calculated based on the neural network model, and the gradient of the first time frame signal is used to represent the probability distribution of the speech signal.

[0161] As an example of S304, S304a, predict the first sampling signal according to the gradient of the first time frame signal.

[0162] Execute the nth time (2≤n≤N, and N and n are both integers):

[0163] As an example of S302, S302b, obtain the nth time frame signal according to the n-1th sampling signal, wherein the n-1th sampling signal is obtained according to the audio feature, 2≤n≤N, and N and n are both integers.

[0164] As an example of S303, S303c, the gradient of the n-th time frame signal is calculated based on the neural network model, and the gradient of the n-th time frame signal is used to represent the probability distribution of the speech signal.

[0165] As an example of S304, S304b, predict the nth sampling signal according to the gradient of the nth time frame signal.

[0166] Regarding the case where S302 to S304 can be executed N times in a loop, as an example of S305 , S305 b , when n=N, obtain a voice activity sequence according to the Nth sampling signal.

[0167] That is, a possible implementation of the diffusion model includes a process of gradient calculation and prediction (or sampling).

[0168] Exemplarily, in combination with S305a, as an example of S302b, S302c calculates the gradient of the n-1th sampling signal, and the gradient of the n-1th sampling signal is used to characterize the probability distribution of the speech signal; according to the gradient of the n-1th sampling signal, the n-1th sampling signal is corrected to obtain the n-1th corrected signal, and the n-1th corrected signal is used as the nth time frame signal.

[0169] Exemplarily, for the case where S302 to S304 can be executed in a loop N times, as an example of S305a, S305c calculates the gradient of the Nth sampling signal, and the gradient of the Nth sampling signal is used to characterize the probability distribution of the speech signal; according to the gradient of the Nth sampling signal, the Nth sampling signal is corrected to obtain the Nth corrected signal, and the Nth corrected signal is used as the speech activity sequence.

[0170] In other words, another possible implementation of the diffusion model includes a process of gradient calculation, prediction (or sampling), recalculation of gradients, and correction. This allows for higher VAD accuracy, and subsequent audio algorithms (such as speech recognition and voice wake-up) can obtain more accurate processing results, improving the user experience.

[0171] For example, in the case where S302 to S304 can be executed N times in a loop, as an example of S304c, S304d:

[0172] The first execution case: based on the stochastic differential equation of the conditional diffusion model, the first drift coefficient is obtained according to the first time frame signal, wherein the condition of the conditional diffusion model is the audio signal; according to the first drift coefficient and the gradient of the first time frame signal, the first inverse drift coefficient is obtained; according to the first time frame signal, the first inverse drift coefficient and the first third Gaussian noise value, the first sampling signal is predicted, wherein the first third Gaussian noise value is generated based on the first second random seed.

[0173] Execution for the nth time: Based on the stochastic differential equation of the conditional diffusion model, the nth drift coefficient is obtained according to the nth time frame signal and the 1st time frame signal, wherein the condition of the conditional diffusion model is the audio signal; based on the nth drift coefficient and the gradient of the nth time frame signal, the nth inverse drift coefficient is obtained; based on the nth time frame signal, the nth inverse drift coefficient and the nth third Gaussian noise value, the nth sampling signal is predicted, wherein the nth third Gaussian noise value is generated based on the nth second random seed.

[0174] Since the inverse drift coefficient is used to characterize the difference between the denoised random signal and the speech signal in the process of denoising the audio signal to be processed by sampling, that is, the use of the inverse drift coefficient for prediction in this application can reduce the noise of the audio signal while obtaining the probability distribution of the speech signal.

[0175] Figure 5 FIG2 shows a schematic diagram of a voice activity detection method 400 provided by an embodiment of the present application. In the method 200, the voice activity detection module 204 based on the diffusion model performs the method 400. That is, the method 400 can be an implementation of the voice activity detection module 204 based on the diffusion model to obtain a voice activity sequence. For example, the voice activity detection module 204 based on the diffusion model includes the following: Figure 5 The initialization module 401 to the judgment module 407 are shown.

[0176] Initialization module 401, initialize sampling times n = 1 (indicates the current first sampling), initialize sampling time s = 0.999, initialize sampling signal The initialized parameters are output to the drift coefficient calculation module 402 .

[0177] Here, x(l) is the initialization audio signal, which is a sequence corresponding to the time frame of the time-frequency spectrum of x(k). x(l) can be the zero-crossing rate sequence c(l) or the energy sequence h(l), or the sum of the time-frequency spectrum z(f,l) according to the frequency index. For example, x(l) = sum(z(f,l),f).

[0178] It can be understood that the value of l is 1, 2, 3...L.

[0179] Alternatively, x(l) may also be other randomly generated numbers, which is not limited in this application.

[0180] in, Indicates the time frame signal.

[0181] It can be understood that the steps performed by the initialization module 401 can be used as a specific example of SS302a.

[0182] Drift coefficient calculation module 402, the input of this module is x(l), s and The function of this module is to calculate the drift coefficient fc according to the stochastic differential equation, such as The drift coefficient calculated by this module is output to the inverse drift coefficient calculation module 404 to calculate the inverse drift coefficient.

[0183] It can be understood that the step performed by the drift coefficient calculation module 402 can be used as a specific example of S304d.

[0184] Gradient calculation module 403, the input of this module is x(l), s, c(l), h(l), z(f,l). The function of this module is to calculate the gradient #1 using the trained neural network model. The gradient calculated by this module is output to the inverse drift coefficient calculation module 404 to calculate the inverse drift coefficient.

[0185] It can be understood that the steps performed by the above-mentioned gradient calculation module 403 can be used as a specific example of S303b and S303c.

[0186] The inverse drift coefficient calculation module 404 takes the drift coefficient fc and the gradient grad as input, calculates the inverse drift coefficient fcr using the formula fcr=-fc+grad, and outputs it to the prediction module 405 for prediction.

[0187] It can be understood that the step performed by the inverse drift coefficient calculation module 404 can be used as a specific example of S304d.

[0188] The prediction module 405, which may also be referred to as the sampling prediction module 405, uses the formula To predict the new sampled signal, z is Gaussian noise with mean zero and variance 1 (i.e., the third Gaussian noise described above). It should be noted that generating Gaussian noise requires setting a random seed. For evidence collection, the random seed is set when the terminal device is reset.

[0189] Among them, the formula Said that Updated to Get a new sampling signal.

[0190] It can be understood that the steps performed by the prediction module 405 can be used as a specific example of S304.

[0191] After the prediction module 405 , the gradient calculation module 403 is also performed.

[0192] Gradient calculation module 403, the input of this module is x(l), s, c(l), h(l), z(f,l). Here, is the new sampled signal updated in the prediction module 405 The function of this module is to calculate the gradient #2 using the trained neural network model. The gradient calculated by this module is output to the correction module 406 for correction.

[0193] It can be understood that calculating the gradient once based on the predicted new sampling signal is equivalent to performing one more calculation by the neural network. The calculated gradient effect is better, and it can better represent the distribution of the speech signal, thereby making the speech activity sequence finally output by this method more accurate.

[0194] It can be understood that the steps performed by the above-mentioned gradient calculation module 403 can be used as a specific example of S305a.

[0195] Correction module 406 uses the annealed Langevin kinetic sampling method to calibrate the predicted new sampling signal according to gradient #2. Perform correction to obtain the correction signal after correction In the case of , it serves as the n+1th input of the drift coefficient calculation module 402, and in the case of n=N, it serves as the final output of this method.

[0196] It can be understood that the steps performed by the correction module 406 can be used as a specific example of S305a.

[0197] The judgment module 407 judges whether n is equal to N (the preset total number of sampling steps, for example, assuming it is 10). If n=N, the correction signal Output as y(l), otherwise update the sampling times and sampling time, n←n+1, s=max(0.03, (n-1)×(0.999-0.03) / N)), and execute the drift coefficient calculation module 402 cyclically.

[0198] The beneficial effects of method 400 can be found in the beneficial effects of the corresponding solution in method 300.

[0199] Figure 6 A schematic diagram of a voice activity detection method 500 provided in an embodiment of the present application is shown.

[0200] It is understood that the neural network model used in the gradient calculation module 403 (ie, the neural network model used in the method 300) needs to be trained in advance (online or offline training is possible), and the training process is as follows: Figure 6 As shown, it includes the following sub-modules.

[0201] The activity detection module 501 obtains a clean speech signal and generates a voice activity detection sequence x0 of the clean speech signal using a voice activity detection algorithm.

[0202] It is understandable that, because the input signal at this time is a noise-free clean speech signal, general voice activity detection algorithms (such as the traditional VAD method and artificial intelligence-based VAD algorithm mentioned above) can obtain a better voice activity detection sequence x0 of the clean speech signal.

[0203] The noisy audio signal processing module 502 mixes the noise signal and the clean speech signal to obtain a noisy audio signal y as a sample audio signal.

[0204] Sampling time generation module 503 requires no input. This module randomly generates a decimal between tmin and tmax as the sample sampling time, where tmin is the minimum value, such as 0.003, and tmax is the maximum value, such as 0.999. The sample sampling time generated by this module is output to sample generation module 505.

[0205] The Gaussian noise generation module 504 first calculates the scale of the Gaussian noise (i.e., the first Gaussian noise mentioned above) according to the variance formula (i.e., the standard deviation σ of the Gaussian noise). t ), Among them, g is a preset parameter, and its physical meaning is the diffusion intensity of the random differential variance. The larger the variance, the larger the Gaussian noise and the greater the randomness. Too small or too large a diffusion intensity may affect the effect. If the diffusion intensity is too small, the randomness is too small, and the final generated gradient effect will be relatively poor; vice versa. Preferably, g = 2. Or, g = 1. Then generate a set of zero-mean unit variance Gaussian noise z (that is, the second Gaussian noise mentioned above), and label To the loss function calculation module.

[0206] Sample generation module 505, according to The generated mean is μ=(1-t)x0+ty and the variance is Gaussian noise. x(t) is called the sample sequence. It should be noted that, It is derived, not simply obtained based on the sample sequence.

[0207] The deep neural network model 506 inputs the noisy audio signal y, sampling time t, x(t), c(l), h(l), z(f,l) into the module, and the output of the module is the sample gradient.

[0208] There are many options for deep neural network models, such as noise conditional scoring network (NCSN)++ network, convolutional neural networks (CNN), long short-term memory network (LSTM), U-net and other network models.

[0209] The loss function calculation module 507 calculates the loss function value according to the loss function, such as And use the calculated loss function value to update the parameters of the deep neural network model.

[0210] The beneficial effects of method 500 can be found in the beneficial effects of the corresponding solution in method 300.

[0211] Figure 7 A hardware structure diagram of an electronic device 1000 provided in an embodiment of the present application is shown.

[0212] See also Figure 7 The electronic device 1000 may include a processor 110, an audio module 120, a microphone 120A, and optionally, a speaker 120B.

[0213] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0214] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc.

[0215] The controller may be the nerve center and command center of the electronic device 1000. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0216] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0217] In the present application, the processor 110 is used to obtain an audio signal to be processed and determine a voice activity sequence of the audio signal based on a diffusion model. The processing of the audio signal based on the diffusion model includes obtaining a time frame signal in a sequence form according to audio features, calculating the gradient of the time frame signal based on a neural network model, the gradient being used to characterize the probability distribution of the voice signal included in the audio signal, predicting a sampling signal according to the gradient, and obtaining a voice activity sequence indicating whether a voice signal exists in each time frame based on the sampling signal. The voice activity detection method provided in the present application is very stable in terms of VAD effect and accurate in terms of high or low signal-to-noise ratio. Therefore, the voice activity detection method provided in the present application can stably obtain accurate voice activity detection results.

[0218] In some embodiments, the processor 110 may include one or more interfaces, such as an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0219] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0220] The electronic device 1000 can implement audio functions such as music playback, recording, voice calls, etc. through the audio module 120, the speaker 120B, the microphone 120A, and the application processor (not shown in the figure).

[0221] In a voice call, voice wake-up, or voice recognition scenario, the microphone 120A is used to record the audio signal to be processed.

[0222] Optionally, during a voice call, the speaker 120B is used to play the voice of the party talking to the user; in a voice recognition or voice wake-up scenario, after the electronic device recognizes that the microphone 120A records the user's voice signal, it transmits a feedback voice signal to the user through the speaker 120B.

[0223] Next, the software system of the electronic device 1000 will be described.

[0224] For example, the electronic device 1000 may be a mobile phone. The software system of the electronic device 1000 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software system of the electronic device 1000.

[0225] Figure 8 FIG1 shows a block diagram of a software system of an electronic device 1000 provided in an embodiment of the present application. Figure 8 The layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers: from top to bottom, the application layer (application layer), the application framework layer (framework layer), the hardware abstraction layer (HAL), the driver layer, and the hardware layer.

[0226] The application layer may include a series of application packages, such as a dialing application, a gallery application, etc. (not shown in the figure). In the embodiment of the present application, the application package may include applications such as recording, voice calling, and voice assistants. These applications all require the use of a microphone to record audio and perform voice activity detection on the audio recorded by the microphone so that subsequent audio algorithms can operate normally. Among them, the voice assistant is used to recognize the user's voice and interact with the user.

[0227] Alternatively, the application layer may also include other applications that require voice activity detection processing on the voice signal recorded by the microphone, which is not limited in this application.

[0228] The framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. In the embodiment of the present application, the framework layer includes a microphone service interface and a voice activity detection service interface. The voice activity detection service interface can provide an API and programming framework for applications that access the voice activity detection service. The microphone service can be used to provide an API and programming framework for applications that call the microphone.

[0229] The Hardware Abstraction Layer (HAL) is an interface layer between the operating system kernel and upper-layer software, providing a virtual hardware platform for the operating system. In embodiments of the present application, the hardware abstraction layer may include a microphone hardware abstraction layer and a voice activity detection algorithm. The microphone hardware abstraction layer may provide virtual hardware for microphone 1, microphone 2, or more microphone devices. The voice activity detection algorithm may include operating code and data that implement the voice signal processing method provided in embodiments of the present application.

[0230] The driver layer is the layer between hardware and software. It includes drivers for various hardware components, including microphone device drivers and digital signal processor drivers. The microphone device driver drives the microphone sensor to collect sound signals and the audio signal processor to pre-process the sound signals to generate digital audio signals. The digital signal processor driver drives the digital signal processor to process the digital audio signals.

[0231] The hardware layer includes sensors and audio signal processors. Among them, the sensors include microphones 1 and 2. The microphones included in the sensors correspond one-to-one to the virtual microphones included in the microphone hardware abstraction layer. The audio signal processor can be used to convert the sound signals collected by the microphones into audio digital signals. The digital signal processor can be used to process audio digital signals. It should be noted that the application provides Figure 8 The software structure diagram of the electronic device shown is only an example and does not limit the specific module division in different layers of the Android operating system. For details, please refer to the introduction of the Android operating system software structure in conventional technology.

[0232] The following describes the method in the embodiment of the present application in detail in combination with the above hardware structure and system structure:

[0233] In response to enabling applications such as voice recognition, calling, and voice assistants, such applications can call the voice activity detection service interface to obtain the application programming interface and programming framework provided by the voice activity detection service.

[0234] On the one hand, the Voice Activity Detection service can call the Microphone service in the framework layer to collect sound signals from the environment. Specifically, the Microphone service can call Microphone 1 in the Microphone Hardware Abstraction Layer to send a command to the Microphone 1 sensor in the hardware layer to collect sound signals. The Microphone Hardware Abstraction Layer sends this command to the Microphone Device Driver in the Driver Layer. Based on this command, the Microphone Device Driver activates Microphone 1, thereby acquiring sound signals from the environment and generating digital audio signals through the audio signal processor.

[0235] The Voice Activity Detection service can initialize the Voice Activity Detection algorithm. The Voice Activity Detection algorithm can obtain the digital audio signal generated by the audio signal processor through the Microphone Hardware Abstraction Layer. Then, based on the voice signal processing method stored in the Voice Activity Detection algorithm, the Voice Activity Detection algorithm can use the digital signal processor to process the obtained digital audio signal to obtain a digital audio signal after voice activity detection.

[0236] Specifically, how to process the digital audio signal to obtain the digital audio signal after voice activity detection can be referred to the previous section. Figures 3 to 6 The method flow chart shown.

[0237] Finally, the voice activity detection algorithm can pass the digital audio signal after voice activity detection back to the voice activity detection service and then back to the application layer.

[0238] The present invention provides a chip system comprising one or more processors configured to retrieve and execute instructions stored in a memory, thereby executing the method of the present invention. The chip system may be composed of a chip or may include a chip and other discrete devices.

[0239] The chip system may include an input circuit or interface for sending information or data, and an output circuit or interface for receiving information or data.

[0240] The present application also provides a computer program product, which, when executed by a processor, implements the method described in any method embodiment of the present application.

[0241] The computer program product can be stored in a memory and finally converted into an executable target file that can be executed by a processor through preprocessing, compilation, assembly and linking.

[0242] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements the method described in any method embodiment of the present application. The computer program can be a high-level language program or an executable target program.

[0243] The computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0244] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and equipment and the technical effects produced can refer to the corresponding processes and technical effects in the aforementioned method embodiments, and will not be repeated here.

[0245] In the several embodiments provided in this application, the disclosed systems, devices and methods can be implemented in other ways. For example, some features of the method embodiments described above can be ignored or not executed. The device embodiments described above are merely schematic, and the division of units is only a logical function division. There may be other division methods in actual implementation, and multiple units or components may be combined or integrated into another system. In addition, the coupling between the units or the coupling between the components may be direct coupling or indirect coupling, and the above coupling includes electrical, mechanical or other forms of connection.

[0246] It should be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0247] It should be understood that the term "plurality" used herein refers to two or more. The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0248] The terms (or numbers) "first", "second", ... etc. that appear in the embodiments of the present application are only used for descriptive purposes, that is, they are only used to distinguish different objects, such as different "coordinates", etc., and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first", "second", ... etc. may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "at least one (item)" refers to one or more. "Multiple" means two or more. "At least one of the following (item)" or similar expressions refers to any combination of these items, including any combination of a single (item) or plural (items).

[0249] In short, the above description is only a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included in the scope of protection of this application.

Claims

1. A voice activity detection method, applied to electronic equipment, characterized in that: The method comprises: Acquire an audio signal to be processed, wherein the audio features of the audio signal include a time-frequency spectrum; Obtaining a time frame signal according to the audio signal, wherein the time frame signal is a sequence corresponding to the time frames of the time spectrum; Calculating a gradient of the time frame signal based on the audio signal and the time frame signal based on a neural network model, wherein the gradient of the time frame signal is used to characterize a probability distribution of a speech signal included in the audio signal; predicting a sampling signal of the time frame signal according to the gradient of the time frame signal; A voice activity sequence of the audio signal is obtained according to the sampling signal, wherein the voice activity sequence corresponds to a time frame of the time spectrum and is used to indicate whether a voice signal exists in each time frame of the time spectrum.

2. The method according to claim 1, wherein The step of obtaining a speech activity sequence of the audio signal to be processed according to the sampling signal includes: Calculating a gradient of the sampled signal based on the neural network model, wherein the gradient of the sampled signal is used to characterize a probability distribution of the speech signal; Correcting the sampling signal according to the gradient of the sampling signal to obtain a correction signal; The voice activity sequence is obtained according to the correction signal.

3. The method according to claim 1 or 2, wherein: The neural network model is trained based on a training data set and a loss function value, wherein the training data set includes a sample sequence and a sample audio signal, the sample audio signal includes a sample speech signal, the sample sequence is generated based on the sample audio signal, and the loss function value is used to characterize the probability distribution of the sample speech signal.

4. The method according to claim 3, wherein The training data set also includes sample audio features of the sample audio signal, the sample audio features include sample time-frequency spectra, and the gradient of the time frame signal is calculated based on the neural network model, including: The time frame signal, the audio signal and the audio feature are input into the neural network model to obtain the gradient of the time frame signal, wherein the sample sequence is a first Gaussian noise value, the first Gaussian noise value is generated based on the sample speech signal, the sample audio signal and a second Gaussian noise value that obeys a standard normal distribution, the loss function value is generated based on the variance of the first Gaussian noise value and the second Gaussian noise value, and the second Gaussian noise value is generated based on a first random seed.

5. The method according to claim 4, wherein The mean value of the first Gaussian noise value is generated according to a sample voice activity sequence and the sample audio signal, wherein the sample voice activity sequence is a result of voice activity detection generated based on the sample voice.

6. The method according to claim 5, wherein The variance of the first Gaussian noise value is generated according to a diffusion strength determined based on a stochastic differential equation.

7. The method according to any one of claims 4 to 6, characterized in that The sample audio features further include: sample time domain features, and the audio features further include: time domain features; Wherein, the sample time domain feature includes a sample zero-crossing rate feature, and the time domain feature includes a zero-crossing rate feature; And / or, the sample time domain feature includes a sample energy feature, and the time domain feature includes an energy feature.

8. The method according to any one of claims 1 to 7, characterized in that Obtaining a time frame signal according to the audio signal includes: obtaining a first time frame signal according to part or all of the audio features of the audio signal; Calculating the gradient of the time frame signal based on the neural network model includes: calculating the gradient of the first time frame signal based on the neural network model, the gradient of the first time frame signal being used to characterize the probability distribution of the speech signal; The predicting the sampling signal of the time frame signal according to the gradient of the time frame signal includes: predicting the first sampling signal according to the gradient of the first time frame signal.

9. The method according to claim 8, wherein Obtaining a time frame signal according to the audio signal includes: obtaining an nth time frame signal according to an n-1th sampling signal, wherein the n-1th sampling signal is obtained according to the audio signal, 2≤n≤N, and N and n are both integers; The calculating the gradient of the time frame signal based on the neural network model includes: calculating the gradient of the nth time frame signal based on the neural network model, the gradient of the nth time frame signal being used to characterize the probability distribution of the speech signal; The predicting the sampling signal of the time frame signal according to the gradient of the time frame signal includes: predicting the nth sampling signal according to the gradient of the nth time frame signal.

10. The method according to claim 9, wherein The step of obtaining a speech activity sequence of the audio signal to be processed according to the sampling signal includes: In the case of n=N, the voice activity sequence is obtained according to the Nth sampling signal.

11. The method according to claim 9 or 10, wherein: The step of obtaining the nth time frame signal according to the n-1th sampling signal includes: Calculating the gradient of the n-1th sampling signal, where the gradient of the n-1th sampling signal is used to characterize the probability distribution of the speech signal; The n-1th sampling signal is corrected according to the gradient of the n-1th sampling signal to obtain an n-1th correction signal, and the n-1th correction signal is used as the nth time frame signal.

12. The method according to claim 11, wherein The step of obtaining the voice activity sequence according to the Nth sampling signal includes: Calculating a gradient of the Nth sampling signal, where the gradient of the Nth sampling signal is used to characterize a probability distribution of the speech signal; The Nth sampling signal is corrected according to the gradient of the Nth sampling signal to obtain an Nth corrected signal, and the Nth corrected signal is used as the voice activity sequence.

13. The method according to any one of claims 1 to 12, characterized in that The predicting the sampling signal of the time frame signal according to the gradient of the time frame signal includes: Obtaining a drift coefficient according to the time frame signal based on a stochastic differential equation of a conditional diffusion model, wherein the condition of the conditional diffusion model is the audio signal; Obtaining an inverse drift coefficient according to the drift coefficient and the gradient of the time frame signal; The sampling signal is predicted according to the time frame signal, the inverse drift coefficient, and a third Gaussian noise value, wherein the third Gaussian noise value is generated based on a second random seed.

14. The method according to claim 8, wherein The step of predicting the first sampling signal according to the gradient of the first time frame signal comprises: Obtaining a first drift coefficient according to the first time frame signal based on a stochastic differential equation of a conditional diffusion model, wherein a condition of the conditional diffusion model is the audio signal; Obtaining a first inverse drift coefficient according to the first drift coefficient and the gradient of the first time frame signal; The first sampling signal is predicted according to the first time frame signal, the first inverse drift coefficient, and the first third Gaussian noise value, wherein the first third Gaussian noise value is generated based on the first second random seed.

15. The method according to claim 9, wherein The predicting of the nth sampling signal according to the gradient of the nth time frame signal comprises: Obtaining an nth drift coefficient based on the nth time frame signal and the first time frame signal using a stochastic differential equation of a conditional diffusion model, wherein the condition of the conditional diffusion model is the audio signal; Obtaining an nth inverse drift coefficient according to the nth drift coefficient and the gradient of the nth time frame signal; The nth sampling signal is predicted according to the nth time frame signal, the nth inverse drift coefficient, and the nth third Gaussian noise value, wherein the nth third Gaussian noise value is generated based on the nth second random seed.

16. An electronic device, characterized in that: The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 15.

17. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 15.