Audio noise reduction method and device, electronic equipment and readable storage medium
By calculating the long-term signal-to-noise ratio and stability index of audio signals, four types of acoustic scenarios are identified for targeted noise reduction. This solves the problem of insufficient universality and practicality of existing acoustic scenario classification methods, and achieves a more efficient audio noise reduction effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2022-08-25
- Publication Date
- 2026-08-04
AI Technical Summary
Existing acoustic scene classification methods are insufficient in terms of versatility and practicality. Traditional methods only select specific scenes, while deep learning-based methods are too complex and difficult to deploy in conjunction with actual noise reduction needs.
By calculating the long-term signal-to-noise ratio (SNR) and long-term stationarity of the target audio signal, the target acoustic scene is determined, and targeted noise reduction processing is performed based on the scene. The long-term SNR and stationarity are used as the essential characteristics of the audio signal for classification, which is divided into four types of acoustic scenes: high SNR non-stationary, low SNR non-stationary, high SNR stationary, and low SNR stationary.
It improves the accuracy and practicality of audio noise reduction, meets the different noise reduction needs of users in different acoustic scenarios, and enhances the sound quality of electronic devices.
Smart Images

Figure CN115995234B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of audio processing technology, specifically relating to an audio noise reduction method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] Acoustic scene classification has wide applications in daily life. It refers to the process of analyzing the acoustic content contained in audio to identify the corresponding acoustic scene.
[0003] Acoustic scene classification in related technologies is mainly achieved through the following two methods: Method 1, based on traditional acoustic scene classification methods, specifically observes the signal characteristics of a specific scene and extracts the corresponding features for acoustic scene classification. Method 2, a scene classification method based on deep learning models, specifically extracts speech features from the input speech signal, such as Mel-frequency cepstral coefficients, logarithmic amplitude spectrum, and phase spectrum, and selects an appropriate deep classification model for supervised learning based on the extracted speech features. Then, the learned deep classification model is used to classify the audio into acoustic scenes.
[0004] However, following the methods described above, traditional acoustic scene classification methods only select specific acoustic scenes, while deep learning-based scene classification methods are too complex and difficult to deploy in conjunction with actual noise reduction needs. Thus, the versatility and practicality of acoustic scene classification methods in related technologies are poor. Summary of the Invention
[0005] The purpose of this application is to provide an audio noise reduction method, apparatus, electronic device, and readable storage medium that can solve the problem of poor universality and practicality of audio noise reduction methods in related technologies.
[0006] In a first aspect, embodiments of this application provide an audio noise reduction method, the method comprising: calculating a target long-time signal-to-noise ratio and a target long-time stationarity index corresponding to a target audio signal, wherein the target long-time stationarity index is used to indicate the degree of stability of noise in the target audio signal; determining a target acoustic scene corresponding to the target audio signal based on the target long-time signal-to-noise ratio and the target long-time stationarity index; and performing noise reduction processing on the target audio signal based on the target acoustic scene.
[0007] Secondly, embodiments of this application provide an audio noise reduction device, which includes a processing module and a determining module. The processing module is used to calculate the target long-time signal-to-noise ratio and the target long-time stability index corresponding to the target audio signal, wherein the target long-time stability index is used to indicate the stability of noise in the target audio signal; the determining module is used to determine the target acoustic scene corresponding to the target audio signal based on the target long-time signal-to-noise ratio and the target long-time stability index calculated by the processing module; the processing module is further used to perform noise reduction processing on the target audio signal based on the target acoustic scene determined by the determining module.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program product, the program product being stored in a storage medium, and the program product being executed by at least one processor to implement the method as described in the first aspect.
[0012] In this embodiment, the target long-term signal-to-noise ratio (SNR) and target long-term stability index corresponding to the target audio signal are calculated. The target long-term stability index indicates the stability of noise in the target audio signal. Based on the target SNR and target long-term stability index, the target acoustic scene corresponding to the target audio signal is determined. Based on the target acoustic scene, the target audio signal is denoised. This scheme, since the long-term SNR and stability index corresponding to the audio signal are two essential characteristics of noise in the audio signal, allows for a more accurate and rapid determination of the target acoustic scene based on the target audio signal. This improves the accuracy of denoising the target audio signal based on the target acoustic scene, and the denoising method has better versatility and practicality. Attached Figure Description
[0013] Figure 1 This is one of the flowcharts illustrating the audio noise reduction method provided in the embodiments of this application;
[0014] Figure 2 This is a second schematic flowchart of the audio noise reduction method provided in the embodiments of this application;
[0015] Figure 3 This is a schematic diagram of the audio noise reduction device provided in the embodiments of this application;
[0016] Figure 4 This is one of the structural schematic diagrams of the electronic device provided in the embodiments of this application;
[0017] Figure 5 This is the second schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0020] For electronic device users, call quality is a crucial indicator of an device's performance. To improve sound quality, noise reduction processing can be applied to the voice signal during calls.
[0021] Currently, speech signals can be classified according to their acoustic environment, allowing for targeted noise reduction based on the specific acoustic environment in which they belong.
[0022] Acoustic scene classification in related technologies mainly employs traditional audio noise reduction methods and scene classification methods based on deep learning models.
[0023] Specifically, traditional audio noise reduction methods can extract corresponding features by observing the signal characteristics of specific scenes to classify acoustic scenes, and then perform corresponding noise suppression after classification. Examples include detecting wind noise, keyboard noise, and mobile phone motor vibration. It is evident that traditional acoustic scene classification methods can only classify specific acoustic scenes.
[0024] Scene classification methods based on deep learning models generally consist of two steps. The first step is to extract speech features from the input speech signal, such as Mel-frequency cepstral coefficients, logarithmic amplitude spectrum, and phase spectrum. The second step is to select an appropriate deep classification model based on these extracted speech features for supervised learning, and then use the learned deep classification model to classify the audio into acoustic scenes.
[0025] However, deep learning-based scene classification methods have the following drawbacks: (1) The network size is generally large, making it difficult to deploy in real time for some scenarios with high power consumption requirements; (2) Labeling noisy scenes one by one is also a large workload; (3) The scene classification is too detailed, which is not conducive to practical use. For example, the scenes are divided into too many different scenes such as subway, bus, coffee shop, canteen, car, airport, etc.
[0026] As discussed above, traditional acoustic scene classification methods only select specific scenarios, while deep learning-based methods are too complex and difficult to deploy in practice for noise reduction. Consequently, the versatility and practicality of acoustic scene classification methods in related technologies are poor.
[0027] The audio noise reduction method provided in this application aims to offer an audio noise reduction method based on the mainstream noise reduction algorithm framework. Based on the long-term signal-to-noise ratio and stability index of the audio signal, the acoustic scene to which the audio signal belongs is determined, which has better versatility and practicality.
[0028] Unlike common acoustic scene classifications, such as the overly diverse multi-classification of scenes like subways, buses, cafes, canteens, cars, and airports, or the simple binary classification of special scenes like wind noise, this application proposes a classification based on two essential characteristics of acoustic scenes: long-term signal-to-noise ratio (SNR) and noise stability (i.e., noise stability). This classification combines both SNR and noise stability to create four acoustic scenes: a first acoustic scene, a second acoustic scene, a third acoustic scene, and a fourth acoustic scene.
[0029] In the first acoustic scenario, the long-term signal-to-noise ratio (SNR) of the audio signal is greater than or equal to the SNR threshold, and the long-term stability index of the audio signal is greater than or equal to the stability index threshold. In the second acoustic scenario, the long-term SNR of the audio signal is greater than or equal to the SNR threshold, and the long-term stability index of the audio signal is less than the stability index threshold. In the third acoustic scenario, the long-term SNR of the audio signal is less than the SNR threshold, and the long-term stability index of the audio signal is greater than or equal to the stability index threshold. In the fourth acoustic scenario, the long-term SNR of the audio signal is less than the SNR threshold, and the long-term stability index of the audio signal is less than the stability index threshold. Thus, since acoustic scenarios can be divided into four categories based on two essential characteristics, the audio noise reduction method provided in this application has significant practicality and versatility.
[0030] Optionally, after determining the acoustic scene of the audio signal, noise suppression is applied to the audio signal based on the noise suppression strategy corresponding to that acoustic scene. This allows for addressing different noise reduction needs of users in different acoustic scenes. For example, in noisy environments, users desire greater noise suppression. In high signal-to-noise ratio scenarios, users prefer to preserve the original voice quality and do not want excessive noise reduction processing. In non-stationary scenarios, users desire effective suppression of sudden noise. In stationary scenarios, users desire more natural noise suppression.
[0031] The audio noise reduction method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0032] This application provides an audio noise reduction method. Figure 1 This paper illustrates a possible flowchart of an audio noise reduction method provided in an embodiment of this application, as shown below. Figure 1 As shown, the audio noise reduction method provided in this application embodiment may include steps 101 to 103 as described below. The following description uses an electronic device performing the method as an example.
[0033] Step 101: The electronic device calculates the target long-term signal-to-noise ratio and the target long-term stability index corresponding to the target audio signal.
[0034] Among them, the target long-time stationarity index can be used to indicate the degree of noise stability in the target audio signal. The target long-time signal-to-noise ratio can characterize the relative signal-to-noise ratio of the target audio signal; in other words, the target long-time signal-to-noise ratio characterizes the relative noise level of the target audio signal relative to the audio signal over a period of time.
[0035] Optionally, the target audio signal can be one or more frames of audio signal collected by an electronic device, or one or more frames of audio signal in an audio file, depending on the actual usage requirements.
[0036] For example, taking the target audio signal as the audio signal collected by an electronic device, during a call, the electronic device can perform frame-based processing on the collected audio signal. For instance, in real-time processing, the voice signal collected by the microphone of the electronic device is sent to the digital processing chip of the electronic device in real time, such as sending a 10ms long audio signal to the digital processing chip at a time. Since the voice signal is a short-term stable signal (approximately considered stable within 30ms) and a long-term non-stable signal, the voice signal within a certain duration can be regarded as a frame of audio signal; for example, every 30ms of voice signal is regarded as a frame of audio signal, that is, a processing frame. Specifically, the digital processing chip reads 10ms of audio signal each time, merges the read 10ms audio signal with the previously buffered audio signal, and after assembling an audio signal of about 30ms, it analyzes and processes the audio signal (i.e., a frame of audio signal).
[0037] For ease of description, unless otherwise specified, the target audio signal in the following embodiments is an audio signal collected by an electronic device.
[0038] Optionally, the "electronic device calculates the target long-time signal-to-noise ratio corresponding to the target audio signal" can be specifically implemented through the following steps A and B.
[0039] Step A: The electronic device determines N first instantaneous signal-to-noise ratios based on the instantaneous signal-to-noise ratios of N sets of historical audio signals.
[0040] Among them, each of the above N sets of historical audio signals may include M historical audio signals, and the N first instantaneous signal-to-noise ratios correspond one-to-one with the N sets of historical audio signals; M and N are both positive integers.
[0041] For example, M can be an integer greater than 1, and N can be any integer from 5 to 10.
[0042] In this embodiment of the application, the electronic device can divide the collected M adjacent audio signals into a group, that is, the audio signals from the 1st frame to the Mth frame are the 1st group; the audio signals from the M+1th frame to the 2Mth frame are the 2nd group, and so on.
[0043] Optionally, the aforementioned N sets of historical audio signals can be the most recently acquired N sets of historical audio signals. Assuming that the electronic device has already acquired Q frames of audio signals before acquiring the target audio signal, where Q = W1*M + W2, W1 is an integer greater than or equal to N, and W2 is a positive integer less than M, meaning the electronic device has acquired W1 sets of historical audio signals, then: the "above N sets of historical audio signals" refers to the most recently acquired N sets of historical audio signals from these W1 sets.
[0044] For example, assuming N is 10, and assuming that before collecting the target audio signal, the electronic device has collected 20 (W1 = 20) sets of historical audio signals, and in chronological order they are: set 1 to set 20, then the above N sets of historical audio signals include set 11 to set 20.
[0045] In this embodiment of the application, "instantaneous signal-to-noise ratio of N groups of historical audio signals" may include: the instantaneous signal-to-noise ratio of M frames of historical audio signals in each of the N groups of historical audio signals, that is, including the instantaneous signal-to-noise ratio of M*N frames of historical audio signals, for a total of M*N instantaneous signal-to-noise ratios.
[0046] Optionally, in one possible implementation, for each of the N sets of historical audio signals, the electronic device can determine the first instantaneous signal-to-noise ratio corresponding to each set of historical audio signals through step 1 below. That is, after the electronic device performs step 1 N times, it can obtain N first instantaneous signal-to-noise ratios.
[0047] Step 1: The electronic device determines the maximum instantaneous signal-to-noise ratio (SNR) among the instantaneous SNRs of each group of historical audio signals, and sets the maximum instantaneous SNR as the first instantaneous SNR corresponding to each group of historical audio signals.
[0048] In this embodiment, the electronic device can determine a first instantaneous signal-to-noise ratio (SNR) for every M frames of audio signals collected. Specifically, assuming the most recently collected set of historical audio signals by the electronic device is the Tth set of historical audio signals, the electronic device can determine the first instantaneous SNR (SNR) * snr corresponding to the Tth set of historical audio signals in one or another manner. M (T), where T is a positive integer.
[0049] In one approach, the first instantaneous signal-to-noise ratio (SNR) corresponding to the T-th group of historical audio signals is... M (T) can be expressed by the following formula (1):
[0050]
[0051] In formula (1), f is a multiple of M, and Let f be the instantaneous signal-to-noise ratio (SNR) of the f-th frame of the T-th historical audio signal (specifically, the smoothed instantaneous SNR).
[0052] As can be seen from Formula 1, in one approach, before determining the maximum instantaneous signal-to-noise ratio among the instantaneous signal-to-noise ratios of each group of historical audio signals, the electronic device needs to store the instantaneous signal-to-noise ratio of each frame of audio signal in each group of historical audio signals.
[0053] In another approach, during the acquisition of the T-th group of historical audio signals, the electronic device can compare the instantaneous signal-to-noise ratio (SNR) of the j-th frame audio signal in the T-th group of historical audio signals with the instantaneous SNR of the (j-1)-th frame audio signal; delete the smaller instantaneous SNR and retain the larger one; then compare the larger instantaneous SNR with the instantaneous SNR of the (j+1)-th frame audio signal in the T-th group of historical audio signals; and so on, until the most recently retained instantaneous SNR is compared with the instantaneous SNR of the M-th frame audio signal in the T-th group of historical audio signals, and the larger of the two is taken as the maximum instantaneous SNR in the T-th group of historical audio signals.
[0054] Thus, in another approach, since the instantaneous signal-to-noise ratios of two adjacent frames of audio signals in the same set of historical audio signals can be compared and a larger instantaneous signal-to-noise ratio can be retained, the number of instantaneous signal-to-noise ratio buffers can be saved.
[0055] It is understandable that the electronic device updates the SNR once every M frames of audio signal collected. M (T). That is, a first instantaneous noise ratio is obtained for every M frames of audio signal collected.
[0056] Optionally, assume that the N first instantaneous signal-to-noise ratios constitute an N-dimensional array snr matrix Then, the above N first instantaneous signal-to-noise ratios can be expressed by the following formula (2):
[0057] snr matrix =[snr M (T-N+1),snr M (T-N+2),…,snr M (T)] (2);
[0058] In formula (2), snr M (T) represents the first instantaneous signal-to-noise ratio corresponding to the Tth group of historical audio signals.
[0059] Optionally, the electronic device can construct an N-dimensional array to store the first instantaneous signal-to-noise ratios (SNRs) corresponding to the most recent N sets of historical speech signals. Thus, each time the electronic device determines a first instantaneous SNR, it can update this N-dimensional array. Specifically, the electronic device can remove the first instantaneous SNR from the N-dimensional array and add the most recently determined first instantaneous SNR to the N-dimensional array. In this way, the electronic device can directly use the N first instantaneous SNRs in the N-dimensional array to determine the target long-term SNR of the current frame of speech signal.
[0060] In the above embodiment, the electronic device determines the largest signal-to-noise ratio among the M instantaneous signal-to-noise ratios of each group of historical audio signals as the first instantaneous signal-to-noise ratio corresponding to each group of historical audio signals. In actual implementation, the second largest signal-to-noise ratio or the average instantaneous signal-to-noise ratio among the M instantaneous signal-to-noise ratios of each group of historical audio signals can also be determined as the first instantaneous signal-to-noise ratio corresponding to each group of historical audio signals.
[0061] Thus, since the maximum instantaneous signal-to-noise ratio of each group of historical audio signals can characterize the relative intensity of the speech signal and noise signal of that group of historical audio signals, and the N first instantaneous signal-to-noise ratios are the maximum instantaneous signal-to-noise ratios of the N groups of historical audio signals, the N first instantaneous signal-to-noise ratios can accurately characterize the quality of the speech signal of the audio sequence to which the target audio signal belongs.
[0062] The following example illustrates how an electronic device determines the instantaneous signal-to-noise ratio (SNR) of an audio signal.
[0063] Specifically, the electronic device can determine the instantaneous signal-to-noise ratio of the target audio signal through the following steps i to iii.
[0064] Step i: The electronic device first performs a Fast Fourier Transform (TTF) on the target audio signal, that is, transforms the target audio signal to the frequency domain to obtain the target time-frequency signal X(t,k) of the target audio signal. Here, t represents the time frame of the target audio signal, and k represents the k-th frequency point in the target audio signal.
[0065] Step ii: The electronic device determines the total signal energy Esignal(t) of the target audio signal based on the target time-frequency signal X(t,k). Esignal(t) can be expressed by the following formula (3):
[0066]
[0067] In formula (3), B is the number of frequency points in the target audio signal, t represents the time frame of the target time-frequency signal, and k represents the k-th frequency point in the target time-frequency signal X(t,k).
[0068] Step iii: The electronic device calculates the noise signal Noise(t,k) of the target audio signal based on the target time-frequency signal X(t,k); and determines the total noise energy Enoise(t) of the target audio signal based on Noise(t,k). Here, k represents the k-th frequency point in the noise signal Noise(t,k).
[0069] For specific methods on determining the noise signal Noise(t,k) and the total noise energy Enoise(t) for electronic devices, please refer to the relevant descriptions in related technologies. For example, electronic devices can determine the noise signal Noise(t,k) of the target audio signal based on methods such as recursive averaging algorithms based on the signal existence probability.
[0070] Step iiii: The electronic device determines the instantaneous signal-to-noise ratio (SNR) of the target audio signal based on the noise signal Noise(t,k) and the total signal energy Esignal(t) of the target audio signal. c (t).
[0071] Among them, the instantaneous signal-to-noise ratio (SNR) of the target audio signal c (t) can be expressed by the following formula (4):
[0072]
[0073] At this point, the electronic device obtains the instantaneous signal-to-noise ratio of the audio signal in frame t.
[0074] Furthermore, electronic devices can measure the instantaneous signal-to-noise ratio (SNR) of the target audio signal. c (t) is smoothed to obtain the final instantaneous signal-to-noise ratio of the target audio signal. The final instantaneous signal-to-noise ratio It can be expressed by the following formula (5):
[0075]
[0076] In formula (5), α is the smoothing factor. Let α be the final instantaneous signal-to-noise ratio of the (t-1)th frame of the audio signal (i.e., the frame preceding the target audio signal). For example, the value of α can range from 0 to 0.3.
[0077] It can be understood that the instantaneous signal-to-noise ratio of each group of historical audio signals can specifically include: the final instantaneous signal-to-noise ratio of the M frames of audio signals in that group of historical audio signals.
[0078] Step B: The electronic device determines the target long-term signal-to-noise ratio based on N first instantaneous signal-to-noise ratios.
[0079] Optionally, the electronic device can determine the average signal-to-noise ratio of N first instantaneous signal-to-noise ratios and determine the average signal-to-noise ratio as the target long-term signal-to-noise ratio.
[0080] Specifically, after the electronic device determines the average signal-to-noise ratio (SNR) of N first instantaneous SNRs, it can first smooth the average SNR, and then determine the smoothed SNR as the target long-time SNR, which is snr. l (t) can be expressed by the following formula (6):
[0081] snr l (t)=(1-μ)*snr l (t-1)+μ*mean(snr matrix (T)) (6);
[0082] In formula (6), snr l (t-1) represents the long-time signal-to-noise ratio (SNR) of the previous frame of the current frame's speech signal, and snr matrix (T) represents the signal-to-noise ratio of N first instants, and μ is a smoothing factor with a value range of 0 to 0.1.
[0083] Of course, electronic devices can also determine the target long-term signal-to-noise ratio based on N first instantaneous signal-to-noise ratios and by using any other possible methods; for example, electronic devices can determine the target long-term signal-to-noise ratio by taking the square root of the signal-to-noise ratios in the first set of signal-to-noise ratios.
[0084] Optionally, the electronic device determines the target long-term signal-to-noise ratio (SNR) based on N first instantaneous SNRs and second instantaneous SNRs; wherein the second instantaneous SNR can be the instantaneous SNR of the target audio signal. Specifically, the electronic device can determine the average SNR between the N first instantaneous SNRs and second instantaneous SNRs, then smooth the average SNR, and determine the smoothed average SNR as the target long-term SNR.
[0085] Thus, since the electronic device can determine the target long-term signal-to-noise ratio (SNR) of the target audio signal based on N first instantaneous SNRs, this target long-term SNR can better characterize the relative stability of noise in the target audio signal. This allows for a more accurate determination of the acoustic scene corresponding to the target audio signal.
[0086] Optionally, the "electronic device estimates the target long-term stationarity index corresponding to the target audio signal" can be achieved through the following steps C and D.
[0087] Step C: The electronic device determines the signal energy difference between the first signal energy and the second signal energy.
[0088] Step D: The electronic device smooths the signal energy difference to obtain the target long-term stability index.
[0089] The first signal energy is the signal energy after the target audio signal has undergone stationary noise reduction processing, and the second signal energy is the signal energy after the target audio signal has undergone deep learning noise reduction processing.
[0090] Optionally, the signal energy difference M in step C st (t) can be expressed by the following formula (7):
[0091] M st (t)=max(10*log10(Es(t))-10*log10(Et(t)), 0) (7);
[0092] In formula (7), Es(t) represents the first signal energy and Et(t) represents the second signal energy.
[0093] It should be noted that for stationary noise, both stationary noise reduction (also known as traditional signal processing) and deep learning noise reduction can effectively suppress stationary noise. In other words, the difference between the two noise reduction methods for stationary noise is relatively small, i.e., M... st (t) is close to 0. Stationary noise reduction has a weaker ability to suppress non-stationary noise, while deep learning noise reduction has a stronger ability to suppress non-stationary noise; that is, for non-stationary noise, the energy of the first signal is greater than the energy of the second signal, thus M... st (t) is relatively large, M st The value of (t) depends on the difference in noise energy suppression between the two.
[0094] Typically, if the target audio signal is a stationary speech signal, then M st (t) is generally close to 0, and the target audio signal is non-stationary speech, then M st (t) can be a few decibels (dB).
[0095] In this embodiment of the application, due to M st (t) represents the instantaneous stability of the target audio signal, which can be measured by electronic devices. st (t) is smoothed to obtain the target long-term stability index, which can indicate the stability of noise over a period of time, i.e. the relative stability of noise in the target audio signal.
[0096] Among them, the electronic device can use the following formula (8) to control M. st (t) is smoothed:
[0097] MS stp (t)=(1-β)*MSstp (t-1)+β*M st (t) (8);
[0098] In formula (8), MS stp (t) represents the target long-term stationarity index, MS stp (t-1) is the long-term stationarity index corresponding to the audio signal of the (t-1)th frame (i.e. the audio signal preceding the target audio signal), and β is the smoothing factor, the value of β can be from 0 to 0.1.
[0099] It should be noted that MS stp The larger (t) is, the more unstable the noise in the target audio signal.
[0100] The following is an exemplary description of a method for determining the first signal energy in an electronic device.
[0101] First, the electronic device can determine the stationary noise floor of the target audio signal based on stationary noise reduction methods (also known as traditional signal processing methods), such as minimum tracking methods and histogram methods. Then, based on this stationary noise floor, the corresponding stationary noise reduction gain (hereinafter referred to as the first frequency gain) Gs(t,k) of the target audio signal is determined. Since Gs(t,k) is calculated based on the stationary noise floor, it can only suppress stationary noise in the target audio signal.
[0102] It is understandable that there are many methods to determine Gs(t,k), such as Wiener filtering, equalization algorithms (Minimum Mean Square Error, MMSE), etc. For details, please refer to the relevant technologies, which will not be introduced in detail here.
[0103] Secondly, the electronic device uses the target time-frequency signal X(t,k) and the first frequency gain G to determine the frequency. s (t,k), determine the first signal energy Es(t), specifically, the first signal energy Es(t) can be expressed by the following formula (9):
[0104]
[0105] In formula (9), B is the number of frequency points in the target audio signal.
[0106] The following is an exemplary description of a method for determining the energy of a second signal in an electronic device.
[0107] First, the electronic device can determine the gain G at the second frequency point based on a deep learning mask noise reduction algorithm. mask (t,k).
[0108] It is understandable that denoising algorithms based on deep learning mask computation are the current mainstream denoising methods, as they have a certain ability to suppress both stationary and non-stationary noise.
[0109] For a detailed description of the noise reduction algorithm based on deep learning mask computation, please refer to the relevant technologies.
[0110] Secondly, the electronic device uses the target time-frequency signal X(t,k) and the second frequency gain G. mask (t,k), determine the second signal energy Et(t), specifically, the second signal energy Et(t) can be expressed by the following formula (10):
[0111]
[0112] In formula (10), B is the number of frequency points in the target audio signal.
[0113] Step 102: The electronic device determines the target acoustic scene to which the target audio signal belongs based on the target long-term signal-to-noise ratio and the target long-term stability index.
[0114] Optionally, in the embodiments of this application, step 102 described above can be implemented by step 102a or step 102d as described below.
[0115] Step 102a: When the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is greater than or equal to the stability index threshold, the electronic device determines the target acoustic scene as the first acoustic scene.
[0116] Step 102b: When the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is less than the stability index threshold, the electronic device determines the target acoustic scene as the second acoustic scene.
[0117] Step 102c: When the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to the stability index threshold, the electronic device determines the target acoustic scene as the third acoustic scene.
[0118] Step 102d: When the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is less than the stability index threshold, the electronic device determines the target acoustic scene as the fourth acoustic scene.
[0119] For example, the signal-to-noise ratio threshold can be 15 dB; the stability index threshold is 2 dB.
[0120] Optionally, both the signal-to-noise ratio threshold and the stability index threshold can be adjusted.
[0121] For example, assuming the signal-to-noise ratio threshold is thr_snr and the stationarity index threshold is thr_ms, then as shown in Table 1:
[0122] Table 1
[0123] <![CDATA[snr l (t)≥thr_snr&MS st (t)≥thr_ms]]> High signal-to-noise ratio non-stationary noise <![CDATA[snr l (t)≥thr_snr&MS st (t)<thr_ms]]> High signal-to-noise ratio stable noise <![CDATA[snr l (t)<thr_snr&MS st (t)≥thr_ms]]> Low signal-to-noise ratio non-stationary noise <![CDATA[snr l (t)<thr_snr&MS st (t)<thr_ms]]> Low signal-to-noise ratio stable noise
[0124] As shown in Table 1, when the target's long-term signal-to-noise ratio (SNR) is... l (t)≥thr_snr, and the target long-term stability index MS st When (t)≥thr_ms, the acoustic scene is a high signal-to-noise ratio, non-stationary noise; that is, the first acoustic scene.
[0125] When the long-term signal-to-noise ratio (SNR) l (t)≥thr_snr and the stationarity measure MS st When (t) < thr_ms, the acoustic scene is a high signal-to-noise ratio, stable noise; that is, the second acoustic scene.
[0126] When the long-term signal-to-noise ratio (SNR) l (t) < thr_snr and the stationarity measure MS st When (t)≥thr_ms, the acoustic scene is a low signal-to-noise ratio, non-stationary noise; that is, the third acoustic scene.
[0127] When the long-term signal-to-noise ratio (SNR) l (t) < thr_snr and the stationarity measure MS st When (t) < thr_ms, the acoustic scene is a low signal-to-noise ratio, with stable noise; that is, the fourth acoustic scene.
[0128] Where thr_snr is an adjustable signal-to-noise ratio threshold, and thr_ms is an adjustable stability index threshold.
[0129] For example, the signal-to-noise ratio threshold can be any value within 15dB ± c; the stability index threshold can be 2dB ± d, where c and d are determined according to actual usage requirements.
[0130] Thus, since the first, second, third, and fourth acoustic scenarios can cover all acoustic scenarios in reality, the universality and applicability of acoustic scenario classification are improved.
[0131] Step 103: The electronic device performs noise reduction processing on the target audio signal based on the target acoustic scene.
[0132] It is understood that in the audio noise reduction method provided in the embodiments of this application, the acoustic scene is classified into four types, namely, the first acoustic scene, the second acoustic scene, the third acoustic scene and the fourth acoustic scene. Since there are few types of acoustic scenes, the electronic device can perform noise reduction processing on the audio signals in each acoustic scene in a targeted manner, that is, perform targeted noise reduction processing on different acoustic scenes.
[0133] For example, assuming that the noise reduction strategy for each acoustic scenario includes: deep learning noise reduction processing and noise stabilization processing, then:
[0134] In the noise reduction strategy corresponding to the first acoustic scene, the weight of the stationary noise reduction processing is less than the weight of the deep learning noise reduction processing, and the noise suppression ratio is the first ratio.
[0135] In the noise reduction strategy corresponding to the second acoustic scene, the weight of the stationary noise reduction processing is greater than that of the deep learning noise reduction processing, and the noise suppression ratio is the second ratio.
[0136] In the noise reduction strategy corresponding to the third acoustic scenario, the weight of the stationary noise reduction processing is less than the weight of the deep learning noise reduction processing, and the noise suppression ratio is the third ratio.
[0137] In the fourth acoustic scenario, the corresponding noise reduction strategy is that the weight of the stationary noise reduction process is greater than that of the deep learning noise reduction process, and the noise suppression ratio is the fourth ratio.
[0138] Among them, the first proportion is less than the third proportion, and the first proportion is less than the fourth proportion; correspondingly, the second proportion is less than the third proportion, and the second proportion is less than the fourth proportion.
[0139] Optionally, the first ratio and the second ratio can be the same, and the third ratio and the fourth ratio can be the same.
[0140] Thus, since the noise reduction strategy corresponding to the target acoustic scene can be used to perform noise reduction processing on the target audio signal, the effect of noise reduction processing on the target audio signal is improved, thereby improving the sound quality of electronic devices.
[0141] Optionally, the electronic device can output the processed target audio signal after performing noise reduction processing on the target audio signal. For example, when the electronic device is making a call with a target device, i.e., the target audio signal is the voice signal acquired by the electronic device during the call, the electronic device can send the processed target audio signal to the target device after performing noise reduction processing on the target audio signal.
[0142] It should be noted that for each frame of audio in the audio file or each frame of audio acquired by the electronic device, the electronic device can perform steps 101 to 103 as described above.
[0143] In the audio noise reduction method provided in this application embodiment, since the long-term signal-to-noise ratio and the stability index corresponding to the audio signal are two essential characteristics of noise in the audio signal, the target acoustic scene corresponding to the target audio signal can be determined more accurately and quickly based on the target long-term signal-to-noise ratio and the target long-term stability index corresponding to the target audio signal, thereby improving the accuracy of target audio noise reduction based on the target acoustic scene.
[0144] The following is in conjunction with the appendix Figure 2 The audio noise reduction method provided in the embodiments of this application will be described by way of example.
[0145] For example, taking the noise reduction processing of audio signals during a call by an electronic device as an example, the electronic device can use the audio noise reduction method provided in the embodiments of this application to perform noise reduction processing on each frame of audio signal during the call. Figure 2 As shown, the audio signal noise reduction processing method may include the following steps 201 to 219.
[0146] Step 201: The electronic device reads in the voice signal and performs frame segmentation on the read in the voice signal.
[0147] For example, in real-time processing, each audio signal captured by the microphone is sent to the digital processing chip of the electronic device in real time, such as 10ms of data at a time. Since audio signals are short-term stationary (approximately considered stationary within 30ms) but long-term non-stationary, the electronic device can analyze relatively short audio signals, for example, taking a frame of approximately 30ms as the processing frame. That is, by reading in 10ms of data at a time and buffering previously read audio data, approximately 30ms of audio data is collected for analysis and processing.
[0148] As can be seen, the purpose of framing the speech signal is to treat each fixed duration (e.g., 30ms) of speech signal as a processing frame, also known as a frame of speech signal.
[0149] Furthermore, each frame of audio signal obtained from the frame-segmentation process is a time-domain signal.
[0150] Step 202: The electronic device performs an FFT transform on the time-domain signal of the current frame audio signal to obtain the time-frequency signal of the current frame audio signal.
[0151] It can be understood that the current frame voice signal is the target audio signal in the above embodiment.
[0152] Step 203: The electronic device determines the total signal energy of the current frame voice signal based on the time-frequency signal of the current frame voice signal.
[0153] Step 204: The electronic device can calculate the noise signal of the current frame speech signal based on the target time-frequency signal; and determine the total noise energy of the current frame speech signal based on the noise signal.
[0154] For example, electronic devices can calculate the noise signal of the current frame's speech signal using methods such as recursive averaging algorithms based on the probability of signal presence.
[0155] Step 205: The electronic device determines the instantaneous signal-to-noise ratio (SNR) of the current frame's audio signal based on the total signal energy and total noise energy. Since the instantaneous SNR does not reflect the signal-noise energy level over a period of time, the long-term SNR is more suitable for acoustic scene classification. The method for calculating the long-term SNR based on the instantaneous SNR is described below.
[0156] Step 206: The electronic device smooths the instantaneous signal-to-noise ratio of the current frame speech signal to obtain the final instantaneous signal-to-noise ratio of the current frame speech signal.
[0157] In this embodiment of the application, the electronic device can divide each M (e.g., 100) frames of voice signal during a call into a group of historical voice signals.
[0158] It is understandable that the instantaneous signal-to-noise ratio (SNR) of the current frame speech signal can be used to determine the long-term SNR of the current frame speech signal, or it can be used to confirm the long-term SNR of the speech signal read after the current frame speech signal.
[0159] The long-time signal-to-noise ratio (SNR) of a speech signal can characterize the relative noise level of the speech signal relative to audio signals over a period of time (such as multiple frames of historical speech signals).
[0160] It should be noted that the electronic device can perform steps 302 to 306 on each frame of voice signal during the call to obtain the final instantaneous signal-to-noise ratio of each frame of voice signal.
[0161] Step 207: Electronically determine the maximum instantaneous signal-to-noise ratio (SNR) among the instantaneous SNRs of each group of historical speech signals, and set the maximum instantaneous SNR as the first instantaneous SNR corresponding to each group of historical speech signals.
[0162] For the method of determining the maximum instantaneous signal-to-noise ratio among the instantaneous signal-to-noise ratios of each group of historical speech signals for electronic devices, please refer to the relevant description in the above embodiments.
[0163] Step 208: Construct an N-dimensional array for the electronic device.
[0164] The N-dimensional array includes N first instantaneous signal-to-noise ratios, which are the first instantaneous signal-to-noise ratios corresponding to the most recent N sets of historical voice signals during the current call.
[0165] In this embodiment of the application, the electronic device can update the N-dimensional array once for every M frames of audio signal collected.
[0166] Step 209: The electronic device determines the average signal-to-noise ratio of the N first instantaneous signal-to-noise ratios in the N-dimensional array, and smooths the average signal-to-noise ratio, and determines the smoothed average signal-to-noise ratio as the target long-term signal-to-noise ratio.
[0167] Generally, current mainstream noise reduction algorithms combine deep learning methods with traditional noise reduction methods. Since deep learning-based mask estimation methods are effective at suppressing non-stationary noise, while traditional noise reduction algorithms have very limited ability to suppress non-stationary noise, this application proposes using the noise suppression differences between these two noise reduction algorithms to determine a measure of the stationarity of the speech signal during a call.
[0168] Step 210: The electronic device estimates the stationary noise floor of the current frame speech signal.
[0169] There are many methods for determining the stationary noise floor of a speech signal, such as the minimum tracking method and the histogram method.
[0170] Step 211: The electronic device determines the stable noise reduction gain corresponding to the current frame speech signal based on the stable noise floor of the current frame speech signal.
[0171] Among them, the smooth noise reduction gain is used to suppress smooth noise in the speech signal.
[0172] The electronic device calculates the corresponding frequency gain based on the stationary noise floor of the current frame's speech signal, i.e., the stationary noise reduction gain G. s The gain (t,k) is calculated based on the stationary noise floor, so it can only suppress stationary noise. There are many methods for calculating stationary noise reduction gain, such as Wiener filtering and MMSE, but these are not the focus of this application and will not be described in detail.
[0173] Step 212: The electronic device uses a stable noise reduction gain to perform noise reduction processing on the current frame speech signal to obtain the first signal energy after stable noise reduction processing.
[0174] Step 213: The electronic device estimates the non-stationary noise floor of the current frame speech signal.
[0175] For example, electronic devices use deep learning noise reduction algorithms to process the current frame of speech signal and obtain the non-stationary noise floor of the current frame of speech signal.
[0176] Step 214: The electronic device determines the non-stationary noise reduction gain corresponding to the current frame speech signal based on the non-stationary noise floor of the current frame speech signal.
[0177] Among them, the non-stationary noise reduction gain has a certain ability to suppress both non-stationary and stationary noise in the current frame speech signal.
[0178] Step 215: The electronic device uses non-stationary noise reduction gain to perform noise reduction processing on the current frame speech signal to obtain the second signal energy after non-stationary noise reduction processing.
[0179] Step 216: The electronic device determines the signal energy difference between the first signal energy and the second signal energy.
[0180] The energy difference of the signal can be used as an indicator of the instantaneous stability of the current frame's speech signal.
[0181] It's understandable that for stationary noise, both stationary noise reduction and deep learning noise reduction can achieve good results, meaning the difference between the two methods is small, with the signal energy difference approaching zero. For non-stationary noise, stationary noise reduction is weaker, while deep learning noise reduction is stronger. Specifically, the energy Et(t) after deep learning noise reduction is less than the energy Es(t) after stationary noise suppression noise reduction, resulting in a signal energy difference greater than zero. This value depends on the difference in noise energy suppression between the two methods. Generally, stationary speech energy is close to zero, while non-stationary speech energy is several dB.
[0182] Since the signal energy difference represents the instantaneous stability of the current frame, it is not conducive to practical use. Therefore, the signal energy difference can be smoothed to obtain the noise stability of the current frame's speech signal over a period of time.
[0183] Step 217: The electronic device smooths the signal energy difference to obtain the long-term stationarity index of the current frame speech signal.
[0184] Thus, we obtain the stationarity metric. If the value of this metric is close to 0, it indicates that the acoustic scene corresponding to the current frame of speech signal is a stationary noise type; when the value of this metric is large, it indicates that the acoustic scene corresponding to the current frame of speech signal is a non-stationary noise type.
[0185] Step 218: The electronic device determines the target acoustic scene corresponding to the current frame speech signal based on the target long-term signal-to-noise ratio and long-term stationarity index of the current frame speech signal.
[0186] Step 219: The electronic device performs noise reduction processing on the current frame speech signal based on the target acoustic scene.
[0187] For other descriptions of steps 201 to 219, please refer to the relevant descriptions in the above embodiments. To avoid repetition, they will not be repeated here.
[0188] The audio noise reduction method provided in this application can be executed by an audio noise reduction device or a control module within that device for performing the audio noise reduction method. This application uses an audio noise reduction device performing the audio noise reduction method as an example to illustrate the audio noise reduction device provided in this application.
[0189] This application provides an audio noise reduction device. Figure 3 This paper illustrates a possible structural diagram of the audio noise reduction device provided in an embodiment of this application, such as... Figure 3 As shown, the audio noise reduction device 300 may include a processing module 301 and a determining module 302. The processing module is used to calculate the target long-term signal-to-noise ratio (SNR) and the target long-term stability index corresponding to the target audio signal, wherein the target long-term stability index indicates the stability of noise in the target audio signal. The determining module is used to determine the target acoustic scene corresponding to the target audio signal based on the target SNR and the target long-term stability index calculated by the processing module. The processing module is also used to perform noise reduction processing on the target audio signal based on the target acoustic scene determined by the determining module.
[0190] In one possible implementation, the determining module is specifically used to: determine the target acoustic scene as the first acoustic scene when the target long-time signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold and the target long-time stability index is greater than or equal to the stability index threshold;
[0191] If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the second acoustic scene.
[0192] If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the third acoustic scene.
[0193] If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the fourth acoustic scene.
[0194] In one possible implementation, the processing module is specifically used to determine N first instantaneous signal-to-noise ratios based on the instantaneous signal-to-noise ratios of N sets of historical audio signals; and to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios.
[0195] Each set of historical audio signals includes M historical audio signals, and the N first instantaneous signal-to-noise ratios correspond one-to-one with the N sets of historical audio signals; M and N are both positive integers.
[0196] In one possible implementation, the processing module is specifically configured to determine the maximum instantaneous signal-to-noise ratio (SNR) among the instantaneous SNRs of each group of historical audio signals, and to determine the maximum instantaneous SNR as the first instantaneous SNR corresponding to each group of historical audio signals.
[0197] In one possible implementation, the processing module is specifically used to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios and the second instantaneous signal-to-noise ratios;
[0198] Wherein, the second instantaneous signal-to-noise ratio is the instantaneous signal-to-noise ratio of the target audio signal.
[0199] In one possible implementation, the processing module is specifically used to determine the average signal-to-noise ratio of the N first instantaneous signal-to-noise ratios, and to determine the average signal-to-noise ratio as the target long-term signal-to-noise ratio.
[0200] In one possible implementation, the processing module is specifically used to determine the signal energy difference between the first signal energy and the second signal energy; and to smooth the signal energy difference to obtain the target long-term stability index.
[0201] Wherein, the first signal energy is the signal energy after the target audio signal has undergone stationary noise reduction processing, and the second signal energy is the signal energy after the target audio signal has undergone deep learning noise reduction processing.
[0202] In this embodiment, since the long-term signal-to-noise ratio and the stability index corresponding to the audio signal are two essential characteristics of noise in the audio signal, the target acoustic scene corresponding to the target audio signal can be determined more accurately and quickly based on the target long-term signal-to-noise ratio and the target long-term stability index corresponding to the target audio signal, thereby improving the accuracy of target audio noise reduction based on the target acoustic scene.
[0203] The audio noise reduction device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0204] The audio noise reduction device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0205] The audio noise reduction device provided in this application embodiment can achieve... Figure 1 and Figure 2 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0206] Optionally, such as Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described audio noise reduction method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0207] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0208] Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0209] The electronic device 500 includes, but is not limited to, at least some of the following components: radio frequency unit 501, network module 502, audio output unit 503, input unit 504, sensor 505, display unit 506, user input unit 507, interface unit 508, memory 509, and processor 510.
[0210] Those skilled in the art will understand that the terminal 500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 5 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0211] The processor 510 is used to calculate the target long-term signal-to-noise ratio and the target long-term stationarity index corresponding to the target audio signal, wherein the target long-term stationarity index is used to indicate the degree of stability of noise in the target audio signal.
[0212] The processor 510 is further configured to determine the target acoustic scene corresponding to the target audio signal based on the target long-term signal-to-noise ratio and the target long-term stability index;
[0213] The processor 510 is further configured to determine the target acoustic scene based on the processor 510 and perform noise reduction processing on the target audio signal.
[0214] In one possible implementation, the processor 510 is specifically configured to: determine the target acoustic scene as a first acoustic scene when the target long-term signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to a stability index threshold;
[0215] If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the second acoustic scene.
[0216] If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the third acoustic scene.
[0217] If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the fourth acoustic scene.
[0218] In one possible implementation, the processor 510 is specifically configured to determine N first instantaneous signal-to-noise ratios based on the instantaneous signal-to-noise ratios of N sets of historical audio signals; and to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios;
[0219] Each set of historical audio signals includes M historical audio signals, and the N first instantaneous signal-to-noise ratios correspond one-to-one with the N sets of historical audio signals; M and N are both positive integers.
[0220] In one possible implementation, the processor 510 is specifically configured to determine the maximum instantaneous signal-to-noise ratio among the instantaneous signal-to-noise ratios of each group of historical audio signals, and to determine the maximum instantaneous signal-to-noise ratio as the first instantaneous signal-to-noise ratio corresponding to each group of historical audio signals.
[0221] In one possible implementation, the processor 510 is specifically configured to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios and second instantaneous signal-to-noise ratios;
[0222] Wherein, the second instantaneous signal-to-noise ratio is the instantaneous signal-to-noise ratio of the target audio signal.
[0223] In one possible implementation, the processor 510 is specifically configured to determine the average signal-to-noise ratio of the N first instantaneous signal-to-noise ratios, and to determine the average signal-to-noise ratio as the target long-term signal-to-noise ratio.
[0224] In one possible implementation, the processor 510 is specifically used to determine the signal energy difference between the first signal energy and the second signal energy; and to smooth the signal energy difference to obtain the target long-term stability index.
[0225] Wherein, the first signal energy is the signal energy after the target audio signal has undergone stationary noise reduction processing, and the second signal energy is the signal energy after the target audio signal has undergone deep learning noise reduction processing.
[0226] In this embodiment, since the long-term signal-to-noise ratio and the stability index corresponding to the audio signal are two essential characteristics of noise in the audio signal, the target acoustic scene corresponding to the target audio signal can be determined more accurately and quickly based on the target long-term signal-to-noise ratio and the target long-term stability index corresponding to the target audio signal, thereby improving the accuracy of target audio noise reduction based on the target acoustic scene.
[0227] It should be understood that, in this embodiment, the input unit 504 may include a graphics processing unit (GPU) 5041 and a microphone 5042. The GPU 5041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 506 may include a display panel 5061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include a touch detection device and a touch controller. Other input devices 5072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0228] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 501 can transmit it to the processor 510 for processing; in addition, the radio frequency unit 501 can send uplink data to the network-side device. Typically, the radio frequency unit 501 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0229] The memory 509 can be used to store software programs or instructions, as well as various data. The memory 509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0230] Processor 510 may include one or more processing units; optionally, processor 510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 510.
[0231] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio noise reduction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0232] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0233] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio noise reduction method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0234] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio noise reduction method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0235] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0236] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0237] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0238] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio noise reduction method, characterized by, The method includes: Calculate the target long-time signal-to-noise ratio and the target long-time stationarity index corresponding to the target audio signal. The target long-time signal-to-noise ratio is used to characterize the relative noise level of the target audio signal over a period of time, and the target long-time stationarity index is used to indicate the degree of stability of noise in the target audio signal. The target acoustic scene corresponding to the target audio signal is determined based on the target long-term signal-to-noise ratio and the target long-term stability index. Based on the target acoustic scene, the target audio signal is subjected to noise reduction processing; The step of determining the target acoustic scene corresponding to the target audio signal based on the target long-term signal-to-noise ratio and the target long-term stability index includes: If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the first acoustic scene. If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the second acoustic scene. If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the third acoustic scene. If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the fourth acoustic scene. Calculate the target long-term stationarity index corresponding to the target audio signal, including: Determine the signal energy difference between the first signal energy and the second signal energy; The signal energy difference is smoothed to obtain the target long-term stability index; Wherein, the first signal energy is the signal energy after the target audio signal has undergone stationary noise reduction processing, and the second signal energy is the signal energy after the target audio signal has undergone deep learning noise reduction processing.
2. The method according to claim 1, characterized in that, Calculate the target long-time signal-to-noise ratio corresponding to the target audio signal, including: N first instantaneous signal-to-noise ratios are determined based on the instantaneous signal-to-noise ratios of N sets of historical audio signals; The target long-term signal-to-noise ratio is determined based on the N first instantaneous signal-to-noise ratios; Each group of historical audio signals includes M historical audio signals, and the N first instantaneous signal-to-noise ratios correspond one-to-one with the N groups of historical audio signals. Both M and N are positive integers.
3. The method according to claim 2, characterized in that, The determination of N first instantaneous signal-to-noise ratios based on the instantaneous signal-to-noise ratios of N sets of historical audio signals includes: The maximum instantaneous signal-to-noise ratio (SNR) among the instantaneous SNRs of each group of historical audio signals is determined, and the maximum instantaneous SNR is determined as the first instantaneous SNR corresponding to each group of historical audio signals.
4. The method according to claim 2, characterized in that, The determination of the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios includes: The target long-term signal-to-noise ratio is determined based on the N first instantaneous signal-to-noise ratios and the second instantaneous signal-to-noise ratios; Wherein, the second instantaneous signal-to-noise ratio is the instantaneous signal-to-noise ratio of the target audio signal.
5. The method according to claim 2, characterized in that, Determining the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios includes: Determine the average signal-to-noise ratio of the N first instantaneous signal-to-noise ratios, and determine the average signal-to-noise ratio as the target long-term signal-to-noise ratio.
6. An acoustic scene classification device, characterized in that, The device includes: a processing module and a determination module; The processing module is used to calculate the target long-time signal-to-noise ratio and the target long-time stationarity index corresponding to the target audio signal. The target long-time signal-to-noise ratio is used to characterize the relative noise level of the target audio signal over a period of time, and the target long-time stationarity index is used to indicate the degree of stability of noise in the target audio signal. The determining module is used to determine the target acoustic scene corresponding to the target audio signal based on the target long-term signal-to-noise ratio and the target long-term stability index calculated by the processing module. The processing module is further configured to perform noise reduction processing on the target audio signal based on the target acoustic scene determined by the determining module; The determining module is specifically used for: If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the first acoustic scene. If the target long-term signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the second acoustic scene. If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is greater than or equal to the stability index threshold, the target acoustic scene is determined to be the third acoustic scene. If the target long-term signal-to-noise ratio is less than the signal-to-noise ratio threshold and the target long-term stability index is less than the stability index threshold, the target acoustic scene is determined to be the fourth acoustic scene. The processing module is specifically used to determine the signal energy difference between the first signal energy and the second signal energy; and to smooth the signal energy difference to obtain the target long-term stability index. Wherein, the first signal energy is the signal energy after the target audio signal has undergone stationary noise reduction processing, and the second signal energy is the signal energy after the target audio signal has undergone deep learning noise reduction processing.
7. The apparatus according to claim 6, characterized in that, The processing module is specifically used to determine N first instantaneous signal-to-noise ratios based on the instantaneous signal-to-noise ratios of N sets of historical audio signals; and to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios. Each set of historical audio signals includes M historical audio signals, and the N first instantaneous signal-to-noise ratios correspond one-to-one with the N sets of historical audio signals; M and N are both positive integers.
8. The apparatus according to claim 7, characterized in that, The processing module is specifically used to determine the maximum instantaneous signal-to-noise ratio among the instantaneous signal-to-noise ratios of each group of historical audio signals, and to determine the maximum instantaneous signal-to-noise ratio as the first instantaneous signal-to-noise ratio corresponding to each group of historical audio signals.
9. The apparatus according to claim 7, characterized in that, The processing module is specifically used to determine the target long-term signal-to-noise ratio based on the N first instantaneous signal-to-noise ratios and second instantaneous signal-to-noise ratios; Wherein, the second instantaneous signal-to-noise ratio is the instantaneous signal-to-noise ratio of the target audio signal.
10. The apparatus according to claim 7, characterized in that, The processing module is specifically used to determine the average signal-to-noise ratio of the N first instantaneous signal-to-noise ratios, and to determine the average signal-to-noise ratio as the target long-term signal-to-noise ratio.
11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio noise reduction method as described in any one of claims 1 to 5.
12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio noise reduction method as described in any one of claims 1 to 5.