Audio noise reduction method, model training method, apparatus, device, and medium

A hybrid method combining traditional and AI noise reduction techniques through integrated activity detection and model training addresses the limitations of existing speech noise reduction methods, enhancing noise reduction capabilities and stability.

JP7860332B2Active Publication Date: 2026-05-15BIGO TECH PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BIGO TECH PTE LTD
Filing Date
2023-07-12
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing speech noise reduction methods, both traditional and AI-based, face limitations in effectively handling transient noise and maintaining performance in unforeseen scenarios with low signal-to-noise ratios, leading to unpredictable outputs and potential system collapse.

Method used

A hybrid approach combining traditional noise reduction methods with AI noise reduction methods, utilizing a pre-configured audio activity detection algorithm to integrate model and algorithm activity detection results, followed by noise estimation and reduction, and training the speech noise reduction network model based on integrated detection results to enhance noise reduction capabilities.

Benefits of technology

The hybrid method achieves superior noise reduction capabilities across various types of noise, improving speech clarity and stability by leveraging combined activity detection information from both traditional and AI models, resulting in enhanced signal-to-noise ratios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860332000001
    Figure 0007860332000001
  • Figure 0007860332000002
    Figure 0007860332000002
  • Figure 0007860332000003
    Figure 0007860332000003
Patent Text Reader

Abstract

A method for reducing audio noise, a model training method, an apparatus, a device, a medium, and a product. The method for reducing audio noise employs a preset voice activity detection algorithm to detect the current audio frame waiting for processing, and obtains a corresponding algorithm activity detection result [101]; performs an integration process on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame, obtains a target activity detection result corresponding to the current audio frame, and the model activity detection result is output from a preset audio noise reduction network model [102]; based on the target activity detection result, performs noise estimation and noise removal on the current audio frame to obtain an initial noise reduction audio frame [103]; inputs the initial noise reduction audio frame into a preset audio noise reduction network model to output a target noise reduction audio frame and a model activity detection result corresponding to the current audio frame [104]. By adopting the above solution, the audio noise reduction effect can be improved, and the stability and robustness of the audio noise reduction solution can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the priority of Chinese Patent Application No. 202210864010.4, filed with the China National Intellectual Property Administration on July 21, 2022, and all of its contents are incorporated herein by reference. This application further claims priority to international application PCT / CN2023 / 106951, filed on 12 July 2023, and all the contents of that application are incorporated into this application by reference.

[0002] This application relates to the field of audio processing technologies, for example, methods for reducing audio noise, model training methods, devices, apparatuses, and medium related thereto.

Background Art

[0003] With the rapid development of multimedia technologies, various conference, social, and entertainment applications have emerged one after another, among which, many scenarios such as voice calls, audio-video live broadcasts, and multi-person conferences are involved. The audio quality has become an important indicator for judging the performance of an application.

[0004] The audio collected by the microphone of a terminal device usually has a certain degree of noise, and the clarity and quality of the audio can be improved by suppressing the noise contained in the audio by an audio noise reduction algorithm.

[0005] Currently, speech noise reduction methods can be broadly divided into two types: traditional noise reduction methods and artificial intelligence (AI) noise reduction methods. Traditional noise reduction methods achieve speech noise reduction through signal processing techniques and cannot remove transient noise; that is, they have weak noise reduction capabilities against sudden noise. AI noise reduction methods have excellent noise reduction capabilities against both steady and transient noise, but because they are data-driven, they are highly dependent on training samples. If there are scenes that are not considered during the model training process (for example, situations with a low signal-to-noise ratio), encountering such scenes in actual application can lead to unpredictable signal outputs and potentially cause system collapse. [Overview of the project]

[0006] Embodiments of the present invention provide a speech noise reduction method, a model training method, an apparatus, a device, a medium, and a product that can effectively combine traditional noise reduction methods and AI noise reduction methods to improve the speech noise reduction effect.

[0007] According to one aspect of the present invention, a method for reducing audio noise is provided, and this method is This involves using a pre-configured audio activity detection algorithm to perform detection on the currently pending audio frame and obtaining the corresponding algorithm activity detection result. The model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame are integrated to obtain the target activity detection result corresponding to the current audio frame, and the model activity detection result is output from a pre-configured audio noise reduction network model. Based on the target activity detection results, noise estimation and noise reduction are performed on the current audio frame to obtain an initial noise-reduced audio frame. By inputting the initial noise reduction audio frame into the pre-configured audio noise reduction network model, the model activity detection results corresponding to the target noise reduction audio frame and the current audio frame are output. Includes.

[0008] In other aspects of this application, a model training method is provided, and this method is A pre-configured audio activity detection algorithm is used to perform detection on the current sample audio frame, and the corresponding sample algorithm activity detection result is obtained, confirming that the current sample audio frame is associated with an activity detection label and a pure audio frame. The sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame are integrated to obtain the target sample activity detection result corresponding to the current sample audio frame, and the sample model activity detection result is output from the audio noise reduction network model. The aforementioned objective Sample Activity Based on the detection results, noise estimation and noise reduction are performed on the current sample audio frame to obtain an initial noise-reduced sample audio frame. By inputting the initial noise-reduced sample audio frame into the speech noise reduction network model, the system outputs sample model activity detection results corresponding to the target sample noise-reduced audio frame and the current sample audio frame. Based on the target sample noise reduction audio frame and the pure audio frame, a first loss relationship is determined; based on the sample model activity detection result and the activity detection label, a second loss relationship is determined; and based on the first loss relationship and the second loss relationship, the speech noise reduction network model is trained. Includes.

[0009] In another aspect of the present invention, a speech noise reduction device is provided, the device comprising a speech activity detection module, a detection result integration module, a noise reduction processing module, and a model input module, The aforementioned audio activity detection module is configured to employ a pre-configured audio activity detection algorithm to perform detection on the currently pending audio frame and obtain the corresponding algorithm activity detection result. The aforementioned detection result integration module is configured to perform integration processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain a target activity detection result corresponding to the current audio frame, and the model activity detection result is output from a pre-configured audio noise reduction network model. The noise reduction processing module is configured to perform noise estimation and noise reduction on the current audio frame based on the target activity detection result, in order to obtain an initial noise-reduced audio frame. The model input module is configured to input the initial noise reduction audio frame to the pre-configured audio noise reduction network model, thereby outputting model activity detection results corresponding to the target noise reduction audio frame and the current audio frame.

[0010] In another aspect of the present invention, a model training device is provided, the device comprising a voice detection module, an integration module, a noise reduction module, a network model input module, and a network model training module, The aforementioned audio detection module is configured to perform detection on the currently pending sample audio frame using a pre-configured audio activity detection algorithm and to obtain the corresponding sample algorithm activity detection result. The current sample audio frame is associated with an activity detection label and a clean audio frame. The integration module is configured to perform integration processing on the sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame to obtain a target sample activity detection result corresponding to the current sample audio frame, and the sample model activity detection result is output from the audio noise reduction network model. The noise reduction module uses the target sample activity Based on the detection results, the system is configured to perform noise estimation and noise reduction on the current sample audio frame to obtain an initial noise-reduced sample audio frame. The network model input module is configured to input the initial noise reduction sample audio frame to the audio noise reduction network model and output sample model activity detection results corresponding to the target sample noise reduction audio frame and the current sample audio frame. The network model training module is configured to determine a first loss relationship based on the target sample noise reduction audio frame and the clean audio frame, determine a second loss relationship based on the sample model activity detection result and the activity detection label, and train the speech noise reduction network model based on the first loss relationship and the second loss relationship.

[0011] According to another aspect of the present application, an electronic device is provided, said electronic device is At least one processor, The memory that is connected to the aforementioned at least one processor and Includes, The memory stores a computer program that can be executed by the at least one processor, and by executing the computer program by the at least one processor, the at least one processor can perform the speech noise reduction method and / or model training method described in any embodiment of the present application.

[0012] In another aspect of the present application, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, the computer program being used to implement a speech noise reduction method and / or model training method described in any embodiment of the present application when executed by a processor.

[0013] In another aspect of the present application, a computer program product is provided, the computer program product includes a computer program, which, when executed by a processor, implements the speech noise reduction method and / or model training method described in any embodiment of the present application.

[0014] The audio noise reduction solution provided in the embodiment of this application employs a pre-configured audio activity detection algorithm to perform detection on the current audio frame awaiting processing, obtains the corresponding algorithm activity detection result, performs integrated processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame, obtains the target activity detection result corresponding to the current audio frame, the model activity detection result is output from a pre-configured audio noise reduction network model, noise estimation and noise reduction are performed on the current audio frame based on the target activity detection result, obtains an initial noise reduction audio frame, and inputs the initial noise reduction audio frame to a pre-configured audio noise reduction network model to output the target noise reduction audio frame and the model activity detection result corresponding to the current audio frame. By adopting the above approach, the pre-configured speech noise reduction network model can output model activity detection results. When processing the current audio frame using a traditional speech noise reduction algorithm, the model activity detection result of the previous audio frame can be combined with the algorithm activity detection result obtained by the traditional speech noise reduction algorithm. This allows the traditional noise reduction algorithm to obtain more activity detection information, enabling a more rational and accurate determination of speech activity detection results. By performing noise estimation and noise reduction based on these results, speech can be better protected, more noise can be removed, and a higher signal-to-noise ratio can be obtained from traditional noise reduction results. Furthermore, by using the traditional noise reduction results as input to the pre-configured speech noise reduction network model, a more effective noise reduction audio frame can be obtained, reducing the possibility of the pre-configured speech noise reduction network model processing poor data. The traditional noise reduction algorithm and the AI ​​noise reduction method mutually enhance each other, resulting in superior noise reduction capabilities against various types of noise, improving the speech noise reduction effect, and increasing the stability and robustness of the overall speech noise reduction solution.

[0015] The following is an explanation of the drawings that need to be used in the description of the embodiments. The drawings described below are only some embodiments of the present application. It is obvious to those skilled in the art that other drawings can also be obtained based on these drawings without any creative work.

Brief Explanation of the Drawings

[0016] [Figure 1] It is a schematic flowchart of the voice noise reduction method provided by the embodiment of the present application. [Figure 2] It is a schematic flowchart of another voice noise reduction method provided by the embodiment of the present application. [Figure 3] It is a schematic diagram of the inference flow of the voice noise reduction method provided by the embodiment of the present application. [Figure 4] It is a schematic flowchart of the model training method provided by the embodiment of the present application. [Figure 5] It is a schematic diagram of the training process of the model training method provided by the embodiment of the present application. [Figure 6] It is a structural block diagram of the voice noise reduction device provided by the embodiment of the present application. [Figure 7] It is a structural block diagram of the model training device provided by the embodiment of the present application. [Figure 8] It is a structural block diagram of the electronic device provided by the embodiment of the present application.

Modes for Carrying Out the Invention

[0017] To enable those skilled in the art to better understand the solution of the present application, the following will describe the embodiments of the present application in combination with the drawings in the embodiments of the present application. The described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without any creative work belong to the protection scope of the present application.

[0018] Furthermore, terms such as "First," "Second," etc., in the specification, claims, and drawings of this application do not necessarily indicate a specific order or sequence, but are intended to distinguish similar subjects. It should be understood that the data used in this manner are interchangeable where appropriate so that the embodiments of this application described herein may be carried out in an order other than that illustrated or described herein. Furthermore, the terms "includes," "has," and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to including those steps or units that are explicitly shown, but may include steps or units that are not explicitly shown, or other steps or units specific to those processes, methods, products, or devices.

[0019] Figure 1 is a schematic flowchart of an audio noise reduction method provided in an embodiment of the present invention, which can be applied to noise reduction of speech and can be applied to various scenarios such as voice calls, audio / video live broadcasts, and multi-person conferences. The method can be carried out by an audio noise reduction device, which is implemented in the form of hardware and / or software, and which can be placed in an electronic device such as an audio noise reduction device. The electronic device may be a mobile device such as a mobile phone, smartwatch, tablet, and personal digital assistant, or it may be another device such as a desktop computer. As shown in Figure 1, the method includes the following steps.

[0020] Step 101: A pre-configured audio activity detection algorithm is used to perform detection on the current audio frame awaiting processing, and the corresponding algorithm activity detection result is obtained.

[0021] For example, a currently awaiting processing audio frame can be understood as an audio frame that currently requires audio noise reduction processing, and the currently awaiting audio frame can be contained within an audio file or audio stream. As an option, the currently awaiting audio frame may be the original audio frame in the audio file or audio stream, or it may be an audio frame obtained by preprocessing the original audio frame.

[0022] In the embodiments of this application, the entire speech noise reduction scheme can be understood as a single speech noise reduction system, and the current audio frame can be understood as the input signal to the speech noise reduction system. The speech noise reduction scheme may include a traditional speech noise reduction algorithm and an AI speech noise reduction model.

[0023] Here, the type of traditional speech noise reduction algorithm may be, for example, an Adaptive Noise Suppression (ANS) algorithm in network real-time communication (Web Real-Time Communication, webRTC), a linear filtering method, spectral subtraction, a statistical model algorithm, or a subspace algorithm. Traditional speech noise reduction algorithms mainly consist of three main parts: voice activity detection (VAD) estimation, noise estimation, and noise reduction. Voice activity detection, also called voice endpoint detection or voice boundary detection, can identify long periods of silence from a voice signal stream. The pre-configured voice activity detection algorithm in the embodiment of this application may be any voice activity detection algorithm in a traditional speech noise reduction algorithm.

[0024] Here, the pre-configured speech noise reduction network model in this application may be an AI speech noise reduction model, and may include, for example, an RNNoise model or a Dual-Signal Transformation LSTM Network for Real-Time Noise Suppression (DTLN) noise reduction model. The pre-configured speech noise reduction network model includes two branches, one of which outputs noise-reduced speech (which may be abbreviated as the noise reduction branch), and the other branch of which outputs speech activity detection results (which may be abbreviated as the detection branch). For AI speech noise reduction models that already include a detection branch, the original model structure can be maintained, and for AI speech noise reduction models that do not include a detection branch, a detection branch can be added based on the core network, and the network structure of the detection branch may include, for example, convolutional layers and / or fully connected layers.

[0025] Here, RNNoise is a noise reduction strategy that employs a combination of audio feature extraction and a deep neural network.

[0026] For example, to make it easier to distinguish between audio activity detection results from different sources, a pre-configured audio activity detection algorithm can be used to perform detection on the currently pending audio frame. The obtained detection result can then be recorded as the algorithm activity detection result, and the activity detection result output by a pre-configured audio noise reduction network model can be recorded as the model activity detection result.

[0027] Step 102: The model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame are integrated to obtain the target activity detection result corresponding to the current audio frame. Here, the model activity detection result is output from a pre-configured audio noise reduction network model.

[0028] For example, the previous audio frame can be understood as the most recent audio frame preceding the current audio frame; that is, the previous audio frame is located before the current audio frame, and the previous audio frame and the current audio frame are adjacent in terms of their frame sequence numbers. When performing audio noise reduction processing on the previous audio frame, a pre-configured audio noise reduction network model can output a noise reduction audio frame and model activity detection results corresponding to the previous audio frame, and cache the model activity detection results for noise reduction processing on the current audio frame.

[0029] In the embodiment of the present invention, when processing the current audio frame, the activity detection result (target activity detection result) used for noise estimation and noise reduction in the traditional speech noise reduction algorithm can be determined by combining the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame. Compared to simply employing the traditional speech noise reduction algorithm to perform speech activity detection, the traditional noise reduction algorithm can obtain more VAD information, thereby obtaining more accurate noise estimation, better protecting speech, and more accurately removing noise, and improving the signal-to-noise ratio (SNR) of the traditional noise reduction algorithm.

[0030] Step 103: Based on the target activity detection result, noise estimation and noise reduction are performed on the current audio frame to obtain an initial noise-reduced audio frame.

[0031] For example, after obtaining the target activity detection result, the noise estimation algorithm and noise reduction algorithm in traditional speech noise reduction algorithms can be used to perform the corresponding processing on the current audio frame, and the resulting audio frame can be recorded as the initial noise-reduced audio frame.

[0032] Step 104: The initial noise reduction audio frame is input to the pre-configured audio noise reduction network model, which outputs the model activity detection results corresponding to the target noise reduction audio frame and the current audio frame.

[0033] For example, after obtaining an initial noise-reduced audio frame, the initial noise-reduced audio frame can be directly fed into a pre-configured speech noise reduction network model, or the initial noise-reduced audio frame can be transformed based on the features of the pre-configured speech noise reduction network model, for example, into a signal of a pre-configured dimension, which may be, for example, the frequency domain, the time domain, or other dimensions.

[0034] The audio noise reduction method provided in the embodiment of the present invention employs a pre-configured audio activity detection algorithm to perform detection on the current audio frame awaiting processing, obtains the corresponding algorithm activity detection result, performs integrated processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame, obtains the target activity detection result corresponding to the current audio frame, the model activity detection result is output from a pre-configured audio noise reduction network model, noise estimation and noise reduction are performed on the current audio frame based on the target activity detection result to obtain an initial noise reduction audio frame, and the initial noise reduction audio frame is input to a pre-configured audio noise reduction network model to output the target noise reduction audio frame and the model activity detection result corresponding to the current audio frame. By adopting the above approach, the pre-configured speech noise reduction network model can output model activity detection results. When processing the current audio frame using a traditional speech noise reduction algorithm, the model activity detection result from the previous audio frame can be combined with the algorithm activity detection result from the traditional speech noise reduction algorithm. This allows the traditional noise reduction algorithm to obtain more activity detection information, enabling a more rational and accurate determination of speech activity detection results. By performing noise estimation and noise reduction based on these results, speech is better protected, more noise is removed, and a higher signal-to-noise ratio can be obtained from traditional noise reduction. Furthermore, by using the traditional noise reduction results as input to the pre-configured speech noise reduction network model, a more effective noise-reduced audio frame is obtained, reducing the possibility of the pre-configured speech noise reduction network model processing poor data. The traditional noise reduction algorithm and the AI ​​noise reduction method mutually enhance each other, resulting in superior noise reduction capabilities against various types of noise and increasing the overall stability and robustness of the approach.

[0035] In the embodiments of the present invention, the audio activity detection may be at the frame level or at the frequency point level, and the detection result may be represented by one or more probability values.

[0036] In some embodiments, the algorithm activity detection result includes a first probability value that audio is present in the corresponding audio frame, and the model activity detection result includes a second probability value that audio is present in the corresponding audio frame. Here, the integration process performed on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain the target activity detection result corresponding to the current audio frame employs a pre-set calculation method and the second probability value in the model activity detection result corresponding to the previous audio frame. 2 The probability value and the number in the algorithm activity detection result corresponding to the current audio frame. 1 This includes calculating a probability value, obtaining a third probability value, and determining the target activity detection result corresponding to the current audio frame based on the third probability value. With this setup, the target activity detection result can be accurately determined for frame-level audio activity detection.

[0037] Here, the first probability value is used to represent the probability that a corresponding audio frame contains audio after detection has been performed on the corresponding audio frame using a pre-configured audio activity detection algorithm. The corresponding audio frame here may be any audio frame, the current audio frame, or the previous audio frame, and the first probability value corresponding to a different audio frame may be different. The second probability value is used to represent the probability that a corresponding audio frame output by a pre-configured audio noise reduction network model contains audio. The corresponding audio frame here may be any audio frame, and the second probability value corresponding to a different audio frame may be different.

[0038] For example, the first probability value in the algorithm activity detection result corresponding to the current audio frame can be used to represent the probability that the current audio frame (let's call it A) contains audio after performing detection on the current audio frame using a pre-configured audio activity detection algorithm, and can be recorded as Pa. The second probability value in the model activity detection result corresponding to the previous audio frame can be used to represent the probability that the previous audio frame (let's call it B) contains audio, as predicted by a pre-configured audio noise reduction network model, when performing audio noise reduction processing on the previous audio frame, and can be recorded as Pb. Pa and Pb can be calculated using a pre-configured calculation method to obtain a third probability value, which can be recorded as Pc. For example, the third probability value can be the target activity detection result corresponding to the current audio frame.

[0039] For example, the pre-set calculation method is one of the following: obtaining the maximum value, obtaining the minimum value, calculating the average value, addition, calculating a weighted sum, or calculating a weighted average value. Taking obtaining the maximum value as an example, Pc = max(Pa, Pb).

[0040] In some embodiments, the algorithm activity detection result includes a fourth probability value indicating that sound is present at each of a preset number of frequency points in the corresponding audio frame, and the model activity detection result includes a fifth probability value indicating that sound is present at each of the preset number of frequency points in the corresponding audio frame. Here, integrating the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain a target activity detection result corresponding to the current audio frame involves, for each of the preset number of frequency points, employing a preset calculation method to calculate the fifth probability value for a single frequency point in the model activity detection result corresponding to the previous audio frame and the fourth probability value for the corresponding single frequency point in the algorithm activity detection result corresponding to the current audio frame, thereby obtaining a sixth probability value, and determining the target activity detection result corresponding to the current audio frame based on the preset number of sixth probability values. With this configuration, frequency point-level audio activity detection can be employed to more accurately determine the target activity detection result.

[0041] For example, a pre-set number (recorded as n) can be set according to the actual needs, for example, based on the number of points used in the Fast Fourier Transform during the preprocessing stage, for example, n is 256. The fourth probability value corresponding to the current audio frame is used to represent the probability that each frequency point among the pre-set number of frequency points in the current audio frame contains speech, after performing detection on the current audio frame (let's call it A) using a pre-set speech activity detection algorithm, and can be recorded as PA[n], which can be understood as a vector containing n elements (nth order), where the value of each element takes a value from 0 to 1, and the value of one element is used to represent the probability that the corresponding frequency point contains speech. The fifth probability value corresponding to the previous audio frame is used to represent the probability that each frequency point among the pre-set number of frequency points in the previous audio frame contains speech, as predicted by a pre-set speech noise reduction network model when performing speech noise reduction processing on the previous audio frame (let's call it B), and can be recorded as PB[n]. PA[n] and PB[n] are calculated using a pre-set calculation method, and a pre-set number of sixth probability values ​​are obtained and recorded, for example, as PC[n]. Exemplarily, a vector containing the sixth probability values ​​can be used as the target activity detection result corresponding to the current audio frame.

[0042] For example, the pre-set calculation method is one of the following: obtaining the maximum value, obtaining the minimum value, calculating the average value, addition, calculating a weighted sum, or calculating a weighted average value. Taking obtaining the maximum value as an example, PC[n] = max(PA[n], PB[n]). For example, for the first frequency point in the current audio frame, the maximum value among the corresponding fourth and fifth probability values ​​becomes the sixth probability value corresponding to the first frequency point in the current audio frame, and the same applies to subsequent frequency points.

[0043] In some embodiments, inputting the initial noise-reduced audio frame into the pre-configured speech noise reduction network model includes performing feature extraction of a pre-configured feature dimension on the initial noise-reduced audio frame to obtain a target input signal, and inputting the target input signal into the pre-configured speech noise reduction network model, or inputting the target input signal and the initial noise-reduced audio frame into the pre-configured speech noise reduction network model. With this configuration, purposeful feature extraction can be performed, improving the predictive accuracy and precision of the pre-configured speech noise reduction network model.

[0044] As one option, the pre-defined feature dimension may include the dominant feature dimension and may be a fundamental frequency feature, for example, the fundamental frequency (Pitch), a per-channel energy normalization (PCEN) feature, or a Mel-frequency cepstrum coefficient (MFCC) feature. The pre-defined feature dimension can be determined based on the network structure or feature points of a pre-defined speech noise reduction network model.

[0045] Figure 2 is a schematic flowchart of another speech noise reduction method provided in the embodiments of the present application, which is optimized based on each of the selectable embodiments described above, and Figure 3 is a schematic flowchart of the reasoning flow of the speech noise reduction method provided in the embodiments of the present application, which can be understood in conjunction with Figures 2 and 3. Herein, as shown in Figure 2, the method may include the following steps:

[0046] Step 201: Obtain the original audio frame, perform preprocessing on the original audio frame, and obtain the current audio frame awaiting processing.

[0047] Exemplary, the original audio frame is contained within an audio file or audio stream. For example, it may be an audio stream in a voice call scene. To ensure call quality, noise reduction must be applied to the call audio. Preprocessing may include, for example, framing, windowing, and Fourier transform. The noisy audio frame after preprocessing is the current audio frame awaiting processing and is the input signal (recorded as S0) for a pre-configured traditional noise reduction algorithm.

[0048] Step 202: A pre-configured audio activity detection algorithm from among the pre-configured traditional noise reduction algorithms is used to perform detection on the current audio frame awaiting processing, and the corresponding algorithm activity detection result is obtained.

[0049] For example, the pre-configured traditional noise reduction algorithm may be an ANS algorithm. By using a pre-configured speech activity detection algorithm corresponding to the VAD estimation function module of the ANS algorithm, and assuming frequency point level detection, it is possible to obtain the speech presence probability Pf

[0256] for 256 frequency points, that is, the algorithm activity detection result corresponding to S0 can be obtained.

[0050] Step 203: For the current audio frame, determine whether a previous audio frame exists. If it exists, perform step 204; otherwise, perform step 206.

[0051] For example, for the first audio frame, since there is no previous audio frame, it is not necessary to obtain the model activity detection result for the previous audio frame. Step 206 can then be executed, and noise estimation and denoising can be performed based on the algorithm activity detection result corresponding to the current audio frame.

[0052] Step 204: Obtain the model activity detection result corresponding to the previous audio frame. Perform an integrated process on the obtained model activity detection result and the algorithm activity detection result corresponding to the current audio frame to obtain the target activity detection result corresponding to the current audio frame.

[0053] For example, the model activity detection result corresponding to the previous audio frame is output from a pre-configured speech noise reduction network model based on artificial intelligence, and may be the speech presence probability PF

[0256] for 256 frequency points in the previous audio frame. By adopting a method that takes the maximum value, an integrated VAD estimation result (target activity detection result): P

[0256] = max(Pf

[0256] , PF

[0256] ) can be obtained.

[0054] In step 205, based on the target activity detection result, noise estimation and noise reduction are performed on the current audio frame using the pre-configured traditional noise reduction algorithm to obtain an initial noise-reduced audio frame, and then step 207 is executed.

[0055] For example, a pre-configured traditional noise reduction algorithm performs noise estimation and noise reduction based on P

[0256] to obtain an audio signal S1 that has undergone traditional noise reduction processing, i.e., an initial noise-reduced audio frame.

[0056] Step 206: Based on the algorithm activity detection result corresponding to the current audio frame, noise estimation and noise reduction are performed on the current audio frame using the pre-configured traditional noise reduction algorithm to obtain an initial noise-reduced audio frame.

[0057] For example, a pre-configured traditional noise reduction algorithm performs noise estimation and noise reduction based on Pf

[0256] to obtain an audio signal S1 that has undergone traditional noise reduction processing, i.e., an initial noise-reduced audio frame.

[0058] Step 207: Feature extraction is performed on the initial noise-reduced audio using a pre-defined feature dimension to obtain the target input signal.

[0059] For example, S1 may be a signal in the frequency domain, time domain, or other dimensional domain as the input signal to a pre-configured speech noise reduction network model, and based on differences in the model design of the pre-configured speech noise reduction network model, there may be a one-step extraction and calculation of dominant features (e.g., fundamental frequency features), and the extracted feature information is recorded as the target input signal S2.

[0060] Step 208: The target input signal and / or initial noise reduction audio frame are input to a pre-configured audio noise reduction network model, which outputs model activity detection results corresponding to the target noise reduction audio frame and the current audio frame.

[0061] As one option, either S1 or S2 may be used as the input to the model, or both S1 and S2 may be used as the input to the model. These inputs are then fed into a pre-configured speech noise reduction network model to perform inference calculations and obtain an output signal. The output signal consists of two parts: the first part is the output S3 of the final noise-reduced speech from the speech noise reduction method, and the second part is the VAD output PF

[0256] of the model, which is used when the traditional speech noise reduction algorithm processes the next audio frame.

[0062] Step 209 determines whether there are any original audio frames waiting to be processed. If there are, the process returns to step 201; otherwise, the process ends.

[0063] For example, when a voice call ends, if all original audio frames have already undergone noise reduction processing, the process can be terminated at this point. If there are original audio frames that have not been noise-reduced, the process can be returned to step 201 and the noise reduction processing can be continued.

[0064] The speech noise reduction method provided in the embodiment of this application utilizes a method in which a pre-configured speech noise reduction network model based on artificial intelligence provides information feedback to a traditional noise reduction algorithm, enabling the traditional noise reduction algorithm to obtain more VAD information. Both the traditional noise reduction and AI noise reduction VAD estimation employ frequency point levels, resulting in more accurate noise estimation. The traditional noise reduction algorithm can better protect speech, remove more noise, and improve the output signal-to-noise ratio of the traditional noise reduction. After the initial noise-reduced speech signal with a high signal-to-noise ratio undergoes feature extraction, it enriches the input to the pre-configured speech noise reduction network model, reducing the possibility of the pre-configured speech noise reduction network model processing poor data, improving the speech noise reduction effect of the model, and enhancing the speech noise reduction performance.

[0065] Figure 4 is a schematic flowchart of the model training method provided in the embodiment of the present application, and Figure 5 is a schematic diagram of the training process of the model training method provided in the embodiment of the present application, and can be understood in conjunction with Figure 4 to represent the embodiment of the present application. This embodiment can be applied when training an artificial intelligence-based speech noise reduction network model, which can be applied to various scenes such as voice calls, audio / video live broadcasts, and multi-person conferences. The method can be performed by a model training device, which can be implemented in the form of hardware and / or software, and which can be placed on an electronic device such as a model training device. The electronic device may be a mobile device such as a mobile phone, smartwatch, tablet, and personal digital assistant, or it may be another device such as a desktop computer. A speech noise reduction network model trained using the embodiment of the present application can be applied to a speech noise reduction method provided in any embodiment of the present application.

[0066] As shown in Figure 4, the method includes the following steps.

[0067] Step 401: A pre-configured audio activity detection algorithm is used to perform detection on the current sample audio frame and obtain the corresponding sample algorithm activity detection result. Here, the current sample audio frame is associated with an activity detection label and a pure audio frame.

[0068] For example, a pure (clean) audio dataset and a noise dataset can be mixed with a noised audio dataset according to a predefined mixing rule, which can be based, for example, on the signal-to-noise ratio or the Room Impulse Response (RIR). As an option, the mixed noised audio dataset and the pure audio dataset can be used together as the training set for the model. The current sample audio frame may be an audio frame in the training set. The current sample audio frame may have an activity detection label, which can be added by a manual markup method. For example, at the frame level, the label may be 1 if it contains audio and 0 if it does not. For example, at the frequency point level, the label may be a vector containing a predefined number of elements, each element taking a value of 1 or 0, where the corresponding frequency point has a value of 1 if it contains audio and a value of 0 if it does not.

[0069] Step 402: The sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame are integrated to obtain the target sample activity detection result corresponding to the current sample audio frame. Here, the sample model activity detection result is output from the speech noise reduction network model.

[0070] For example, the integration process of activity detection results in this step can be similar to the integration process in the speech noise reduction method provided in the embodiments of the present application, and may be, for example, integration at the frequency point level or integration at the frame level, and a similar preset calculation method can be used to perform integration for the corresponding frequency values, and specific details can be found in the relevant parts of the present application, and a detailed explanation is omitted here.

[0071] Step 403, the aforementioned objective Sample Activity Based on the detection results, noise estimation and noise reduction are performed on the current sample audio frame to obtain an initial noise-reduced sample audio frame.

[0072] Step 404: The initial noise reduction sample audio frame is input to the audio noise reduction network model to output the sample model activity detection results corresponding to the target noise reduction sample audio frame and the current sample audio frame.

[0073] Step 405: Determine a first loss relationship based on the target sample noise reduction audio frame and the pure audio frame; determine a second loss relationship based on the sample model activity detection result and the activity detection label; and train the speech noise reduction network model based on the first loss relationship and the second loss relationship.

[0074] Exemplary, loss relationships can be used to characterize the difference between two types of data, can be expressed in terms of loss values, and can be calculated, for example, by employing a loss function. A first loss relationship is used to characterize the difference between a target sample noise-reduced audio frame and a pure audio frame, and a second loss relationship is used to characterize the difference between a sample model activity detection result and an activity detection label. Here, the function types of the first loss function for calculating the first loss relationship and the second loss function for calculating the second loss relationship can be set according to the actual needs.

[0075] For example, a target loss relationship can be calculated based on the first loss relationship and the second loss relationship, and the calculation method may be, for example, a weighted sum.

[0076] For example, a speech noise reduction network model can be trained based on a target loss relationship. During the training process, the goal is to minimize the target loss relationship, and the weight parameter values ​​in the speech noise reduction network model can be optimized using training methods such as backpropagation until a pre-set training stop condition is met. The training stop condition can be set according to the actual needs, for example, based on the iteration order, the degree of convergence of the loss value, or the model accuracy.

[0077] The model training method provided in the embodiment of the present invention avoids the risk of data mismatch that can occur when traditional noise reduction algorithms and speech noise reduction network models are trained as a single unit during the training process, by serializing the traditional noise reduction algorithm with a speech noise reduction network model that is trained independently. The model obtained after training is used for speech noise reduction and has excellent noise reduction capabilities against various types of noise, improving the noise reduction effect.

[0078] As one option, the sample algorithm activity detection result includes a first sample probability value indicating that audio is present in the corresponding sample audio frame, and the sample model activity detection result includes a second sample probability value indicating that audio is present in the corresponding sample audio frame. Here, integrating the sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame to obtain a target sample activity detection result corresponding to the current sample audio frame includes using a pre-set calculation method to calculate a second sample probability value in the sample model activity detection result corresponding to the previous sample audio frame and a first sample probability value in the sample algorithm activity detection result corresponding to the current sample audio frame, obtaining a third sample probability value, and determining the target sample activity detection result corresponding to the current sample audio frame based on the third sample probability value.

[0079] As one option, the sample algorithm activity detection result includes a fourth sample probability value in the corresponding audio frame where sound is present at each of the preset number of frequency points, and the model activity detection result includes a fifth sample probability value in the corresponding audio frame where sound is present at each of the preset number of frequency points. Here, integrating the sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame to obtain a target sample activity detection result corresponding to the current sample audio frame includes, for each frequency point among the predetermined number of frequency points, employing a predetermined calculation method to calculate the 5th sample probability value of a single frequency point in the sample model activity detection result corresponding to the previous sample audio frame and the 4th sample probability value of the corresponding single frequency point in the sample algorithm activity detection result corresponding to the current sample audio frame, thereby obtaining a 6th sample probability value, and determining the target sample activity detection result corresponding to the current sample audio frame based on the predetermined number of 6th sample probability values.

[0080] As one option, inputting the initial noise reduction sample audio frame into the speech noise reduction network model includes performing feature extraction of a preset feature dimension on the initial noise reduction sample audio frame to obtain a target input signal, and inputting the target input signal into the speech noise reduction network model, or inputting the target input signal and the initial noise reduction sample audio frame into the speech noise reduction network model.

[0081] Figure 6 is a structural block diagram of an audio noise reduction device provided in an embodiment of the present invention, which can be implemented by software and / or hardware and can generally be integrated into an electronic device such as an audio noise reduction device, and can perform audio noise reduction by executing an audio noise reduction method. As shown in Figure 6, the device includes an audio activity detection module 601, a detection result integration module 602, a noise reduction processing module 603, and a model input module 604. The audio activity detection module 601 is configured to employ a pre-configured audio activity detection algorithm to perform detection on the currently pending audio frame and obtain the corresponding algorithm activity detection result. The detection result integration module 602 is configured to perform integration processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain the target activity detection result corresponding to the current audio frame, and the model activity detection result is output from a pre-configured audio noise reduction network model. The noise reduction processing module 603 is configured to perform noise estimation and noise reduction on the current audio frame based on the target activity detection result, in order to obtain an initial noise-reduced audio frame. The model input module 604 is configured to input the initial noise reduction audio frame to the pre-configured audio noise reduction network model, thereby outputting model activity detection results corresponding to the target noise reduction audio frame and the current audio frame.

[0082] The audio noise reduction device provided in the embodiment of the present invention employs a pre-configured audio activity detection algorithm to perform detection on the current audio frame awaiting processing, obtains the corresponding algorithm activity detection result, performs integrated processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame, obtains the target activity detection result corresponding to the current audio frame, the model activity detection result is output from a pre-configured audio noise reduction network model, noise estimation and noise reduction are performed on the current audio frame based on the target activity detection result, obtains an initial noise reduction audio frame, and inputs the initial noise reduction audio frame to a pre-configured audio noise reduction network model to output the target noise reduction audio frame and the model activity detection result corresponding to the current audio frame. By adopting the above approach, the pre-configured speech noise reduction network model can output model activity detection results. When processing the current audio frame using a traditional speech noise reduction algorithm, the model activity detection result from the previous audio frame can be combined with the algorithm activity detection result from the traditional speech noise reduction algorithm. This allows the traditional noise reduction algorithm to obtain more activity detection information, enabling a more rational and accurate determination of speech activity detection results. By performing noise estimation and noise reduction based on these results, speech is better protected, more noise is removed, and a higher signal-to-noise ratio can be obtained from traditional noise reduction. Furthermore, by using the traditional noise reduction results as input to the pre-configured speech noise reduction network model, a more effective noise-reduced audio frame is obtained, reducing the possibility of the pre-configured speech noise reduction network model processing poor data. The traditional noise reduction algorithm and the AI ​​noise reduction method mutually enhance each other, resulting in superior noise reduction capabilities against various types of noise and increasing the overall stability and robustness of the approach.

[0083] As one option, the algorithm activity detection result includes a first probability value indicating that audio is present in the corresponding audio frame, and the model activity detection result includes a second probability value indicating that audio is present in the corresponding audio frame. Here, the detection result integration module 602 is configured to perform integration processing on the model activity detection result and the algorithm activity detection result in the following manner to obtain the target activity detection result corresponding to the current audio frame. A pre-set calculation method is used to calculate a second probability value in the model activity detection result corresponding to the previous audio frame and a first probability value in the algorithm activity detection result corresponding to the current audio frame, thereby obtaining a third probability value, and the target activity detection result corresponding to the current audio frame is determined based on the third probability value.

[0084] As one option, the algorithm activity detection result includes a fourth probability value indicating that sound is present at each of the preset number of frequency points in the corresponding audio frame, and the model activity detection result includes a fifth probability value indicating that sound is present at each of the preset number of frequency points in the corresponding audio frame. Here, the detection result integration module 602 is further configured to perform integration processing on the model activity detection result and the algorithm activity detection result in the following manner to obtain the target activity detection result corresponding to the current audio frame. For each of the aforementioned number of pre-set frequency points, a pre-set calculation method is used to calculate the fifth probability value of a single frequency point in the model activity detection result corresponding to the previous audio frame and the fourth probability value of the corresponding single frequency point in the algorithm activity detection result corresponding to the current audio frame, thereby obtaining a sixth probability value. Based on the aforementioned number of pre-set sixth probability values, the target activity detection result corresponding to the current audio frame is determined.

[0085] One of the options is that the pre-set calculation method is one of the following: obtaining the maximum value, obtaining the minimum value, calculating the average value, addition, calculating a weighted sum, or calculating a weighted average value.

[0086] As one option, the model input module includes a feature extraction unit and a signal input unit. The feature extraction unit is installed to perform feature extraction of a preset feature dimension from the initial noise-reduced audio and obtain the target input signal. The signal input unit is configured to input the target input signal to the preset audio noise reduction network model, or to input the target input signal and the initial noise reduction audio frame to the preset audio noise reduction network model, thereby outputting model activity detection results corresponding to the target noise reduction audio frame and the current audio frame.

[0087] Figure 7 is a structural block diagram of a model training device provided in an embodiment of the present invention, which can be implemented by software and / or hardware, and can generally be integrated into an electronic device such as a model training device, and can perform model training by executing a model training method. As shown in Figure 7, the device includes a voice detection module 701, an integration module 702, a noise reduction module 703, a network model input module 704, and a network model training module 705. The voice detection module 701 employs a pre-configured voice activity detection algorithm to perform detection on the currently pending sample audio frame and obtain the corresponding sample algorithm activity detection result. The current sample audio frame is associated with an activity detection label and a clean audio frame. The integration module 702 is configured to perform integration processing on the sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame to obtain a target sample activity detection result corresponding to the current sample audio frame, and the sample model activity detection result is output from the audio noise reduction network model. The noise reduction module 703 targets the aforementioned objective Sample Activity Based on the detection results, the system is configured to perform noise estimation and noise reduction on the current sample audio frame to obtain an initial noise-reduced sample audio frame. The network model input module 704 is configured to input the initial noise reduction sample audio frame to the audio noise reduction network model and output sample model activity detection results corresponding to the target sample noise reduction audio frame and the current sample audio frame. The network model training module 705 is configured to determine a first loss relationship based on the target sample noise reduction audio frame and the clean audio frame, determine a second loss relationship based on the sample model activity detection result and the activity detection label, and train the speech noise reduction network model based on the first loss relationship and the second loss relationship.

[0088] The model training apparatus provided in the embodiment of the present invention avoids the risk of data mismatch that can occur when traditional noise reduction algorithms and speech noise reduction network models are trained as a single unit during the training process, by serializing the traditional noise reduction algorithm with a speech noise reduction network model that is trained independently. The model obtained after training is used for speech noise reduction and has excellent noise reduction capabilities against various types of noise, improving the noise reduction effect.

[0089] Embodiments of the present application provide an electronic device, which can integrate a speech noise reduction device and / or a model training device provided in the embodiments of the present application. Figure 8 is a structural block diagram of an electronic device provided in the embodiments of the present application. The electronic device 800 includes a processor 801 and a memory 802 that communicates with the processor 801. The memory 802 stores a computer program that can be executed by the processor 801, and when the computer program is executed by the processor 801, the processor 801 can execute the speech noise reduction method and / or model training method described in any embodiment of the present application. Here, the number of processors may be one or more, and in Figure 8, one processor is used as an example.

[0090] Embodiments of the present application further provide a computer-readable storage medium in which a computer program is stored, and which, when executed by a processor, is used to implement the speech noise reduction method and / or model training method described in any embodiment of the present application.

[0091] Embodiments of the present application further provide a computer program product which includes a computer program that, when executed by a processor, implements, for example, a speech noise reduction method and / or a model training method provided in embodiments of the present application.

[0092] The voice noise reduction apparatus, model training apparatus, electronic device, storage medium, and product provided in the above embodiments can perform the voice noise reduction method or model training method provided in the corresponding embodiments of the present application and comprises a corresponding functional module and beneficial effects for performing such method. Technical details not described in detail in the above embodiments can be referenced from the voice noise reduction method or model training method provided in any embodiment of the present application.

Claims

1. This involves using a pre-configured audio activity detection algorithm to perform detection on the currently pending audio frame and obtaining the corresponding algorithm activity detection result. The model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame are integrated to obtain the target activity detection result corresponding to the current audio frame, and the model activity detection result corresponding to the previous audio frame is output from the pre-configured audio noise reduction network model. Based on the target activity detection result corresponding to the current audio frame, noise estimation and noise reduction are performed on the current audio frame to obtain an initial noise-reduced audio frame. By inputting the initial noise reduction audio frame into the pre-configured audio noise reduction network model, the model activity detection results corresponding to the target noise reduction audio frame and the current audio frame are output, and the model activity detection results corresponding to the current audio frame are used for integration processing corresponding to the next audio frame. including, Methods for reducing audio noise.

2. The algorithm activity detection result includes a first probability value indicating that audio is present in the corresponding audio frame, and the model activity detection result includes a second probability value indicating that audio is present in the corresponding audio frame. Performing an integrated process on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain the target activity detection result corresponding to the current audio frame is: This includes using a pre-set calculation method to calculate a second probability value in the model activity detection result corresponding to the previous audio frame and a first probability value in the algorithm activity detection result corresponding to the current audio frame, obtaining a third probability value, and determining the target activity detection result corresponding to the current audio frame based on the third probability value. The method according to claim 1.

3. The algorithm activity detection result includes a fourth probability value indicating that sound is present at each of the predetermined number of frequency points in the corresponding audio frame, and the model activity detection result includes a fifth probability value indicating that sound is present at each of the predetermined number of frequency points in the corresponding audio frame. Performing an integrated process on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain the target activity detection result corresponding to the current audio frame is: For each of the aforementioned number of pre-set frequency points, a pre-set calculation method is used to calculate the fifth probability value of a single frequency point in the model activity detection result corresponding to the previous audio frame, and the fourth probability value of the corresponding single frequency point in the algorithm activity detection result corresponding to the current audio frame, thereby obtaining a sixth probability value. Based on the aforementioned number of predetermined sixth probability values, the target activity detection result corresponding to the current audio frame is determined, including, The method according to claim 1.

4. The aforementioned pre-set calculation method is one of the following: obtaining the maximum value, obtaining the minimum value, calculating the average value, addition, calculating a weighted sum, or calculating a weighted average value. The method according to claim 2.

5. Inputting the initial noise-reduced audio frame into the pre-configured audio noise reduction network model outputs the model activity detection results corresponding to the target noise-reduced audio frame and the current audio frame. The process involves performing feature extraction on the initial noise-reduced audio frame with a predetermined feature dimension to obtain the target input signal. The target input signal is input to the preset audio noise reduction network model, or the target input signal and the initial noise reduction audio frame are input to the preset audio noise reduction network model, thereby outputting model activity detection results corresponding to the target noise reduction audio frame and the current audio frame. including, The method according to claim 1.

6. A model training method performed on at least one processor, A pre-configured audio activity detection algorithm is used to perform detection on the current sample audio frame, and the corresponding sample algorithm activity detection result is obtained, confirming that the current sample audio frame is associated with an activity detection label and a pure audio frame. The sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame are integrated to obtain the target sample activity detection result corresponding to the current sample audio frame, and the sample model activity detection result corresponding to the previous sample audio frame is output from the audio noise reduction network model. Based on the target sample activity detection result corresponding to the current sample audio frame, noise estimation and noise reduction are performed on the current sample audio frame to obtain an initial noise-reduced sample audio frame. By inputting the initial noise reduction sample audio frame into the audio noise reduction network model, the system outputs the sample model activity detection results corresponding to the target sample noise reduction audio frame and the current sample audio frame. The sample model activity detection results corresponding to the current sample audio frame are used for the integration process corresponding to the next sample audio frame. Based on the target sample noise reduction audio frame and the pure audio frame, a first loss relationship is determined; based on the sample model activity detection result and the activity detection label, a second loss relationship is determined; and based on the first loss relationship and the second loss relationship, the speech noise reduction network model is trained. including, Model training methods.

7. It includes an audio activity detection module, a detection result integration module, a noise reduction processing module, and a model input module. The aforementioned audio activity detection module is configured to employ a pre-configured audio activity detection algorithm to perform detection on the currently pending audio frame and obtain the corresponding algorithm activity detection result. The aforementioned detection result integration module is configured to perform integration processing on the model activity detection result corresponding to the previous audio frame and the algorithm activity detection result corresponding to the current audio frame to obtain a target activity detection result corresponding to the current audio frame, wherein the model activity detection result corresponding to the previous audio frame is output from a pre-configured audio noise reduction network model. The noise reduction processing module is configured to perform noise estimation and noise reduction on the current audio frame based on the target activity detection result corresponding to the current audio frame, in order to obtain an initial noise-reduced audio frame. The model input module inputs the initial noise reduction audio frame to the pre-configured audio noise reduction network model, outputting model activity detection results corresponding to the target noise reduction audio frame and the current audio frame. The model activity detection results corresponding to the current audio frame are configured to be used for integration processing corresponding to the next audio frame. Audio noise reduction device.

8. It includes a voice detection module, an integration module, a noise reduction module, a network model input module, and a network model training module. The aforementioned audio detection module is configured to perform detection on the currently pending sample audio frame using a pre-configured audio activity detection algorithm and to obtain the corresponding sample algorithm activity detection result. The current sample audio frame is associated with an activity detection label and a clean audio frame. The integration module is configured to perform integration processing on the sample model activity detection result corresponding to the previous sample audio frame and the sample algorithm activity detection result corresponding to the current sample audio frame to obtain the target sample activity detection result corresponding to the current sample audio frame, and the sample model activity detection result corresponding to the previous sample audio frame is output from the audio noise reduction network model. The noise reduction module is configured to perform noise estimation and noise reduction on the current sample audio frame based on the target sample activity detection result corresponding to the current sample audio frame, in order to obtain an initial noise-reduced sample audio frame. The network model input module inputs the initial noise reduction sample audio frame to the audio noise reduction network model, outputting sample model activity detection results corresponding to the target sample noise reduction audio frame and the current sample audio frame. The sample model activity detection results corresponding to the current sample audio frame are configured to be used for integration processing corresponding to the next sample audio frame. The network model training module is configured to determine a first loss relationship based on the target sample noise reduction audio frame and the clean audio frame, determine a second loss relationship based on the sample model activity detection result and the activity detection label, and train the audio noise reduction network model based on the first and second loss relationships. Model training device.

9. At least one processor, A memory that is communicated with at least one processor, Includes, The memory stores a computer program that can be executed by the at least one processor, and the execution of the computer program by the at least one processor enables the at least one processor to perform the speech noise reduction method and / or the model training method described in any one of claims 1 to 5. Electronic devices.

10. A computer program is stored, and the computer program is used to implement the speech noise reduction method and / or the model training method described in any one of claims 1 to 5 when executed by a processor. A computer-readable storage medium.