An audio noise reduction method and related apparatus
By combining adaptive beamforming and audio noise reduction models, audio noise reduction is achieved using an ideal ratio mask and beam parameters, which solves the problem of poor noise reduction effect in existing technologies and achieves a significant improvement in signal quality and effective noise suppression.
Patent Information
- Application Number
- CN202411661147.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-11-20
AI Technical Summary
Existing audio noise reduction methods are not effective and cannot meet the needs of some application scenarios.
An adaptive beamforming and audio denoising model is combined to perform beamforming by determining the ideal ratio mask and beam parameters of the target audio frame, and then denoising is performed using a pre-trained audio denoising model.
It significantly improves the quality of audio signals, effectively suppresses noise, and achieves better noise reduction.
Smart Images

Figure CN119446174B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio noise reduction technology, and in particular to an audio noise reduction method and related device. BACKGROUND
[0002] In many scenarios, a microphone array is used to collect audio signals. Due to various factors, the audio signals collected by the microphone array usually contain noise. In order to improve the quality of the audio, it is usually necessary to perform noise reduction processing on the audio to suppress the noise in the audio, so as to highlight the target speech in the audio.
[0003] The current audio noise reduction method is mostly a traditional noise reduction method (such as using a filtering algorithm to reduce noise in the audio). Although the traditional noise reduction method can reduce noise in the audio to some extent, the noise reduction effect is not good, and cannot meet the needs of some application scenarios. SUMMARY
[0004] Therefore, the present application provides an audio noise reduction method and related device to solve the problem of poor noise reduction effect of the current audio noise reduction method. The technical solutions are as follows:
[0005] The first aspect of the present application provides an audio noise reduction method, comprising:
[0006] determining an ideal ratio mask corresponding to a previous audio frame of the target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, wherein the target audio frame is an audio frame to be noise-reduced in a noisy multi-channel audio;
[0007] determining a beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame;
[0008] performing beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame to obtain an enhanced single-channel audio signal corresponding to the target audio frame;
[0009] performing noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame using a pre-trained audio noise reduction model to obtain a noise-reduced audio frame corresponding to the target audio frame.
[0010] In a possible implementation manner, the determining the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame comprises:
[0011] According to the frequency domain amplitude spectrum of the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame and the amplitude spectrum of the enhanced single-channel audio signal obtained by performing beamforming processing on the previous audio frame of the target audio frame, the ideal ratio mask corresponding to the previous audio frame of the target audio frame is determined.
[0012] In a possible implementation, the determining, according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, of the beam parameter corresponding to the target audio frame comprises: determining, according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, a speech presence probability of the target audio frame.
[0013] According to the ideal ratio mask corresponding to the previous audio frame of the target audio frame and the speech covariance matrix of the previous audio frame of the target audio frame, the speech covariance matrix of the target audio frame is determined.
[0014] In a possible implementation, the performing, according to the beam parameter corresponding to the target audio frame, of beamforming processing on the target audio frame to obtain the enhanced single-channel audio signal corresponding to the target audio frame comprises:
[0015] discriminating, for the target audio frame, a speech frame and a noise frame;
[0016] If the target audio frame is a speech frame, updating a relative transfer function in a beamforming algorithm according to the beam parameter corresponding to the target audio frame, and if the target audio frame is a noise frame, updating an adaptive noise canceller in the beamforming algorithm according to the beam parameter corresponding to the target audio frame, to obtain an updated beamforming algorithm.
[0017] performing, on the target audio frame, beamforming processing by using the updated beamforming algorithm to obtain the enhanced single-channel audio signal corresponding to the target audio frame.
[0018] In a possible implementation, the discriminating, for the target audio frame, of a speech frame and a noise frame comprises:
[0019] discriminating, for the target audio frame, a speech frame and a noise frame according to the ideal ratio mask corresponding to each audio frame in a forward adjacent frame sequence of the target audio frame, wherein the forward adjacent frame sequence of the target audio frame is a frame sequence composed of continuous N audio frames before the target audio frame, and N is an integer greater than 1.
[0020] In a possible implementation, the discriminating, for the target audio frame, of a speech frame and a noise frame according to the ideal ratio mask corresponding to each audio frame in a forward adjacent frame sequence of the target audio frame comprises:
[0021] If the ideal ratio mask corresponding to each audio frame in the sequence of forward adjacent frames of the target audio frame is less than or equal to a set threshold, the target audio frame is determined to be a noise frame.
[0022] If the ideal ratio mask corresponding to at least one audio frame in the sequence of forward adjacent frames of the target audio frame is greater than the set threshold, the target audio frame is determined to be a speech frame.
[0023] In a possible implementation, the enhanced single-channel audio signal corresponding to the target audio frame is a frequency domain signal.
[0024] The noise reduction processing of the enhanced single-channel audio signal corresponding to the target audio frame by using the pre-trained audio noise reduction model to obtain a noise-reduced audio frame corresponding to the target audio frame includes:
[0025] According to the enhanced single-channel audio signal corresponding to the target audio frame, an amplitude spectrum or a logarithmic amplitude spectrum and a phase spectrum are obtained.
[0026] The obtained amplitude spectrum or logarithmic amplitude spectrum is input into the pre-trained audio noise reduction model to obtain a noise-reduced amplitude spectrum output by the audio noise reduction model.
[0027] An inverse short-time Fourier transform is performed on a frequency domain signal composed of the phase spectrum and the noise-reduced amplitude spectrum to obtain a noise-reduced audio frame corresponding to the target audio frame.
[0028] In a possible implementation, the audio noise reduction model is trained by using a training noisy multi-channel audio and a training clean multi-channel audio corresponding to the training noisy multi-channel audio, and the training noisy multi-channel audio is obtained by superimposing noise on the training clean multi-channel audio.
[0029] The training process of the audio noise reduction model includes:
[0030] According to the training clean multi-channel audio and the noise, an ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio is determined.
[0031] According to the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio, a beam parameter corresponding to each audio frame of the training noisy multi-channel audio is determined.
[0032] According to the beam parameter corresponding to each audio frame of the training noisy multi-channel audio, a beamforming processing is performed on the training noisy multi-channel audio and the training clean multi-channel audio respectively and frame by frame to obtain an enhanced single-channel audio signal corresponding to the training noisy multi-channel audio and the training clean multi-channel audio respectively.
[0033] The audio noise reduction model is trained so that a noise-reduced audio signal obtained by performing frame-by-frame noise reduction on an enhanced single-channel audio signal corresponding to the training noisy multi-channel audio approaches an enhanced single-channel audio signal corresponding to the training clean multi-channel audio.
[0034] The second aspect of the present application provides an audio noise reduction device, comprising: an ideal ratio mask determination module, a beam parameter determination module, a beam forming processing module, and an audio noise reduction module.
[0035] The ideal ratio mask determination module is configured to determine an ideal ratio mask corresponding to a previous audio frame of a target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, wherein the target audio frame is an audio frame to be noise-reduced in noisy multi-channel audio.
[0036] The beam parameter determination module is configured to determine a beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame.
[0037] The beam forming processing module is configured to perform beam forming processing on the target audio frame according to the beam parameter corresponding to the target audio frame, to obtain an enhanced single-channel audio signal corresponding to the target audio frame.
[0038] The audio noise reduction module is configured to perform noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using a pre-trained audio noise reduction model, to obtain a noise-reduced audio frame corresponding to the target audio frame.
[0039] The third aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0040] The memory is configured to store a computer program.
[0041] The processor is configured to execute the computer program, so that the electronic device can implement the steps of the audio noise reduction method described in any one of the above aspects.
[0042] The fourth aspect of the present application provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the audio noise reduction method described in any one of the above aspects.
[0043] The fifth aspect of the present application provides a computer program product, comprising computer readable instructions, which, when executed on an electronic device, cause the electronic device to implement the steps of the audio noise reduction method described in any one of the above aspects.
[0044] By means of the technical scheme, the audio noise reduction method provided in the application first determines an ideal ratio mask corresponding to a previous audio frame of a target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, then determines beam parameters corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, then performs beamforming processing on the target audio frame according to the beam parameters corresponding to the target audio frame to obtain an enhanced single-channel audio signal corresponding to the target audio frame, and finally performs noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using a pre-trained audio noise reduction model to obtain a noise-reduced audio frame corresponding to the target audio frame. The audio noise reduction method provided in the application is a noise reduction method combining adaptive beamforming and an audio noise reduction model. The ideal ratio mask determined according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame is used to determine the beam parameters corresponding to the target audio frame. Since the ideal ratio mask is determined based on the noise-reduced audio frame corresponding to the previous audio frame, the beam parameters determined based on the ideal ratio mask are relatively accurate. Then, the target audio frame is processed by beamforming based on the relatively accurate beam parameters, and an audio signal with significantly improved signal quality can be obtained. The audio signal with significantly improved signal quality is further subjected to noise reduction by the audio noise reduction model, and an audio signal with further improved signal quality can be obtained. The audio noise reduction method provided in the application fully combines the advantages of adaptive beamforming and the audio noise reduction model, can effectively suppress noise in the audio, and has good noise reduction effect. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0046] Figure 1 A schematic diagram of a system architecture related to the present application;
[0047] Figure 2 A hardware structure schematic diagram of a terminal provided in an embodiment of the present application;
[0048] Figure 3 A hardware structure schematic diagram of a server provided in an embodiment of the present application;
[0049] Figure 4 A flowchart of an audio noise reduction method provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of noise reduction on the tth audio frame provided in an embodiment of the present application;
[0051] Figure 6 A structural schematic diagram of an audio noise reduction device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0052] The embodiments of the present application are described below in conjunction with the accompanying drawings. The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0053] The embodiments of the present application are described below in conjunction with the accompanying drawings. It is known to those skilled in the art that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0054] The terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or equipment containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0055] In one possible implementation, as shown in Figure 1 The system architecture involved in the present application can include a terminal 101 and a server 102, and the terminal 101 can interact with the server 102 through a network (wired network or wireless network). The server 102 can include one or more servers (as an example, one server is included in the server 102). Figure 1 The terminal 101 can obtain noisy multi-channel audio, transmit the noisy multi-channel audio to the server 102 through the network, and the server 102 can reduce noise of the noisy multi-channel audio by using the audio noise reduction method provided by the present application, and transmit the noise reduction result to the terminal 101 through the network.
[0056] In another possible implementation, the system architecture involved in the present application can include a terminal. The terminal has strong data processing capability. The terminal can obtain noisy multi-channel audio, and reduce noise of the noisy multi-channel audio by using the audio noise reduction method provided by the present application.
[0057] Next, the product form of the terminal described above is described.
[0058] The terminal described above can be a mobile phone, a tablet computer, a wearable device, a robot, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like, and embodiments of the present application do not limit the same.
[0059] Figure 2 An optional hardware structure diagram of the terminal is shown.
[0060] Reference Figure 2 As shown, the terminal can include a radio frequency unit 210, a memory 220, an input unit 230, a display unit 240, a camera 250 (optional), an audio circuit 260 (optional), a speaker 261 (optional), a microphone 262 (optional), a headphone jack 263 (optional), a processor 270, an external interface 280, a power supply 290, and the like. Those skilled in the art can understand that the terminal can include more or fewer components than those shown, or combine some components, or different components. Figure 2 The above is merely an example of the terminal and does not limit the terminal, which can include more or fewer components than those shown, or combine some components, or different components.
[0061] The input unit 230 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the terminal. Specifically, the input unit 230 can include a touch screen 231 (optional) and / or other input devices 232. The touch screen 231 can collect user touch operations (such as user operations on or near the touch screen using a finger, a joint, a stylus, or any suitable object) and drive the corresponding connection device according to the pre-set program. The touch screen can detect the user's touch action on the touch screen, convert the touch action into a touch signal and send it to the processor 270, and can receive commands from the processor 270 and execute them; the touch signal at least includes touch point coordinate information. The touch screen 231 can provide an input interface and an output interface between the terminal and the user. In addition, the touch screen can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 231, the input unit 230 can also include other input devices. Specifically, the other input devices 232 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, on-off buttons, etc.), trackballs, mice, joysticks, and the like.
[0062] The display unit 240 can be used to display information input by a user or information provided to the user, various menus of the terminal, an interactive interface, file display, and / or playing of any kind of multimedia file.
[0063] The memory 220 can be used to store instructions and data, and can mainly include a storage instruction area and a storage data area. The storage data area can store various data such as multimedia files, texts, etc. The storage instruction area can store software units such as an operating system, applications, instructions required by at least one function, etc., or their subsets, extended sets. It can also include a non-volatile random access memory, and provide the processor 270 with software and applications including management of hardware, software, and data resources in the computing processing device, support for control software, and support for applications. It is also used for storage of multimedia files, and storage of running programs and applications.
[0064] The processor 270 is the control center of the terminal, and connects various parts of the terminal through various interfaces and lines. It performs various functions of the terminal and processes data by running or executing instructions stored in the memory 220 and calling data stored in the memory 220, thereby performing overall control of the terminal. Optionally, the processor 270 can include one or more processing units. Preferably, the processor 270 can integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and an application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 270. In some embodiments, the processor, the memory, and the like can be implemented on a single chip, and in some embodiments, they can also be implemented on separate chips. The processor 270 can also be used to generate corresponding operation control signals to be sent to corresponding components of the computing processing device, read and process data in the software, especially read and process data and programs in the memory 220, so that each functional module therein performs corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0065] The memory 220 can be used to store software codes related to the SQL statement generation method and the data analysis method, and the processor 270 can execute the software codes in the memory 220, or can also schedule other units (such as the above-mentioned input unit 230 and the display unit 240) to realize corresponding functions.
[0066] The radio frequency unit 210 (optional) can be used for receiving and sending signals in the process of information or communication, for example, receiving the downlink information of the base station and processing by the processor 270; in addition, sending the uplink data to the base station. Generally, the radio frequency unit 210 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 210 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0067] In the embodiments of the present application, the radio frequency unit 210 can send data to other devices, and can also receive data sent by other devices. It should be understood that the radio frequency unit 210 is optional, which can be replaced by other communication interfaces, for example, can be a network interface.
[0068] The terminal also includes a power supply 290 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 270 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system.
[0069] The terminal also includes an external interface 280, which can be a standard Micro USB interface, or can be a multi-pin connector, and can be used for connecting the terminal with other devices for communication, or can be used for connecting a charger to charge the terminal.
[0070] Although not shown, the terminal can also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, different function sensors, etc., which will not be described here.
[0071] Next, the product form of the above server is described.
[0072] Figure 3 A structural diagram of the above server is provided, as shown in Figure 3As shown, the server can include a bus 301, a processor 302, a communication interface 303, and a memory 304. The processor 302, the memory 304, and the communication interface 303 communicate through the bus 301.
[0073] The bus 301 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the middle, but it does not mean that there is only one bus or one type of bus.
[0074] The processor 302 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0075] The memory 304 can include a volatile memory, such as a random access memory (RAM). The memory 304 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), or a solid state drive (SSD).
[0076] The memory 304 can be used to store software code related to the SQL statement generation method and the data analysis method. The processor 302 can call the software code stored in the memory 304, or can schedule other units to implement the corresponding functions.
[0077] The processor in the terminal and the server (for example, the processor 270 and the processor 302) can be a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, or the like), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, a DSP, or the like, or a hardware system without an instruction execution function, such as an ASIC, an FPGA, or the like, or a combination of the hardware system without the instruction execution function and the hardware system with the instruction execution function.
[0078] In view of the poor noise reduction effect of the current audio noise reduction method, the present inventor has conducted research, and has conceived a noise reduction scheme based on adaptive beamforming and a neural network model. The noise reduction process of the scheme is as follows: key beam parameters are estimated based on a neural network model, and beamforming processing is performed on the audio to be reduced in noise according to the key beam parameters, to obtain a noise-reduced signal. The present inventor has found, through research on the above noise reduction scheme, that the noise reduction effect of the scheme is subject to the capability of the adaptive beamforming algorithm itself. Although the signal can be accurately directed and enhanced, the noise reduction performance is still unsatisfactory in the case of burst noise and a low signal-to-noise ratio environment. In view of the unsatisfactory noise reduction effect of the above scheme, the present inventor has continued to conduct research, and through continuous research, has finally proposed an audio noise reduction method with better effect. The audio noise reduction method can effectively suppress noise in audio. Next, the audio noise reduction method provided in the present application will be introduced through the following embodiments.
[0079] Referring to Figure 4 , a flowchart of an audio noise reduction method provided in an embodiment of the present application is shown. The method can include the following steps.
[0080] Step S401: An ideal ratio mask (IRM) corresponding to a previous audio frame of a target audio frame is determined according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame.
[0081] In the above step S401, the target audio frame is an audio frame to be reduced in noise in a noisy multi-channel audio, and the noisy multi-channel audio is obtained by means of a microphone array.
[0082] As Figure 5As shown, assuming the target audio frame is the t-th audio frame of the noisy multi-channel audio, the ideal ratio mask IRM(t-1) corresponding to the (t-1)-th audio frame is determined based on the denoised audio frame corresponding to the (t-1)-th audio frame of the noisy multi-channel audio.
[0083] Step S402: Determine the beam parameters corresponding to the target audio frame based on the ideal ratio mask corresponding to the previous audio frame of the target audio frame.
[0084] Among them, beam parameters are the parameters on which beamforming processing is based. The beam parameters corresponding to the target audio frame may include: the speech presence probability and speech covariance matrix of the target audio frame.
[0085] Assuming the target audio frame is the t-th audio frame of a noisy multi-channel audio stream, then based on the ideal ratio mask IRM(t-1) corresponding to the (t-1)-th audio frame of the noisy multi-channel audio stream, determine the speech presence probability SPP(t) and speech covariance matrix R of the t-th audio frame. SS (t).
[0086] Step S403: Based on the beam parameters corresponding to the target audio frame, perform beamforming processing on the target audio frame to obtain the enhanced single-channel audio signal corresponding to the target audio frame.
[0087] like Figure 5 As shown, assuming the target audio frame is the t-th audio frame of a noisy multi-channel audio stream, then based on the speech presence probability SPP(t) and speech covariance matrix R of the t-th audio frame... SS (t), using the set beamforming algorithm, beamforming processing is performed on the t-th audio frame to obtain the enhanced single-channel audio signal corresponding to the t-th audio frame. Beamforming processing can improve the signal-to-noise ratio of the audio signal.
[0088] Beamforming combines signals from multiple microphones, suppresses signals from non-target directions, and enhances signals from the target direction, thereby enabling focused sound pickup in a specific direction. This effectively improves the signal-to-noise ratio of audio signals and also serves as a noise reduction mechanism.
[0089] In this embodiment, the beamforming algorithm can be the GSC algorithm (Generalized Sidelobe Elimination Algorithm). Of course, this embodiment is not limited to this. The beamforming algorithm can also be the MVDR algorithm (Minimum Variance Distortionless Response Algorithm), the LCMV algorithm (Linear Constraint Minimum Variance Algorithm), etc.
[0090] Additionally, it should be noted that the ideal ratio mask corresponding to an audio frame is a vector. The elements of this vector are the ideal ratio masks corresponding to each frequency point of the audio frame. That is, the ideal ratio masks corresponding to each frequency point of an audio frame constitute the ideal ratio mask corresponding to that audio frame. If the target audio frame is the first audio frame of a noisy multi-channel audio, the ideal ratio mask is initialized to a vector of all 1s, and beamforming processing is performed on the first audio frame based on the initialized ideal ratio mask.
[0091] Step S404: Using the pre-trained audio denoising model, perform denoising processing on the enhanced single-channel audio signal corresponding to the target audio frame to obtain the denoised audio frame corresponding to the target audio frame.
[0092] like Figure 5 As shown, after obtaining the enhanced single-channel audio signal corresponding to the target audio frame, the pre-trained audio denoising model is further used to denoise the enhanced single-channel audio signal corresponding to the target audio frame in order to obtain the denoised audio frame corresponding to the target audio frame.
[0093] The audio denoising model is trained using training data from a pre-built training dataset. Each training data point in the dataset includes a noisy training multi-channel audio file and a corresponding clean training multi-channel audio file. The noisy training multi-channel audio file is obtained by superimposing noise onto the clean training multi-channel audio file. The training process of the audio denoising model will be described in subsequent embodiments.
[0094] Assuming the target audio frame is the t-th audio frame of a noisy multi-channel audio stream, after obtaining the denoised audio frame corresponding to the t-th audio frame, the ideal ratio mask IRM(t) corresponding to the t-th audio frame can be determined based on the denoised audio frame. Then, the beam parameters SPP(t+1) and R corresponding to the (t+1)-th audio frame of the noisy multi-channel audio stream can be determined based on the ideal ratio mask IRM(t). SS (t+1), based on the beam parameters SPP(t+1) and R corresponding to the (t+1)th audio frame. SS (t+1) Beamforming is performed on the (t+1)th audio frame to obtain the enhanced single-channel audio signal corresponding to the (t+1)th audio frame. Using the pre-trained audio denoising model, the enhanced single-channel audio signal corresponding to the (t+1)th audio frame is denoised to obtain the denoised audio frame corresponding to the (t+1)th audio frame. This process is repeated until all audio frames of the noisy multi-channel audio are processed.
[0095] The audio noise reduction method provided in the embodiments of the present application first determines an ideal ratio mask corresponding to a previous audio frame of a target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, then determines beam parameters corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, then performs beamforming processing on the target audio frame according to the beam parameters corresponding to the target audio frame to obtain an enhanced single-channel audio signal corresponding to the target audio frame, and finally performs noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using a pre-trained audio noise reduction model to obtain a noise-reduced audio frame corresponding to the target audio frame. The audio noise reduction method provided in the embodiments of the present application is a noise reduction method combining adaptive beamforming and an audio noise reduction model. The method determines the beam parameters corresponding to the target audio frame by means of the ideal ratio mask determined according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame. Since the ideal ratio mask is determined based on the noise-reduced audio frame corresponding to the previous audio frame, the beam parameters determined based on the ideal ratio mask are relatively accurate. Then, the target audio frame is processed by beamforming based on the relatively accurate beam parameters, and an audio signal with significantly improved signal quality can be obtained. The audio signal with significantly improved signal quality is further reduced by the audio noise reduction model, and an audio signal with further improved signal quality can be obtained. The audio noise reduction method provided in the embodiments of the present application fully combines the advantages of adaptive beamforming and the audio noise reduction model, can effectively suppress noise in the audio, and has a good noise reduction effect.
[0096] In another embodiment of the present application, the specific implementation process of "step S401: determining an ideal ratio mask corresponding to a previous audio frame of a target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame" in the above embodiment is introduced.
[0097] In a possible implementation manner, the process of determining an ideal ratio mask corresponding to a previous audio frame of a target audio frame according to a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame can include: determining the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to a frequency domain amplitude spectrum of the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame and an amplitude spectrum of an enhanced single-channel audio signal obtained by performing beamforming processing on the previous audio frame of the target audio frame.
[0098] Suppose that the target audio frame is the tth audio frame of a noisy multi-channel audio, then the ideal ratio mask IRM(t-1) corresponding to the (t-1)th audio frame is determined according to a frequency domain amplitude spectrum of a noise-reduced audio frame corresponding to the (t-1)th audio frame of the noisy multi-channel audio and an amplitude spectrum of an enhanced single-channel audio signal obtained by performing beamforming processing on the (t-1)th audio frame.
[0099] More specifically, the ideal ratio mask IRM(t-1) corresponding to the t-1th audio frame can be calculated in the manner shown in the following formula (1):
[0100] (1)
[0101] wherein, denotes the frequency domain amplitude spectrum of the noise-reduced audio frame corresponding to the t-1th audio frame, denotes the difference between the amplitude spectrum of the enhanced single-channel audio signal obtained by performing beamforming processing on the t-1th audio frame and the frequency domain amplitude spectrum of the noise-reduced audio frame corresponding to the t-1th audio frame, and it should be noted that the above formula is the calculation manner of IRM(t-1) assuming that the GSC algorithm is used to perform beamforming processing in the noise reduction process, and the calculation manner of IRM(t-1) remains unchanged when other algorithms are used to perform beamforming processing.
[0102] In another embodiment of the present application, the specific implementation process of "determining the beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame" in the above embodiment is introduced.
[0103] The process of determining the beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame can include: determining the speech presence probability of the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame; and determining the speech covariance matrix of the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame and the speech covariance matrix of the previous audio frame of the target audio frame.
[0104] Assuming that the target audio frame is the tth audio frame of the noisy multi-channel audio, the speech presence probability SPP(t) of the tth audio frame is determined according to the ideal ratio mask IRM(t-1) corresponding to the t-1th audio frame, and the speech covariance matrix R SS (t) of the tth audio frame is determined according to the ideal ratio mask IRM(t-1) corresponding to the t-1th audio frame and the speech covariance matrix R SS (t-1) of the t-1th audio frame.
[0105] More specifically, the speech presence probability SPP(t) of the tth audio frame can be calculated in the manner shown in the following formula (2), and the speech covariance matrix R SS (t) of the tth audio frame can be calculated in the manner shown in the following formula (3):
[0106] (2)
[0107] (3)
[0108] wherein, a u and a s is a smoothing coefficient, X mic is a raw noisy audio signal obtained by a main microphone of a microphone array (the microphone array includes multiple microphones, and usually one of the microphones is designated as the main microphone).
[0109] In another embodiment of the present application, the implementation process of "step S403: performing beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame to obtain the enhanced single-channel audio signal corresponding to the target audio frame" in the above embodiment is introduced.
[0110] In a possible implementation manner, the process of performing beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame to obtain the enhanced single-channel audio signal corresponding to the target audio frame can include:
[0111] Step a1, distinguishing the target audio frame into a speech frame and a noise frame.
[0112] In a possible implementation manner, the target audio frame can be distinguished into a speech frame and a noise frame according to the ideal ratio mask corresponding to each audio frame in a forward adjacent frame sequence of the target audio frame. The forward adjacent frame sequence of the target audio frame is the continuous N audio frames before the target audio frame, and N is an integer greater than 1. The specific value of N can be determined according to an actual application scenario.
[0113] Specifically, the process of distinguishing the target audio frame into a speech frame and a noise frame according to the ideal ratio mask corresponding to each audio frame in the forward adjacent frame sequence of the target audio frame can include: if the ideal ratio mask corresponding to each audio frame in the forward adjacent frame sequence of the target audio frame is less than or equal to a set threshold, it is determined that the target audio frame is a noise frame; if the ideal ratio mask corresponding to at least one audio frame in the forward adjacent frame sequence of the target audio frame is greater than the set threshold, it is determined that the target audio frame is a speech frame. It should be noted that if the number of audio frames before the target audio frame is less than N, the target audio frame is determined to be a speech frame. The above embodiment mentions that the ideal ratio mask corresponding to an audio frame includes the ideal ratio mask corresponding to the audio frame at multiple frequency points, and if the average of the ideal ratio masks corresponding to the audio frame at multiple frequency points is less than or equal to the set threshold, it is determined that the ideal ratio mask corresponding to the audio frame is less than or equal to the set threshold.
[0114] For example, N is 5, assuming that the target audio frame is the 8th audio frame of the noisy multi-channel audio, the forward adjacent frame sequence of the target audio frame is the frame sequence of the 7th, 6th, 5th, 4th and 3rd audio frames of the noisy multi-channel audio. If the ideal ratio masks corresponding to the 7th, 6th, 5th, 4th and 3rd audio frames of the noisy multi-channel audio are all less than or equal to the set threshold, it is determined that the 8th audio frame is a noise frame. If at least one of the ideal ratio masks corresponding to the 7th, 6th, 5th, 4th and 3rd audio frames of the noisy multi-channel audio is greater than the set threshold, it is determined that the 8th audio frame is a speech frame.
[0115] In step a2, if the target audio frame is a speech frame, the relative transfer function in the beamforming algorithm is updated according to the beam parameter corresponding to the target audio frame. If the target audio frame is a noise frame, the adaptive noise canceller in the beamforming algorithm is updated according to the beam parameter corresponding to the target audio frame, to obtain the updated beamforming algorithm.
[0116] The relative transfer function and the adaptive noise canceller are two parts of the beamforming algorithm. The relative transfer function is a manifestation of the amplitude difference and phase difference of the microphone signals of the microphone array relative to the main microphone signal. The adaptive noise canceller is an algorithm for noise cancellation using an adaptive filter. Updating the relative transfer function and the adaptive noise canceller is to track the changes of speech and noise in real time during the change of the audio signal, so as to realize adaptive noise reduction.
[0117] In step a3, the target audio frame is subjected to beamforming processing by using the updated beamforming algorithm, to obtain the enhanced single-channel audio signal corresponding to the target audio frame.
[0118] In another embodiment of the present application, the specific implementation process of "step S404: using the pre-trained audio noise reduction model to perform noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame, to obtain the noise-reduced audio frame corresponding to the target audio frame" is introduced.
[0119] In a possible implementation manner, the process of using the pre-trained audio noise reduction model to perform noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame, to obtain the noise-reduced audio frame corresponding to the target audio frame can include:
[0120] In step b1, the amplitude spectrum or the logarithmic amplitude spectrum and the phase spectrum are obtained according to the enhanced single-channel audio signal corresponding to the target audio frame.
[0121] It should be noted that the enhanced single-channel audio signal obtained by performing beamforming processing on the target audio frame is a frequency domain signal, i.e., the enhanced single-channel audio signal corresponding to the target audio frame is a frequency domain signal, and the amplitude spectrum and the phase spectrum of the enhanced single-channel audio signal corresponding to the target audio frame can be extracted. Optionally, after obtaining the amplitude spectrum, the obtained amplitude spectrum can be logarithmized to obtain a logarithmic amplitude spectrum.
[0122] Step b2, input the obtained amplitude spectrum or logarithmic amplitude spectrum into the pre-trained audio noise reduction model to obtain a noise-reduced amplitude spectrum output by the audio noise reduction model.
[0123] The obtained amplitude spectrum or logarithmic amplitude spectrum is input into the pre-trained audio noise reduction model, and the audio noise reduction model performs noise reduction processing on the input amplitude spectrum or logarithmic amplitude spectrum to output a noise-reduced amplitude spectrum.
[0124] Step b3, performing inverse short-time Fourier transform on the frequency domain signal composed of the phase spectrum obtained in step b1 and the noise-reduced amplitude spectrum to obtain a noise-reduced audio frame corresponding to the target audio frame.
[0125] The frequency domain signal composed of the phase spectrum obtained in step b1 and the noise-reduced amplitude spectrum output by the audio noise reduction model is subjected to inverse short-time Fourier transform to obtain a time-domain audio signal, i.e., a noise-reduced audio frame corresponding to the target audio frame.
[0126] The above embodiment mentions that the audio noise reduction model is trained using training noisy multi-channel audio and training clean multi-channel audio corresponding to the training noisy multi-channel audio. Next, the training process of the audio noise reduction model is introduced.
[0127] Before introducing the training process of the audio noise reduction model, the process of obtaining the training noisy multi-channel audio and the training clean multi-channel audio corresponding to the training noisy multi-channel audio is introduced.
[0128] The process of obtaining the training noisy multi-channel audio and the training clean multi-channel audio corresponding to the training noisy multi-channel audio can include: first, using a microphone array (such as a 6-channel microphone array) to collect clean multi-channel audio (the clean multi-channel audio here can be understood as not pure audio, and usually also includes some noise) and noise (such as diffuse noise, point source noise, etc.); obtaining the VAD-processed clean multi-channel audio by performing voice activity detection (VAD) processing on the collected clean multi-channel audio; simulating room impulse responses of different scenes by using a mirror source method; convolving the VAD-processed clean multi-channel audio with the room impulse responses to obtain audio, and superimposing noise (such as diffuse noise, point source noise) on the training clean multi-channel audio to obtain the training noisy multi-channel audio, wherein the training clean multi-channel audio and the noise can be mixed according to a certain signal-to-noise ratio to obtain the training noisy multi-channel audio.
[0129] Next, the training process of the audio noise reduction model is introduced.
[0130] The training process of the audio noise reduction model includes:
[0131] Step c1, determining the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio according to the training clean multi-channel audio and the noise.
[0132] The training clean multi-channel audio and the noise in step c1 are signals used to construct the training noisy multi-channel audio.
[0133] The ideal ratio mask IRM(i, f) corresponding to the i-th audio frame of the training noisy multi-channel audio at frequency point f can be calculated by the following formula:
[0134] IRM(i, f) = |S(i, f)| 2 / (|S(i, f)| 2 +|N(i, f)| 2 ) (4)
[0135] Where |S(i, f)| represents the frequency domain amplitude spectrum of the i-th audio frame of the training clean multi-channel audio, and |N(i, f)| represents the frequency domain amplitude spectrum of the i-th audio frame of the noise.
[0136] The ideal ratio masks corresponding to each frequency point of the i-th audio frame of the training noisy multi-channel audio form the ideal ratio mask IRM(i) corresponding to the i-th audio frame of the training noisy multi-channel audio.
[0137] Step c2, determining the beam parameter corresponding to each audio frame of the training noisy multi-channel audio according to the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio.
[0138] For the i-th audio frame of the training noisy multi-channel audio: according to the ideal ratio mask corresponding to the i-th audio frame, the beam parameter corresponding to the i-th audio frame is determined.
[0139] Step c3, according to the beam parameter corresponding to each audio frame of the training noisy multi-channel audio, the training noisy multi-channel audio and the training clean multi-channel audio are respectively processed frame by frame by beam forming to obtain the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio and the training clean multi-channel audio respectively.
[0140] For the i-th audio frame of the training noisy multi-channel audio: the speech frame and the noise frame are distinguished, if the i-th audio frame is a speech frame, the relative transfer function in the beam forming algorithm is updated according to the beam parameter corresponding to the i-th audio frame, if the i-th audio frame is a noise frame, the adaptive noise canceller in the beam forming algorithm is updated according to the beam parameter corresponding to the i-th audio frame, to obtain the updated beam forming algorithm, the i-th audio frame of the training noisy multi-channel audio is processed by beam forming using the updated beam forming algorithm to obtain the enhanced single-channel audio signal corresponding to the i-th audio frame of the training noisy multi-channel audio. At the same time, the i-th audio frame of the training clean multi-channel audio is processed by beam forming using the updated beam forming algorithm to obtain the enhanced single-channel audio signal corresponding to the i-th audio frame of the training clean multi-channel audio.
[0141] Step c4, the audio noise reduction model is trained with the goal that the noise-reduced audio signal obtained by the audio noise reduction model performing frame-by-frame noise reduction on the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio tends to approach the enhanced single-channel audio signal corresponding to the training clean multi-channel audio.
[0142] Specifically, according to the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio, the amplitude spectrum or the logarithmic amplitude spectrum is obtained, and the amplitude spectrum (or the logarithmic amplitude spectrum) of the enhanced single-channel audio signal corresponding to the training clean multi-channel audio is obtained, the amplitude spectrum or the logarithmic amplitude spectrum obtained according to the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio is input into the audio noise reduction model to obtain the noise-reduced amplitude spectrum output by the audio noise reduction model, the first loss function (such as the mean square error loss function) is determined according to the noise-reduced amplitude spectrum output by the audio noise reduction model (which can be logarithm to obtain the logarithmic amplitude spectrum) and the amplitude spectrum (or the logarithmic amplitude spectrum) of the enhanced single-channel audio signal corresponding to the training clean multi-channel audio, and the parameters of the audio noise reduction model are updated according to the first loss function.
[0143] In another possible implementation, in addition to determining the first loss function according to the amplitude spectrum (which can be logarithmized to obtain a logarithm amplitude spectrum) of the noise-reduced signal output by the audio noise reduction model and the amplitude spectrum (or the logarithm amplitude spectrum) of the enhanced single-channel audio signal corresponding to the clean multi-channel audio in training, an inverse short-time Fourier transform can be performed on a frequency domain signal composed of the amplitude spectrum of the noise-reduced signal output by the audio noise reduction model and the phase spectrum of the enhanced single-channel audio signal corresponding to the noisy multi-channel audio in training, and an inverse short-time Fourier transform can be performed on the enhanced single-channel audio signal corresponding to the clean multi-channel audio in training to obtain two time domain signals, a second loss function (such as a scale-invariant signal-to-noise ratio loss function) can be calculated according to the two time domain signals, and then the audio noise reduction model can be updated in parameters according to the first loss function and the second loss.
[0144] Of course, only the second loss function described above can be determined, and then the audio noise reduction model can be updated in parameters according to the second loss function.
[0145] The audio noise reduction model is iteratively trained multiple times according to the training manner described above by using the training data in the training data set until a training end condition (such as model convergence, reaching a set number of training iterations, and the like) is met.
[0146] The audio noise reduction model in this embodiment can be, but is not limited to, an FSMN (Feedforward Sequential Memory Networks, feedforward sequential memory neural network). The FSMN is improved relative to a traditional DNN (Deep Neural Networks, deep neural network), a memory module is added in a hidden layer, and node information at past time points is stored, so that related information of a time sequence can be effectively used, and feedback connection does not need to be added compared to an RNN network, and therefore causality of the network can be maintained.
[0147] The audio noise reduction method provided in the embodiments of the present application is a noise reduction scheme combining adaptive beamforming and an audio noise reduction model (neural network). When performing noise reduction on a target audio frame, the scheme determines an ideal ratio mask corresponding to a previous audio frame of the target audio frame by means of the audio noise reduction model based on a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, and then determines beam parameters corresponding to the target audio frame based on the ideal ratio mask corresponding to the previous audio frame of the target audio frame, and then performs beamforming processing on the target audio frame based on the beam parameters corresponding to the target audio frame. The audio signal obtained through the beamforming processing is subjected to deep post-filtering processing by the audio noise reduction model, and the result of the processing is used for noise reduction on a next audio frame. In this way, a closed-loop system is established, which promotes close cooperation between adaptive beamforming and neural network noise reduction. Since the ideal ratio mask corresponding to the previous audio frame of the target audio frame is determined based on the noise reduction result of the previous audio frame, the ideal ratio mask corresponding to the previous audio frame of the target audio frame can be used to determine relatively accurate beam parameters. Then, the beamforming processing based on the relatively accurate beam parameters can obtain an audio signal with significantly improved signal quality. The audio noise reduction method provided in the embodiments of the present application fully combines the advantages of adaptive beamforming and neural network technology, and can significantly improve the audio noise reduction effect.
[0148] The audio noise reduction method provided in the embodiments of the present application is introduced above, and the device corresponding to the audio noise reduction method is introduced below.
[0149] Please refer to Figure 6 , Figure 6 FIG. 1 is a structural schematic diagram of an audio noise reduction device provided in the embodiments of the present application. The audio noise reduction device can include an ideal ratio mask determination module 601, a beam parameter determination module 602, a beamforming processing module 603, and an audio noise reduction module 604.
[0150] The ideal ratio mask determination module 601 is configured to determine an ideal ratio mask corresponding to a previous audio frame of a target audio frame based on a noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, where the target audio frame is an audio frame to be noise reduced in a noisy multi-channel audio.
[0151] The beam parameter determination module 602 is configured to determine beam parameters corresponding to the target audio frame based on the ideal ratio mask corresponding to the previous audio frame of the target audio frame.
[0152] The beamforming processing module 603 is configured to perform beamforming processing on the target audio frame based on the beam parameters corresponding to the target audio frame, to obtain an enhanced single-channel audio signal corresponding to the target audio frame.
[0153] The audio noise reduction module 604 is configured to perform noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using the pre-trained audio noise reduction model, to obtain a noise-reduced audio frame corresponding to the target audio frame.
[0154] In a possible implementation, the ideal ratio mask determination module 601, when determining the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, is specifically configured to:
[0155] determine the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to the frequency domain amplitude spectrum of the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame and the amplitude spectrum of the enhanced single-channel audio signal obtained by performing beamforming processing on the previous audio frame of the target audio frame.
[0156] In a possible implementation, the beam parameter determination module 602, when determining the beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, is specifically configured to:
[0157] determine the speech presence probability of the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame;
[0158] In a possible implementation, the beamforming processing module 603, when performing beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame to obtain the enhanced single-channel audio signal corresponding to the target audio frame, is specifically configured to:
[0159] perform speech frame and noise frame discrimination on the target audio frame;
[0160] if the target audio frame is a speech frame, update the relative transfer function in the beamforming algorithm according to the beam parameter corresponding to the target audio frame, and if the target audio frame is a noise frame, update the adaptive noise canceller in the beamforming algorithm according to the beam parameter corresponding to the target audio frame, to obtain an updated beamforming algorithm;
[0161] perform beamforming processing on the target audio frame by using the updated beamforming algorithm, to obtain the enhanced single-channel audio signal corresponding to the target audio frame.
[0162] In a possible implementation, the beamforming processing module 603, when performing speech frame and noise frame discrimination on the target audio frame, is specifically configured to:
[0163] The target audio frame is determined as a noise frame if the ideal ratio mask corresponding to each audio frame in the sequence of forward adjacent frames of the target audio frame is less than or equal to a set threshold value.
[0164] In a possible implementation, when the beamforming processing module 603 determines the target audio frame as a speech frame or a noise frame according to the ideal ratio mask corresponding to each audio frame in the sequence of forward adjacent frames of the target audio frame, the beamforming processing module 603 is specifically configured to:
[0165] The target audio frame is determined as a noise frame if the ideal ratio mask corresponding to each audio frame in the sequence of forward adjacent frames of the target audio frame is less than or equal to a set threshold value.
[0166] The target audio frame is determined as a speech frame if the ideal ratio mask corresponding to at least one audio frame in the sequence of forward adjacent frames of the target audio frame is greater than the set threshold value.
[0167] The enhanced single-channel audio signal corresponding to the target audio frame is a frequency domain signal.
[0168] In a possible implementation, when the audio noise reduction module 604 performs noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using the pre-trained audio noise reduction model to obtain a noise-reduced audio frame corresponding to the target audio frame, the audio noise reduction module 604 is specifically configured to:
[0169] The amplitude spectrum or the logarithmic amplitude spectrum and the phase spectrum are obtained according to the enhanced single-channel audio signal corresponding to the target audio frame.
[0170] The obtained amplitude spectrum or logarithmic amplitude spectrum is input into the pre-trained audio noise reduction model to obtain a noise-reduced amplitude spectrum output by the audio noise reduction model.
[0171] The frequency domain signal composed of the phase spectrum and the noise-reduced amplitude spectrum is subjected to inverse short-time Fourier transform to obtain the noise-reduced audio frame corresponding to the target audio frame.
[0172] In a possible implementation, the audio noise reduction model is trained by using a training noisy multi-channel audio and a training clean multi-channel audio corresponding to the training noisy multi-channel audio, and the training noisy multi-channel audio is obtained by superimposing noise on the training clean multi-channel audio.
[0173] The audio noise reduction apparatus can further include a model training module. The model training module is configured to:
[0174] The ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio is determined according to the training clean multi-channel audio and the noise.
[0175] determine the beam parameter corresponding to each audio frame of the training noisy multi-channel audio according to the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio;
[0176] perform beamforming processing on the training noisy multi-channel audio and the training clean multi-channel audio respectively according to the beam parameter corresponding to each audio frame of the training noisy multi-channel audio, to obtain the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio and the training clean multi-channel audio respectively;
[0177] so that the audio noise reduction model is trained to make the noise-reduced audio signal obtained by performing noise reduction on the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio tend to the enhanced single-channel audio signal corresponding to the training clean multi-channel audio.
[0178] The audio noise reduction device provided in the embodiments of the present application first determines the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, then determines the beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, then performs beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame, to obtain the enhanced single-channel audio signal corresponding to the target audio frame, and finally performs noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using the audio noise reduction model trained in advance, to obtain the noise-reduced audio frame corresponding to the target audio frame. The audio noise reduction device provided in the embodiments of the present application is a noise reduction device combining adaptive beamforming and an audio noise reduction model. The device accurately determines the beam parameter corresponding to the target audio frame by means of the ideal ratio mask determined according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, and then performs beamforming processing on the target audio frame according to the beam parameter corresponding to the target audio frame. By performing beamforming processing on the target audio frame, an audio signal with directional enhancement and improved signal-to-noise ratio can be obtained. The audio signal is further reduced by the audio noise reduction model, and an audio signal with higher directional enhancement and signal-to-noise ratio can be obtained. It can be seen that the audio noise reduction device provided in the present application can effectively suppress the noise in the audio and has good noise reduction effect.
[0179] The embodiments of the present application further provide an electronic device, which can include at least one processor, at least one communication interface, at least one memory and at least one communication bus.
[0180] In the embodiments of the present application, the number of processors, communication interfaces, memories and communication buses is at least one, and the processors, communication interfaces, memories and communication buses complete communication with each other through the communication bus;
[0181] The processor can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the application.
[0182] The memory can include a high-speed RAM memory, and can also include a non-volatile memory, such as at least one disk memory.
[0183] The memory stores a program, and the processor can call the program stored in the memory, and the program is used to implement the steps of the audio noise reduction method provided by the above embodiments.
[0184] The embodiments of the application further provide a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the steps of the audio noise reduction method provided by the above embodiments.
[0185] The embodiments of the application further provide a computer program product, which includes computer readable instructions, and when the computer readable instructions run on an electronic device, the electronic device can implement the steps of the audio noise reduction method provided by the above embodiments.
[0186] In addition, it should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiments. In addition, the connection relationship between the modules in the apparatus embodiments provided by the application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0187] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware, and of course can also be implemented by means of dedicated hardware including special integrated circuit, special CPU, special memory, special component, etc. Generally, any function completed by computer program can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, software program implementation is a better embodiment. Based on such understanding, the technical solution of the application or the part of the application which makes contribution to the prior art can be embodied in the form of software product, which is stored in readable storage medium, such as computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a plurality of instructions for making a computer device (which can be personal computer, training device or network device, etc.) execute the method described in various embodiments of the application.
[0188] In the above embodiments, the implementation can be achieved by software, hardware, firmware or any combination thereof, entirely or partially. When implemented by software, the implementation can be achieved in the form of a computer program product, entirely or partially.
[0189] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function described in the embodiments of the application is generated entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as training device, data center, etc. integrated with one or more available media sets. The available medium can be magnetic medium (such as floppy disk, hard disk, magnetic tape), optical medium (such as DVD) or semiconductor medium (such as solid state disk (SSD)) etc.
Claims
1. An audio noise reduction method, characterized by, Comprise: According to the corresponding noise reduction of the previous audio frame of the target audio frame, determine the ideal ratio mask corresponding to the previous audio frame of the target audio frame, wherein the target audio frame is the audio frame to be de-noised in the noisy multi-channel audio; According to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, determine the beam parameter corresponding to the target audio frame; According to the beam parameter corresponding to the target audio frame, the target audio frame is beamformed to obtain the enhanced single-channel audio signal corresponding to the target audio frame; Using the pre-trained audio noise reduction model, the enhanced single-channel audio signal corresponding to the target audio frame is de-noised to obtain the de-noised audio frame corresponding to the target audio frame.
2. The audio noise reduction method of claim 1, wherein, According to the corresponding noise reduction of the previous audio frame of the target audio frame, determine the ideal ratio mask corresponding to the previous audio frame of the target audio frame, comprising: According to the frequency domain amplitude spectrum of the de-noised audio frame corresponding to the previous audio frame of the target audio frame and the amplitude spectrum of the enhanced single-channel audio signal obtained by beamforming processing on the previous audio frame of the target audio frame, the ideal ratio mask corresponding to the previous audio frame of the target audio frame is determined.
3. The audio noise reduction method of claim 1, wherein, According to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, determine the beam parameter corresponding to the target audio frame, comprising: according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame, determine the speech existence probability of the target audio frame; According to the ideal ratio mask corresponding to the previous audio frame of the target audio frame and the speech covariance matrix of the previous audio frame of the target audio frame, determine the speech covariance matrix of the target audio frame.
4. The audio noise reduction method of claim 1, wherein, According to the beam parameter corresponding to the target audio frame, the target audio frame is beamformed to obtain the enhanced single-channel audio signal corresponding to the target audio frame, comprising: Discriminate the speech frame and noise frame of the target audio frame; If the target audio frame is a speech frame, update the relative transfer function in the beamforming algorithm according to the beam parameter corresponding to the target audio frame, if the target audio frame is a noise frame, update the adaptive noise canceller in the beamforming algorithm according to the beam parameter corresponding to the target audio frame, to obtain the updated beamforming algorithm; Using the updated beamforming algorithm, the target audio frame is beamformed to obtain the enhanced single-channel audio signal corresponding to the target audio frame.
5. The audio noise reduction method of claim 4, wherein, The discrimination of the speech frame and noise frame of the target audio frame, comprising: According to the ideal ratio mask corresponding to each audio frame of the previous adjacent frame sequence of the target audio frame, the speech frame and noise frame of the target audio frame are discriminated, wherein the previous adjacent frame sequence of the target audio frame is a frame sequence composed of the previous N audio frames of the target audio frame, and N is an integer greater than 1.
6. The audio noise reduction method of claim 5, wherein, According to the ideal ratio mask corresponding to each audio frame of the previous adjacent frame sequence of the target audio frame, the speech frame and noise frame of the target audio frame are discriminated, comprising: If the ideal ratio mask corresponding to each audio frame in the sequence of forward adjacent frames of the target audio frame is less than or equal to a set threshold, the target audio frame is determined to be a noise frame; If the ideal ratio mask corresponding to at least one audio frame in the sequence of forward adjacent frames of the target audio frame is greater than the set threshold, the target audio frame is determined to be a speech frame.
7. The audio noise reduction method of claim 1, wherein, The enhanced single-channel audio signal corresponding to the target audio frame is a frequency domain signal; The noise reduction processing of the enhanced single-channel audio signal corresponding to the target audio frame by using the pre-trained audio noise reduction model to obtain the noise-reduced audio frame corresponding to the target audio frame includes: According to the enhanced single-channel audio signal corresponding to the target audio frame, the amplitude spectrum or the logarithmic amplitude spectrum and the phase spectrum are obtained; The obtained amplitude spectrum or logarithmic amplitude spectrum is input into the pre-trained audio noise reduction model to obtain the noise-reduced amplitude spectrum output by the audio noise reduction model; The inverse short-time Fourier transform is performed on the frequency domain signal composed of the phase spectrum and the noise-reduced amplitude spectrum to obtain the noise-reduced audio frame corresponding to the target audio frame.
8. The audio noise reduction method of claim 1, wherein, The audio noise reduction model is trained by using a training noisy multi-channel audio and a training clean multi-channel audio corresponding to the training noisy multi-channel audio, and the training noisy multi-channel audio is obtained by superimposing noise on the training clean multi-channel audio; The training process of the audio noise reduction model includes: According to the training clean multi-channel audio and the noise, the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio is determined; According to the ideal ratio mask corresponding to each audio frame of the training noisy multi-channel audio, the beam parameter corresponding to each audio frame of the training noisy multi-channel audio is determined; According to the beam parameter corresponding to each audio frame of the training noisy multi-channel audio, the training noisy multi-channel audio and the training clean multi-channel audio are respectively frame-by-frame beamformed to obtain the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio and the training clean multi-channel audio, respectively; The audio noise reduction model is trained to make the noise-reduced audio signal obtained by the audio noise reduction model on the enhanced single-channel audio signal corresponding to the training noisy multi-channel audio approach the enhanced single-channel audio signal corresponding to the training clean multi-channel audio.
9. An audio noise reduction device, comprising: It includes: An ideal ratio mask determination module, a beam parameter determination module, a beamforming processing module, and an audio noise reduction module; The ideal ratio mask determination module is configured to determine the ideal ratio mask corresponding to the previous audio frame of the target audio frame according to the noise-reduced audio frame corresponding to the previous audio frame of the target audio frame, wherein the target audio frame is an audio frame to be noise-reduced in a noisy multi-channel audio; The beam parameter determination module is configured to determine the beam parameter corresponding to the target audio frame according to the ideal ratio mask corresponding to the previous audio frame of the target audio frame; The beamforming processing module is configured to perform beamforming processing on the target audio frame according to a beam parameter corresponding to the target audio frame, to obtain an enhanced single-channel audio signal corresponding to the target audio frame. The audio noise reduction module is configured to perform noise reduction processing on the enhanced single-channel audio signal corresponding to the target audio frame by using a pre-trained audio noise reduction model, to obtain a noise-reduced audio frame corresponding to the target audio frame.
10. An electronic device, comprising: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is configured to store a computer program; The processor is configured to execute the computer program, so that the electronic device can implement the steps of the audio noise reduction method according to any one of claims 1-8.
11. A computer storage medium, characterized in that The storage medium carries one or more computer programs, which can enable the electronic device to implement the steps of the audio noise reduction method according to any one of claims 1-8 when the one or more computer programs are executed by the electronic device.
12. A computer program product, characterised in that, The computer readable instructions enable the electronic device to implement the steps of the audio noise reduction method according to any one of claims 1-8 when the computer readable instructions are run on the electronic device.
Citation Information
Patent Citations
Foreground voice detection method and device based on microphone array
CN111613247A
Speech enhancement method based on neural network detection and related device thereof
CN118098255A