Array microphone noise suppression method based on adaptive beam forming and deep learning
By combining adaptive beamforming and deep learning technology to process multi-channel audio signals, the problem of insufficient directional noise suppression effect in the prior art is solved, and efficient noise suppression and target signal protection are achieved in complex noise environments.
Patent Information
- Application Number
- CN202510337618.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, deep learning models do not have enough effect on directional noise suppression, and it is difficult to effectively suppress a variety of interference noises in complex acoustic environments, affecting the protection and analysis of target signals.
The array microphone noise suppression method based on adaptive beamforming and deep learning is adopted. By obtaining multi-channel audio signals to generate beamforming weights, combining deep learning models to process signals, and optimizing the deep learning model with the joint optimization model and enhanced joint loss function to achieve effective suppression of directional noise.
It significantly improves the suppression effect of directional noise, effectively protects the target signal in complex noise environments, and improves the signal-to-noise ratio and the quality of the audio signal.
Smart Images

Figure CN120199256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing and noise suppression, and particularly to an array microphone noise suppression method based on adaptive beamforming and deep learning. Background Art
[0002] In the field of audio signal processing, especially in complex acoustic environments, noise suppression technology has always been a hot and difficult research topic. Traditional noise suppression methods, such as fixed beamforming and simple filtering techniques, although they can improve the signal quality to a certain extent, have obvious limitations in dealing with variable noise environments and protecting the integrity of target signals.
[0003] With the rise of deep learning technology, its successful applications in fields such as image recognition and natural language processing have brought new ideas to audio signal processing. Deep learning algorithms can automatically learn features from a large amount of data and have the ability to handle complex non-linear relationships, which makes it show potential advantages in noise suppression.
[0004] However, directly applying deep learning to audio signal processing also faces many challenges. For example, deep learning models based on time-frequency masks (such as U-Net) achieve noise reduction by separating the time-frequency features of speech and noise. However, most existing methods are based on single-channel signals and ignore the spatial information provided by the microphone array, resulting in insufficient suppression effect on directional noise (such as mechanical noise at a fixed position).
[0005] In addition, in specific application scenarios, such as environmental monitoring in poultry farms, it is necessary to accurately collect and analyze poultry calls. However, the sound environment in the farm is extremely complex. In addition to the normal calls of poultry, there are also various interference factors such as mechanical sounds, environmental sounds, wind noise, and possible abnormal calls of poultry. These noises not only mask the target signal but may also mislead subsequent call analysis, affecting the judgment of the health status and behavior patterns of poultry. Therefore, there is an urgent need for an advanced method that can effectively suppress these complex noises while protecting the characteristics of poultry calls. Summary of the Invention
[0006] The present invention provides an array microphone noise suppression method based on adaptive beamforming and deep learning to solve the defect that the deep learning model in the prior art has insufficient suppression effect on directional noise.
[0007] The present invention provides an array microphone noise suppression method based on adaptive beamforming and deep learning, and the method includes:
[0008] Obtain multi-channel audio signals of a sound source, and generate beamforming weights according to the multi-channel audio signals;
[0009] Process the multi-channel audio signal using a deep learning model to obtain an enhanced multi-channel audio signal;
[0010] Process the beamforming weights and the enhanced multi-channel audio signal using a joint optimization model to obtain a target audio signal;
[0011] Optimize the deep learning model using an enhanced joint loss function until the target audio signal meets the noise suppression condition.
[0012] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention, the spatial direction information includes the azimuth angle and elevation angle of the sound source. Calculate the spatial direction information of the target sound source of the multi-channel audio signal through a time delay estimation algorithm, and generate beamforming weights according to the spatial direction information.
[0013] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention, the process of using a deep learning model to process the multi-channel audio signal to obtain an enhanced multi-channel audio signal includes:
[0014] Convert the beamforming weights into time-frequency domain signals to obtain an amplitude spectrum and a phase spectrum. Perform noise suppression processing on the amplitude spectrum through a deep learning network to generate a predicted time-frequency mask, and reconstruct and output the enhanced multi-channel audio signal using the predicted time-frequency mask and the phase spectrum.
[0015] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention, the process of performing noise suppression processing on the amplitude spectrum through a deep learning network to generate a predicted time-frequency mask includes:
[0016] Process the amplitude spectrum using a U-Net network and output a predicted speech-to-noise ratio mask;
[0017] Combine the predicted time-frequency mask and the phase spectrum, perform weighted processing on the amplitude spectrum, generate an enhanced time-domain audio signal through inverse short-time Fourier transform, and dynamically update the noise spectrum estimation to adapt to environmental noise changes.
[0018] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention, the process of using a joint optimization model to process the beamforming weights and the enhanced multi-channel audio signal to obtain a target audio signal includes:
[0019] Construct a loss function,
[0020] L = α·L SNR + β·L SI-SDR + γ·L Mask
[0021] Among them, L is the total loss, L SNR is the SNR loss between the output and the clean speech, L SI-SDR represents the scale-invariant signal-to-noise ratio loss, L Mask represents the cross-entropy between the mask and the ideal binary mask (IBM);
[0022] Calculate the gradient of the total loss L with respect to the beamforming weight w
[0023]
[0024] Among them, is the gradient of the loss function with respect to the beamforming output signal, is the gradient of the beamforming output signal with respect to the weight;
[0025] Calculate the gradient of the beamforming output signal with respect to the weight:
[0026] y BF = w H x(t)
[0027] Among them, y BF is the beamforming output signal, w is the weight, x(t) is a function of the received signal, w H represents the conjugate transpose of w;
[0028] Calculate the chain rule transfer gradient:
[0029]
[0030] represents the rate of change of the loss function L with respect to the weight w to minimize the joint loss function L.
[0031] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention,
[0032] Add a frequency-domain penalty term L freq to the joint optimization model to protect specific frequency bands of the sound source,
[0033]
[0034] Among them, γ is a weight factor, S ∧ (f,τ) is the frequency-domain signal, ||·||2 represents the L2 norm, which is used to calculate the frequency band energy;
[0035] Add a time-domain sparse regularization term L sparse to the joint optimization model to encourage the joint optimization model to produce a sparse output in the time domain,
[0036]
[0037] Among them, L sparse is the time-domain sparse regularization term, λ is the weight factor, and ||·||1 represents the L1 norm, which is used to encourage sparsity.
[0038] According to an array microphone noise suppression method combining adaptive beamforming and deep learning provided by the present invention, the step of optimizing the deep learning model by using the enhanced joint loss function until the target audio signal meets the noise suppression condition includes:
[0039] Convergence of the total loss function:
[0040] L total = L main + γL freq + λL sparse
[0041] The frequency-domain penalty term L freq , the time-domain sparse regularization term L sparse , and other loss terms L main are weighted and summed to obtain the total loss L total . When the change range of the total loss during training is lower than the preset threshold, it is considered that the model converges;
[0042] Verification of frequency band protection:
[0043] On the validation set, if the frequency band energy loss of the enhanced audio signal within a specific frequency band is lower than the set threshold, it is determined that the frequency-domain penalty term is effective, the target audio signal meets the noise suppression condition, and there is no phenomenon of excessive suppression of the signal in the specific frequency band, ensuring that the frequency-domain penalty term effectively avoids signal over-suppression;
[0044] Verification of time-domain sparsity:
[0045] The enhanced audio signal meets the time-domain sparsity index, that is, the proportion of samples whose signal amplitude exceeds the preset amplitude threshold needs to reach or exceed the preset proportion threshold;
[0046] Verification of signal-to-noise ratio improvement:
[0047] On the validation set, the signal-to-noise ratio improvement of the enhanced audio signal compared to the original noisy signal is ≥ the preset improvement threshold, and the target audio signal meets the noise suppression condition.
[0048] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the noise suppression method described in any one of the above technical solutions.
[0049] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the noise suppression method described in any one of the above technical solutions are implemented.
[0050] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the noise suppression method described in any one of the above technical solutions are implemented.
[0051] The adaptive beamforming and deep learning-based array microphone noise suppression method provided by the present invention first obtains multi-channel audio signals of a sound source and generates beamforming weights based on the signals, thereby achieving precise enhancement of the target sound source direction and effective suppression of noise in non-target directions. By using a deep learning model to process the multi-channel audio signals, not only is the spatial information of the microphone array fully exploited, but also the beamforming weights and the enhanced multi-channel audio signals are processed through joint optimization of the model to obtain high-quality target audio signals. In addition, the deep learning model is optimized by enhancing the joint loss function to ensure that the target audio signals meet the noise suppression conditions, thus effectively solving the problem in the prior art that single-channel signal processing methods ignore the spatial information of the microphone array in a complex noise environment and significantly improving the suppression effect on directional noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a schematic flowchart of the adaptive beamforming and deep learning-based array microphone noise suppression method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0055] The following will describe Figure 1 a method for suppressing noise in an array microphone based on adaptive beamforming and deep learning according to the present invention. The method includes:
[0056] Obtain the multi-channel audio signal of the sound source, and generate beamforming weights according to the multi-channel audio signal;
[0057] Process the multi-channel audio signal using a deep learning model to obtain an enhanced multi-channel audio signal;
[0058] Process the beamforming weights and the enhanced multi-channel audio signal using a joint optimization model to obtain a target audio signal;
[0059] Optimize the deep learning model using an enhanced joint loss function until the target audio signal meets the noise suppression condition.
[0060] First, obtain the multi-channel audio signal of the sound source and generate beamforming weights accordingly, which realizes the precise enhancement of the target sound source direction and the effective suppression of noise in non-target directions. Processing the multi-channel audio signal using a deep learning model not only fully exploits the spatial information of the microphone array but also processes the beamforming weights and the enhanced multi-channel audio signal through a joint optimization model to obtain a high-quality target audio signal. Additionally, optimizing the deep learning model using an enhanced joint loss function ensures that the target audio signal meets the noise suppression condition, thus effectively solving the problem in the prior art that the single-channel signal processing method ignores the spatial information of the microphone array in a complex noise environment and significantly improving the suppression effect on directional noise.
[0061] Deploy a linear six-microphone array at a fixed position in the poultry farm to collect sound data. However, the collected sound data is mixed with multiple noises such as mechanical sounds, environmental sounds, wind noises, and non-interfering abnormal poultry sounds, which will interfere with the collection of normal poultry calls. Therefore, it is necessary to process this sound data and establish a poultry sound signal model to achieve the effective extraction and enhancement of the target poultry calls.
[0062] Specifically, the spatial direction information includes the azimuth angle θ of the target sound source s and the elevation angle Φ of the target sound source s , calculate the spatial direction information of the target sound source of the multi-channel audio signal through a time delay estimation algorithm, and generate beamforming weights according to the spatial direction information;
[0063] In order to obtain the poultry target signal s(t), where the time-frequency characteristics of poultry are concentrated in [1 kHz to 8 kHz]
[0064] According to the array signal model, obtain the target signal s(t):
[0065]
[0066] where x(t) ∈ C Mis the received signal vector, a(θ) is the direction vector, and θ s is the azimuth angle of the target sound source, and θ i is the direction of the noise signal, and v(t) is the sensor noise;
[0067] The time difference of the signal arriving at each array element can be converted into a phase difference. The time delay t of the m-th array element relative to the reference array element m is:
[0068]
[0069] where the element spacing is d, the angle between the incident direction of the target sound signal source and the normal of the array is θ s , and c is the speed of light;
[0070] Then, the time difference is converted into a phase difference:
[0071]
[0072] where f0 is the center frequency of the signal and λ is the wavelength of the signal
[0073] The steering vector a(θ s ) is required. The steering vector describes the target sound signal incident on the array from the direction θ. Since the linear six-microphone array is a uniform linear array, the linear array equation is used:
[0074]
[0075] The final array signal model is
[0076]
[0077] x(t) = [x1(t), x2(t),..., x M (t)] T
[0078] where x(t) ∈ C M is the received signal vector;
[0079] Find the azimuth angle θ of the target sound source s :
[0080]
[0081]
[0082] where Δt y and Δt x are the time differences of the target sound source in the y-axis direction and the x-axis direction, respectively.
[0083] Top viewφ s :
[0084]
[0085] Next, the adaptive beamforming technology is used to generate a beamforming weight vector according to the direction of the sound source, suppress noise in non-target directions, and output a beamforming signal.
[0086] The received signal x(t) is weighted and summed by the weight vector, and the output signal is:
[0087] y BF (t) = w H x(t)
[0088] Among them, w H Indicates w MVDR The conjugate transpose of y BF (t) is its output signal.
[0089] The conjugate transpose w of the weight vector obtained by the MVDR beamforming algorithm is obtained H :
[0090] w H =a H (θ s )R nn -1a(θ s )a(θ s ) H R nn -1a(θ s ) H (θ s )R nn -1a(θ s )a(θ s ),
[0091] in, is the noise covariance matrix.
[0092] In order to minimize the power output of the array while keeping the directivity of the desired signal undistorted, two conditions need to be optimized:
[0093] min w w H R mm w
[0094] subject to w H a(θ s )=1
[0095] Among them, w is the weight vector, R mm is the autocorrelation matrix of the array received signal, a(θ s ) is from the desired signal direction θs The steering vector incident on the array.
[0096] To solve this constrained optimization problem, it is necessary to introduce the Lagrange multiplier λ and construct the Lagrangian function:
[0097] L(w,λ) = w H R nn w + λ(1 - w H a(θ s ))
[0098] where λ is the Lagrange multiplier, and to find the optimal weight, we need to take the derivative of w and set it to 0:
[0099]
[0100]
[0101]
[0101] w H a(θ s ) = 1,
[0102]
[0103] Simplify to obtain the Lagrange multiplier λ as:
[0104]
[0105] Then substitute λ into the expression of w:
[0106]
[0107] Use the obtained conjugate transpose w H , and solve for the output signal for beamforming:
[0108] y BF (t) = w H x(t),
[0109] Simplify to obtain:
[0110]
[0111] Furthermore, the processing of the multi-channel audio signal by the deep learning model to obtain the enhanced multi-channel audio signal includes:
[0112] Convert the beamforming weights into time-frequency domain signals to obtain the amplitude spectrum and phase spectrum. Perform noise suppression processing on the amplitude spectrum through a deep learning network to generate a predicted time-frequency mask, and use the predicted time-frequency mask and phase spectrum to reconstruct and output the enhanced multi-channel audio signal.
[0113] Specifically, for deep learning enhanced training, first, the output signal y BF (t) is transformed into the time-frequency domain through the short-time Fourier transform:
[0114]
[0115] where ω(t - τ) is the window function, τ is the time variable representing the center position of the current window, and ω is the frequency variable, which is used later to generate a graph of frequency varying with time.
[0116] The amplitude spectrum |y BF (f, τ)| and the phase spectrum ∠y BF (f, τ)
[0117] Next, a U-Net network, a popular convolutional neural network architecture particularly suitable for image segmentation tasks, is designed. Taking the amplitude spectrum as the input, it outputs a mask M(f, τ) ∈ [0, 1] related to the signal-to-noise ratio of the speech and makes a prediction for the mask:
[0118] M(f, τ) = U-N(y BF (f, τ))
[0119] The Net network module is used to process the amplitude spectrum to predict the signal-to-noise ratio mask of the speech. In this network, the encoder part receives the amplitude spectrum input with the shape of (B, C, H, W) and extracts and compresses features through a series of convolutional layers, activation functions, batch normalization layers, and pooling layers.
[0120] Subsequently, the decoder part gradually restores the spatial dimension of the feature map through deconvolutional layers and upsampling layers, while using skip connections to retain detailed information. Finally, the output layer uses the Sigmoid activation function to limit the prediction result between 0 and 1, generating a mask representing the signal-to-noise ratio of the speech, where a larger value means that this frequency domain component is more likely to be an audio signal.
[0121] After obtaining the mask, it is weighted and then transformed back to the time-domain signal S^(t) through the inverse Fourier transform:
[0122] |S^(f, τ)| = M(f, τ) Θ |Y BF (f, τ)|
[0123] ∠S^(f, τ) = ∠Y BF (f, τ)
[0124] Reconstruct the frequency-domain signal s^(f, τ):
[0125] s^(f, τ) = |S^(f, τ)|e j∠S^(f,τ)
[0126] Then, convert the frequency-domain signal s^(f,τ) into a time-domain signal S^(t):
[0127]
[0128] Preferably, processing the beamforming weights and the enhanced multi-channel audio signal by using the joint optimization model to obtain a target audio signal includes:
[0129] Construct a loss function,
[0130] L = α·L SNR + β·L SI-SDR + γ·L Mask
[0131] where L is the total loss, L SNR is the SNR loss between the output and the clean speech, L SI-SDR represents the scale-invariant signal-to-noise ratio loss, and L Mask represents the cross-entropy between the mask and the ideal binary mask (IBM);
[0132] Calculate the gradient of the total loss L with respect to the beamforming weight w
[0133]
[0134] where, is the gradient of the loss function with respect to the beamforming output signal, is the gradient of the beamforming output signal with respect to the weight;
[0135] Calculate the gradient of the beamforming output signal with respect to the weight:
[0136] y BF = w H x(t)
[0137] where y BF is the beamforming output signal, w is the weight, x(t) is a function of the received signal, and w H represents the conjugate transpose of w;
[0138] Calculate the chain rule transfer gradient:
[0139]
[0140] represents the rate of change of the loss function L with respect to the weight w, ensuring that the beamforming weight w and the deep learning parameters are effectively updated during training to minimize the joint loss function L, not only improving the effects of noise suppression and speech enhancement, but also making the entire system more robust and adaptable to different noise environments.
[0141] Furthermore, a frequency-domain penalty term \(L\) is added to the joint optimization model freq to protect specific frequency bands of the sound source,
[0142]
[0143] where \(\gamma\) is a weight factor used to balance the relationship between the frequency-domain penalty term and other loss terms, and \(S(f,\tau)\) is the frequency-domain signal, and \(\|\cdot\|_2\) represents the \(L_2\) norm, which is used to calculate the frequency-band energy; ∧ (f,τ) is the frequency-domain signal, ||·||2 represents the L2 norm, which is used to calculate the frequency-band energy;
[0144] A time-domain sparse regularization term \(L\) is added to the joint optimization model sparse to encourage the joint optimization model to produce a sparse output in the time domain,
[0145]
[0146] where \(L\) sparse is the time-domain sparse regularization term, \(\lambda\) is the weight factor, and \(\|\cdot\|_1\) represents the \(L_1\) norm, which is used to encourage sparsity.
[0147] Optimizing the deep learning model using the enhanced joint loss function until the target audio signal meets the noise suppression conditions includes:
[0148] Convergence of the total loss function:
[0149] \(L\) total = \(L\) main + \(\lambda L\) freq + \(\lambda L\) sparse
[0150] Adding the frequency-domain penalty term \(L\) freq , the time-domain sparse regularization term \(L\) sparse , and other loss terms \(L\) main and summing them with weights to obtain the total loss \(L\) total . When the change range of the total loss during training is lower than the preset threshold, the model is considered to have converged;
[0151] Frequency-band protection verification:
[0152] On the validation set, if the frequency-band energy loss of the enhanced audio signal within a specific frequency band is lower than the set threshold, it is determined that the frequency-domain penalty term is effective, the target audio signal meets the noise suppression conditions, and there is no over-suppression of the specific frequency-band signal, ensuring that the frequency-domain penalty term effectively avoids signal over-suppression;
[0153] Time-domain sparsity verification:
[0154] The enhanced audio signal meets the time-domain sparsity index, that is, the proportion of samples whose signal amplitude exceeds the preset amplitude threshold needs to reach or exceed the preset proportion threshold;
[0155] Verification of SNR improvement:
[0156] On the validation set, the improvement in the signal-to-noise ratio of the enhanced audio signal compared to the original noisy signal is ≥ the preset improvement threshold, and the target audio signal meets the noise suppression condition.
[0157] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the noise suppression method described in any one of the above technical solutions are implemented.
[0158] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0159] On the other hand, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the noise suppression method described in any one of the above technical solutions are implemented.
[0160] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the noise suppression method described in any one of the above technical solutions are implemented.
[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for noise suppression of an array microphone based on adaptive beamforming and deep learning, characterized in that: The method comprises: Acquire a multi-channel audio signal of a sound source, and generate beamforming weights according to the multi-channel audio signal; Processing the multi-channel audio signal using a deep learning model to obtain an enhanced multi-channel audio signal; Processing the beamforming weights and the enhanced multi-channel audio signal using a joint optimization model to obtain a target audio signal; The deep learning model is optimized using an enhanced joint loss function until the target audio signal meets the noise suppression conditions.
2. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 1, characterized in that: The spatial direction information includes the azimuth angle and the elevation angle of the sound source. The spatial direction information of the target sound source of the multi-channel audio signal is calculated by a time delay estimation algorithm, and a beamforming weight is generated according to the spatial direction information.
3. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 1, characterized in that: The method of processing the multi-channel audio signal by using a deep learning model to obtain an enhanced multi-channel audio signal includes: The beamforming weights are converted into time-frequency domain signals to obtain an amplitude spectrum and a phase spectrum, the amplitude spectrum is subjected to noise suppression processing through a deep learning network to generate a predicted time-frequency mask, and the predicted time-frequency mask and phase spectrum are used to reconstruct and output an enhanced multi-channel audio signal.
4. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 3, characterized in that: The step of performing noise suppression processing on the amplitude spectrum through a deep learning network to generate a predicted time-frequency mask includes: Use the U-Net network to process the amplitude spectrum and output the predicted speech-to-noise ratio mask; The predicted time-frequency mask and phase spectrum are combined to perform weighted processing on the amplitude spectrum, an enhanced time-domain audio signal is generated by inverse short-time Fourier transform, and the noise spectrum estimation is dynamically updated to adapt to the change of environmental noise.
5. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 1, characterized in that: The method of processing the beamforming weights and the enhanced multi-channel audio signal by using a joint optimization model to obtain a target audio signal includes: Construct the loss function, L=α·L SNR +β·L SI-SDR +γ·L Mask Among them, L is the total loss, L SNR is the SNR loss between the output and the clean speech, L SI-SDR represents the scale-invariant signal-to-noise ratio loss, L Mask represents the cross entropy between the mask and the ideal binary mask (IBM); Calculate the gradient of the total loss L with respect to the beamforming weights w in, is the gradient of the loss function to the beamforming output signal, is the gradient of the beamforming output signal with respect to the weight; Calculate the gradient of the beamforming output signal with respect to the weights: <h2 style=";text-align:left;direction:ltr">y<h2 style=";text-align:left;direction:ltr"> BF <h2 style=";text-align:left;direction:ltr"> =w<h2 style=";text-align:left;direction:ltr"> H <h2 style=";text-align:left;direction:ltr"> xt Among them, y BF is the beamforming output signal, w is the weight, xt is the function of the received signal, w H represents the conjugate transpose of w; The chain rule transfer gradient is calculated: It represents the rate of change of the loss function L with respect to the weight w to minimize the joint loss function L.
6. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 5, characterized in that: Add a frequency domain penalty term L in the joint optimization model freq , used to protect specific frequency bands of sound sources, Among them, γ is the weight factor, S ∧ f,τ is the frequency domain signal, ∥·∥2 represents the L2 norm, which is used to calculate the frequency band energy; Add the time domain sparse regularization term L in the joint optimization model sparse , encouraging the joint optimization model to produce sparse output in the time domain, Among them, L sparse is the time-domain sparse regularization term, λ is the weight factor, and ∥·∥1 represents the L1 norm, which is used to encourage sparsity.
7. The method for noise suppression of an array microphone using adaptive beamforming and deep learning according to claim 6, characterized in that: The method of optimizing the deep learning model by using the enhanced joint loss function until the target audio signal meets the noise suppression condition includes: The total loss function converges: THE total =L main +γL freq +λL sparse The frequency domain penalty term L freq , time domain sparse regularization term L sparse , and other loss terms L main The weighted summation gives the total loss L total , when the total loss changes during training and is lower than the preset threshold, the model is considered to have converged; Frequency band protection verification: On the validation set, if the frequency band energy loss of the enhanced audio signal in a specific frequency band is lower than the set threshold, the frequency domain penalty term is considered effective, the target audio signal meets the noise suppression conditions, and there is no over-suppression of the signal in a specific frequency band, ensuring that the frequency domain penalty term effectively avoids over-suppression of the signal; Time domain sparsity verification: The enhanced audio signal meets the time domain sparsity index, that is, the proportion of samples whose signal amplitude exceeds the preset amplitude threshold must reach or exceed the preset ratio threshold; Signal-to-noise ratio improvement verification: On the validation set, the improvement in the signal-to-noise ratio of the enhanced audio signal compared to the original noisy signal is ≥ the preset improvement threshold, and the target audio signal meets the noise suppression conditions.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the noise suppression method according to any one of claims 1 to 7 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the noise suppression method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the noise suppression method according to any one of claims 1 to 7 are implemented.