Multi-channel microphone array acoustic signal enhancement method and system
By degrading the beamformer in the microphone array acoustic signal processing method into a fusion structure of two beamformers, the problem of waste of computing resources and interference suppression amounts in the number of large microphones is solved, and efficient voice signal enhancement is achieved.
Patent Information
- Application Number
- CN202211372548.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-11-03
AI Technical Summary
The existing microphone array acoustic signal processing method is wasted heavily in the case of large microphone numbers, and it is difficult to flexibly control the amount of interference suppression.
The typical beamformer is degraded into a fusion structure of two beamformers, and preliminary estimation and enhancement of the source signal and the interfering signal is achieved by determining the observed signal vector, covariance matrix and key parameters.
While reducing the amount of computing, the amount of interference suppression can be flexibly controlled and the efficiency of voice signal processing can be improved.
Smart Images

Figure CN116013347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal enhancement, and in particular to a method and system for enhancing acoustic signals of a multi-channel microphone array. Background Art
[0002] Using microphone arrays for speech enhancement is a hot topic in speech signal processing. Microphone arrays consist of a number of microphones arranged in a specific spatial configuration. The collected speech signals contain not only time-frequency information but also spatial information. Therefore, the spatial directional characteristics of the microphone array can be used to track the spatial location of the sound source signal and suppress noise and interference from other directions, thereby enhancing the target speech and improving speech quality. Microphone arrays have been widely used in fields such as human-computer interaction, video conferencing, and speaker recognition.
[0003] When using a microphone array to observe a sound source signal, various surrounding noise and interference are often observed. This wide variety of interference requires the proposed method to enhance the sound signal. Existing methods use a typical beamformer, which requires solving the matrix at every moment and frequency. When the number of microphones is large, matrix inversion is computationally intensive, and even after noise reduction, the amount of interference suppression cannot be effectively controlled. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-channel microphone array acoustic signal enhancement method and system, which reduces the amount of computation while flexibly controlling the amount of interference suppression by degenerating a typical beamformer into a fusion structure of two beamformers.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A method for enhancing a multi-channel microphone array acoustic signal, the method comprising:
[0007] Obtaining sound source signals;
[0008] Based on the sound source signal, an observation signal vector is determined; the observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector;
[0009] determining a covariance matrix based on the noise vector;
[0010] Identify the classical beamformer;
[0011] Inputting the covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal into the classical beamformer to obtain two beamformers to be fused;
[0012] Determining a preliminary estimate of a source signal and a preliminary estimate of an interference signal based on the two beamformers to be fused and the observed signal vector;
[0013] Identify key parameters;
[0014] An enhanced signal is determined based on the key parameters, the preliminary estimate of the source signal, and the preliminary estimate of the interfering signal.
[0015] Optionally, the acquiring of the sound source signal is specifically to observe the sound source signal through a microphone array.
[0016] Optionally, the expression of the observation signal vector is as follows:
[0017] y(f,t)=a(f,t)X(f,t)+b(f,t)ξ(f,t)+v(f,t)
[0018] Where y(f, t) represents the observed signal vector, a(f, t) represents the steering vector of the source signal, b(f, t) represents the steering vector of the interference signal, X(f, t) represents the sound source signal, ξ(f, t) represents the interference signal, and v(f, t) represents the noise vector.
[0019] Optionally, the covariance matrix is determined based on the noise vector using the following formula:
[0020] Φ v (f,t)=E[v(f,t)v H (f,t)]
[0021] Among them, Φ v (f,t) represents the covariance matrix, v(f,t) represents the noise vector, v H (f, t) represents the conjugate transpose of the noise vector, and E represents the mathematical expectation.
[0022] Optionally, the expressions of the two beamformers to be fused are as follows:
[0023]
[0024]
[0025] Among them, W a (f, t) represents the beamformer pointing to the sound source, W b (f,t) represents the beamformer pointing towards the interference, Φ v -1 (f, t) represents the inverse matrix of the noise covariance matrix, a(f, t) represents the steering vector of the source signal, a H (f, t) represents the conjugate transpose of the sound source steering vector, b(f, t) represents the steering vector of the interference signal, and bH (f, t) represents the conjugate transpose of the interference signal steering vector.
[0026] Optionally, the expression of the preliminary estimate of the source signal is as follows:
[0027]
[0028] Among them, z a (f, t) represents the initial estimate of the source signal, represents the conjugate transpose of the beamformer pointing to the sound source, and y(f, t) represents the observation signal vector.
[0029] Optionally, the key parameters are determined using the following formula:
[0030]
[0031]
[0032]
[0033]
[0034]
[0035] Among them, a(f, t) represents the steering vector of the source signal, a H (f, t) represents the conjugate transpose of the sound source steering vector, b(f, t) represents the steering vector of the interference signal, and b H (f, t) represents the conjugate transpose of the interference signal steering vector, Φ v -1 (f, t) represents the inverse matrix of the noise covariance matrix, β represents the interference suppression ratio, k(f, t) is a parameter constructed by the sound source, the interference guidance amount and the noise covariance matrix, ρ(f, t) represents the spatial correlation coefficient, It is a parameter determined by β and the signal, c1(f, t) represents the combination coefficient of the source-directed beamformer, and c2(f, t) represents the combination coefficient of the interference-directed beamformer. If x>0, Relu(x)=x; if x<=0, Relu(x)=0.
[0036] Optionally, the expression of the enhanced signal is as follows:
[0037]
[0038] Where c1(f, t) represents the combination coefficient of the beamformer pointing to the sound source, represents the conjugate of c2(f, t), Z a (f, t) and Zb (f, t) represent the preliminary estimates of the source signal and interference, respectively, and * represents conjugation.
[0039] Based on the above method in the present invention, the present invention further provides a multi-channel microphone array acoustic signal enhancement system, the enhancement system comprising:
[0040] A sound source signal acquisition module, used to acquire a sound source signal;
[0041] An observation signal vector determination module is used to determine an observation signal vector based on the sound source signal; the observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector;
[0042] a covariance matrix determination module, configured to determine a covariance matrix based on the noise vector;
[0043] A classical beamformer determination module, used for determining a classical beamformer;
[0044] a beamformer determination module to be fused, configured to input the covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal into the classical beamformer to obtain two beamformers to be fused;
[0045] A module for determining a preliminary estimate of the source signal and a preliminary estimate of the interference signal, configured to determine a preliminary estimate of the source signal and a preliminary estimate of the interference signal based on the two beamformers to be fused and the observation signal vector;
[0046] A key parameter determination module, used for determining key parameters;
[0047] The enhanced signal determination module is configured to determine an enhanced signal based on the key parameters, a preliminary estimate of the source signal, and a preliminary estimate of the interference signal.
[0048] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0049] The above method in the present invention degenerates a typical beamformer into a fusion structure of two beamformers, which reduces the amount of computation while being able to flexibly control the amount of interference suppression. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1This is a flow chart of a multi-channel microphone array acoustic signal enhancement method according to embodiment 1 of the present invention;
[0052] Figure 2 This is a structural diagram of a multi-channel microphone array sound signal enhancement system according to embodiment 2 of the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] The purpose of the present invention is to provide a multi-channel microphone array acoustic signal enhancement method and system, which reduces the amount of computation while flexibly controlling the amount of interference suppression by degenerating a typical beamformer into a fusion structure of two beamformers.
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] Example 1
[0057] Figure 1 This is a flow chart of a multi-channel microphone array acoustic signal enhancement method according to embodiment 1 of the present invention. Figure 1 As shown, the above method in the present invention includes:
[0058] Step 101: Acquire a sound source signal.
[0059] Specifically, the present invention is a system that uses a microphone array to observe sound source signals, wherein the microphone array refers to an arrangement of microphones, consisting of a certain number of acoustic sensors (generally microphones), and is used to sample and process the spatial characteristics of the sound field.
[0060] Step 102: Based on the sound source signal, determine an observation signal vector; the observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector.
[0061] Assume that the observation signal of the mth (m=0, 1, 2, ..., M-1) microphone in the time-frequency domain is Y m (f, t), where f is frequency and t is time. By listing all observed signals into a vector, we can get:
[0062] y(f,t)=[Y0(f,t) Y1(f,t)…Y M-1(f, t)] T
[0063] =a(f,t)X(f,t)+b(f,t)ξ(f,t)+v(f,t)
[0064] The superscript T stands for transpose. a(f, t) and b(f, t) are two vectors, representing the steering vector of the source signal and the steering vector of the interference signal, respectively. X(f, t) and ξ(f, t) are the time-frequency representations of the source signal and the interference signal, respectively. v(f, t) is a noise vector with the same dimensions as a(f) and b(f). The goal of array processing is to recover X(f, t) from y(f, t) while suppressing the interference components ξ(f, t) and v(f, t).
[0065] Step 103: Determine a covariance matrix based on the noise vector.
[0066] Specifically, the noise vector in the observed signal is defined as a covariance matrix, and the interference suppression ratio is set.
[0067] The covariance matrix of the noise vector v(f, t) is defined as:
[0068] Φ v (f, t) = E[v(f, t)v H (f, t)]
[0069] Set the interference suppression ratio to β, a number between 0 and 1.
[0070] Step 104: Determine the classical beamformer.
[0071] The expression of the classical beamformer is:
[0072]
[0073] Step 105: Input the covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal into the classical beamformer to obtain two beamformers to be fused.
[0074]
[0075]
[0076] Step 106: Determine a preliminary estimate of the source signal and a preliminary estimate of the interference signal based on the two beamformers to be fused and the observed signal vector.
[0077] That is, by applying the two beamformers to be fused to the observed signal vector y(f, t), a preliminary estimate of the source signal and the interference signal can be obtained. The specific expression is as follows:
[0078]
[0079]
[0080] Step 107: Determine key parameters.
[0081] The expressions of key parameters are as follows:
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] Among them, a(f,t) represents the steering vector of the source signal, a H (f, t) represents the conjugate transpose of the sound source steering vector, b(f, t) represents the steering vector of the interference signal, and b H (f, t) represents the conjugate transpose of the interference signal steering vector, Φ v -1 (f, t) represents the inverse matrix of the noise covariance matrix, β represents the interference suppression ratio, k(f, t) is a parameter constructed by the sound source, the interference guidance amount and the noise covariance matrix, ρ(f, t) represents the spatial correlation coefficient, is a parameter determined by β and the signal, c1(f,t) represents the combination coefficient of the source beamformer, and c2(f,t) represents the combination coefficient of the interference beamformer. If x>0, Relu(x)=x; if x<=0, Relu(x)=0.
[0088] Step 108: Determine an enhanced signal based on the key parameters, the preliminary estimate of the source signal, and the preliminary estimate of the interference signal.
[0089] That is, the results of the two beamformers are fused to obtain the overall output of the array, which is recorded as the enhanced signal and is expressed as follows:
[0090]
[0091] Example 2
[0092] Based on the above method in the present invention, the present invention further provides a multi-channel microphone array acoustic signal enhancement system, the system comprising:
[0093] The sound source signal acquisition module 201 is used to acquire a sound source signal.
[0094] The observation signal vector determination module 202 is configured to determine an observation signal vector based on the sound source signal. The observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector.
[0095] The covariance matrix determination module 203 is configured to determine a covariance matrix based on the noise vector.
[0096] The classical beamformer determination module 204 is configured to determine a classical beamformer.
[0097] The beamformer determination module 205 to be fused is configured to input the covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal into the classical beamformer to obtain two beamformers to be fused.
[0098] The module 206 for determining a preliminary estimate of the source signal and a preliminary estimate of the interference signal is configured to determine a preliminary estimate of the source signal and a preliminary estimate of the interference signal based on the two beamformers to be fused and the observed signal vector.
[0099] The key parameter determination module 207 is used to determine the key parameters.
[0100] The enhanced signal determination module 208 is configured to determine an enhanced signal based on the key parameters, the preliminary estimate of the source signal, and the preliminary estimate of the interference signal.
[0101] The multi-channel microphone array acoustic signal enhancement method and system in the present invention have the following beneficial effects:
[0102] The method of the present invention determines a classical beamformer and degenerates the typical beamformer into a fusion structure of two beamformers, thereby reducing the amount of computation and flexibly controlling the amount of interference suppression.
[0103] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0104] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for enhancing acoustic signals of a multi-channel microphone array, characterized in that: The enhancement method comprises: Obtaining sound source signals; Based on the sound source signal, an observation signal vector is determined; the observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector; the expression of the observation signal vector is as follows: y(f,t)=a(f,t)X(f,t)+b(f,t)ξ(f,t)+v(f,t) Where y(f, t) represents the observed signal vector, a(f, t) represents the steering vector of the source signal, b(f, t) represents the steering vector of the interference signal, X(f, t) represents the sound source signal, ξ(f, t) represents the interference signal, and v(f, t) represents the noise vector. The covariance matrix is determined based on the noise vector using the following formula: Φ v (f,t)=E[v(f,t)v H (f,t)] Among them, Φ v (f,t) represents the noise covariance matrix, v H (f, t) represents the conjugate transpose of the noise vector, and E represents the mathematical expectation; Determine the classical beamformer, the expression is: Where W(f) is the classical beamformer, Φ y -1 (f, t) represents the inverse matrix of the observed signal covariance matrix, a H (f, t) represents the conjugate transpose of the source signal steering vector; The noise covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal are input into the classical beamformer to obtain two beamformers to be fused. The expressions of the two beamformers to be fused are as follows: Among them, W a (f, t) represents the beamformer pointing to the sound source, W b (f,t) represents the beamformer pointing towards the interference, Φ v -1 (f, t) represents the inverse matrix of the noise covariance matrix, b H (f, t) represents the conjugate transpose of the interference signal steering vector; A preliminary estimate of the source signal and a preliminary estimate of the interference signal are determined based on the two beamformers to be fused and the observed signal vector; the expression of the preliminary estimate of the source signal is as follows: Among them, z a (f,t) represents the initial estimate of the source signal, represents the conjugate transpose of the beamformer pointing toward the sound source; The expression of the preliminary estimate of the interference signal is as follows: Among them, z b (f,t) represents the preliminary estimate of the interference signal, represents the conjugate transpose of the beamformer pointing toward the interference; To determine the key parameters, use the following formula: Among them, β represents the interference suppression ratio, k(f,t) is a parameter constructed by the sound source, the interference guidance amount and the noise covariance matrix, ρ(f,t) represents the spatial correlation coefficient, is a parameter determined by β and the signal, c1(f,t) represents the combination coefficient of the beamformer pointing to the sound source, and c2(f,t) represents the combination coefficient of the beamformer pointing to the interference; An enhanced signal is determined based on the key parameters, the preliminary estimate of the source signal, and the preliminary estimate of the interference signal. The expression of the enhanced signal is as follows: Where c1(f,t) represents the combination coefficient of the beamformer pointing to the sound source, represents the conjugate of c2(f,t), Z a (f,t) and Z b (f, t) represent the preliminary estimates of the source signal and interference, respectively, and * represents conjugation.
2. The multi-channel microphone array acoustic signal enhancement method according to claim 1, characterized in that: The acquiring of the sound source signal specifically involves observing the sound source signal through a microphone array.
3. A multi-channel microphone array sound signal enhancement system, characterized in that: The enhancement system comprises: A sound source signal acquisition module, used to acquire a sound source signal; An observation signal vector determination module is configured to determine an observation signal vector based on the sound source signal; the observation signal vector includes: a steering vector of the source signal, a steering vector of the interference signal, and a noise vector; the expression of the observation signal vector is as follows: y(f,t)=a(f,t)X(f,t)+b(f,t)ξ(f,t)+v(f,t) Where y(f, t) represents the observed signal vector, a(f, t) represents the steering vector of the source signal, b(f, t) represents the steering vector of the interference signal, X(f, t) represents the sound source signal, ξ(f, t) represents the interference signal, and v(f, t) represents the noise vector. A covariance matrix determination module is used to determine the covariance matrix based on the noise vector, using the following formula: Φ v (f,t)=E[v(f,t)v H (f,t)] Among them, Φ v (f,t) represents the noise covariance matrix, v H (f, t) represents the conjugate transpose of the noise vector, and E represents the mathematical expectation; The classical beamformer determination module is used to determine the classical beamformer, and the expression is: Where W(f) is the classical beamformer, Φ y -1 (f, t) represents the inverse matrix of the observed signal covariance matrix, a H (f, t) represents the conjugate transpose of the source signal steering vector; The beamformer determination module to be fused is configured to input the noise covariance matrix, the steering vector of the source signal, and the steering vector of the interference signal into the classical beamformer to obtain two beamformers to be fused. The expressions of the two beamformers to be fused are as follows: Among them, W a (f, t) represents the beamformer pointing to the sound source, W b (f,t) represents the beamformer pointing towards the interference, Φ v -1 (f, t) represents the inverse matrix of the noise covariance matrix, b H (f, t) represents the conjugate transpose of the interference signal steering vector; A module for determining a preliminary estimate of the source signal and a preliminary estimate of the interference signal is configured to determine a preliminary estimate of the source signal and a preliminary estimate of the interference signal based on the two beamformers to be fused and the observed signal vector; the expression for the preliminary estimate of the source signal is as follows: Among them, z a (f,t) represents the initial estimate of the source signal, represents the conjugate transpose of the beamformer pointing toward the sound source; The expression of the preliminary estimate of the interference signal is as follows: Among them, z b (f,t) represents the preliminary estimate of the interference signal, represents the conjugate transpose of the beamformer pointing toward the interference; The key parameter determination module is used to determine the key parameters using the following formula: Among them, β represents the interference suppression ratio, k(f,t) is a parameter constructed by the sound source, the interference guidance amount and the noise covariance matrix, ρ(f,t) represents the spatial correlation coefficient, is a parameter determined by β and the signal, c1(f,t) represents the combination coefficient of the beamformer pointing to the sound source, and c2(f,t) represents the combination coefficient of the beamformer pointing to the interference; The enhanced signal determination module is configured to determine an enhanced signal based on the key parameters, a preliminary estimate of the source signal, and a preliminary estimate of the interference signal. The expression of the enhanced signal is as follows: Where c1(f,t) represents the combination coefficient of the beamformer pointing to the sound source, represents the conjugate of c2(f,t), Z a (f,t) and Z b (f, t) represent the preliminary estimates of the source signal and interference, respectively, and * represents conjugation.
Citation Information
Patent Citations
Neural network based time-frequency mask estimation and beamforming for speech pre-processing
CN110503971A
Steering vector estimation method and system
CN114609587A