A sound source localization method based on distributed array
By combining distributed and centralized microphone arrays and optimizing beamforming using the CLEAN-SC algorithm, the problem of limited positioning accuracy of distributed arrays in large environments is solved, achieving higher sound source positioning accuracy and resolution.
Patent Information
- Application Number
- CN202411712660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-27
AI Technical Summary
In existing technologies, the spatial resolution of a single microphone array is limited by its physical size, making it difficult to apply in large environments. While distributed microphone arrays have advantages, they face problems such as variable structure, different gains, and asynchronous recording, which affect positioning accuracy.
By combining distributed and centralized arrays, and through preprocessing, the CLEAN-SC algorithm, and beamforming technology, the horizontal and vertical positions of the sound source are determined respectively. The CLEAN-SC algorithm is used to optimize beamforming and improve positioning accuracy.
It achieves more accurate sound source localization in large environments, overcomes the challenges of distributed arrays, and improves the accuracy and resolution of localization, especially in low-frequency signals.
Smart Images

Figure CN119716736B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization, and in particular to a sound source localization method based on a distributed microphone array. Background Technology
[0002] When a single microphone acquires sound data, the spatial distribution characteristics of interference and noise mean that the movement of the sound source signal can cause changes in the spatial location of the signal source, leading to a decrease in sound signal pickup quality. In contrast to a single microphone, a microphone array acquires spatial information about the sound source signal by utilizing the small time differences between the multi-channel data collected by multiple microphones. While current mainstream sound source localization technologies, such as beamforming, can effectively target the main sound source, the spatial resolution of a single array is inevitably limited by its physical size because beamforming relies on pre-designed array distances to calculate the small time delays or frequency domain phase shifts between multi-channel data. In beamforming, to maintain excellent spatial resolution over a larger target area, the array size must be large enough, inevitably leading to expensive and bulky arrays, significantly limiting practical applications. Therefore, in large environments, multiple microphone arrays, i.e., distributed microphone arrays, are needed to cover the entire relevant space.
[0003] In distributed microphone arrays, the microphone positions are no longer limited to traditional fixed geometries, providing a wide and flexible spatial coverage. Multiple arrays overlap in sound source localization, and in these areas, integrating multiple arrays yields more accurate localization results than a single array. When the sound source is far from the microphone array, the array's localization accuracy decreases due to sound wave attenuation. Therefore, when two arrays have the same geometry and background signal-to-noise ratio, the closer array will provide a more accurate location estimate and have a larger weighting than the farther array. Compared to a single centralized array, distributed arrays offer advantages in acoustic coverage and are currently widely used in fields such as volcanic activity monitoring and earthquake early warning. However, they also present some challenges, including variable array structures, varying gain due to different distances from the sound source to the array, and asynchronous recording. Summary of the Invention
[0004] This invention proposes a method for three-dimensional spatial localization of a sound source, which combines distributed array and centralized array. The distributed array algorithm determines the horizontal position information of the sound source, and the centralized array algorithm determines the vertical position information of the sound source.
[0005] The technical solution of this invention is a sound source localization method based on a distributed array: the method includes:
[0006] Step 1: Arrange multiple microphone sensors into a single microphone array, and place multiple microphone arrays in space;
[0007] Step 2: Use multiple microphone arrays to collect the total sound signal in the space, and preprocess the collected sound data signal, including: (1) sampling filtering to ensure correct reconstruction of the sound signal; (2) audio validity detection to detect valid endpoints of the received data and remove invalid sound data to ensure the real-time performance of the algorithm; (3) noise reduction to perform spectral subtraction noise reduction on the collected data; (4) windowing and framing to decompose the long audio signal into audio frames with a duration of milliseconds for smooth approximation processing.
[0008] Step 3: Perform large-scale sound source localization to estimate the horizontal position of the sound source;
[0009] Step 4: Determine the vertical position of the sound source; determine whether the frequency of the collected sound signal is lower than the set threshold. If yes, proceed to step 5; otherwise, output the position directly.
[0010] Step 5: Use the CLEAN-SC algorithm for processing and localization to obtain the accurate location of the sound source.
[0011] Furthermore, the specific method for step 3 is as follows:
[0012] Step 3.1: The signal from the i-th sound source is used Indicates that t represents time. This indicates that at position l Position in space is represented by a three-dimensional vector, s i It is a function of time and position, representing the sound signal of the i-th sound source at position l at time t; the signal received by the m-th microphone array is represented as:
[0013]
[0014] in, Indicates sound signal from arrive Impulse response of the propagation path, This represents the location of the m-th microphone array, which is represented in space by a three-dimensional vector. n m (t) represents all noise sources;
[0015] Step 3.2: For x m Perform a Fourier transform on (t) to obtain
[0016]
[0017] in, This represents the Fourier transform of the i-th sound source signal. Indicates from arrive Fourier transform of the impulse response of the propagation path ω represents the Fourier transform of the ambient noise signal at the m-th microphone array, where ω represents the frequency.
[0018] Step 3.3: For Perform weighted processing to obtain the weighted result.
[0019]
[0020] in, Let represent the definition, where β is a parameter controlling the degree of spectral whitening, taking values [0,1], |.| β β represents the norm 2 raised to the power of β; weighting is a process of deconvolving the spectrum to make the contribution of each frequency region to the coherent sum of steering power more uniform, and the β parameter limits the enhancement of the low-level spectrum region dominated by noise.
[0021] Step 3.4: Represent the weighted result from Step 3.3 in the frequency domain as follows:
[0022]
[0023] Among them, S l This represents the estimated total coherence at estimation point l, where M indicates that the distributed array has a total of M microphone subarrays. This represents the weighted result of the Fourier transform of the signal received by the m-th microphone, divided by its own modulus raised to the power of beta. q represents the sequence number of the microphone array, and m also represents the sequence number of the microphone array. Assume there are a total of M microphone arrays, denoted as array 1, 2, ..., m...M. This calculates the correlation between signals received from any two microphone arrays.
[0024] Furthermore, the specific method for step 4 is as follows:
[0025] Step 4.1: For a microphone array, there are Q microphone elements, each microphone is represented by a sub-index q, q∈{1,…,Q}, and each microphone outputs a discrete-time signal s. q (n), at the point in space x = [x, y, z] T The steering response power P0(x) is expressed as:
[0026]
[0027] Where Z represents s q(n) is the set of all points. This macro equals, meaning that P0(x) is equal to its definition. Similar to τ q (x) represents the time delay of the signal from sound source x to the qth microphone;
[0028] Step 4.2: After weighting P0(x), we obtain the weighted result P(x):
[0029]
[0030] in, It is a frequency domain weighting function. For each microphone array, there are a total of Q microphones. This indicates that the q1th microphone receives the signal s. q1 The Fourier transform of (n), also S q2 (e jω ) represents the discrete signal s received by the q2th microphone. q2 The Fourier transform of (n), where the superscript "*" indicates the conjugate transpose. The definition of is:
[0031]
[0032] Among them, f s The system's sampling frequency is represented by v, the speed of sound is x. q The position of the q-th microphone is represented by ||·|, where || represents the absolute value and "" represents the rounding operator.
[0033] in,
[0034] Estimate P(x) at each candidate sound source grid point G, and the grid point corresponding to the largest value is the spatial location of the sound source. Right now:
[0035]
[0036] Furthermore, the specific method for step 5 is as follows:
[0037] Step 5.1: Scan the grid points of the sound source plane get The power of the sound source at that location for:
[0038]
[0039] in, The cross-power spectrum matrix, This represents a vector, calculated by taking the guide vector as a unit vector, assuming that for the focal point... The guide vector is The vector is The superscript "T" indicates transpose; find the location of the global maximum value from it. Subtract from the global value To eliminate the influence of peak values and obtain sound power unaffected by peak values.
[0040]
[0041] in, It is by The cross-spectral matrix induced by the sound source at a given location; requiring arbitrary scan points Peak position The acoustic mutual power must be determined by Decision, i.e., arbitrary exist:
[0042]
[0043] in, yes The weighted vector, This represents the steering vector with the peak value as the sound source in the i-th iteration. Let represent the cross-spectral matrix of the microphone sound pressure signal obtained in the (i-1)th iteration. CLEAN-SC is a multi-iteration algorithm, and the superscript indicates the iteration number; when satisfying And assume It consists of a single coherent sound source component The resulting solution can be uniquely obtained:
[0044]
[0045] in, This represents the maximum sound intensity value of the scanned region in the (i-1)th iteration;
[0046]
[0047] (m,n) where m and n are the indices of the microphones, and S represents all possible combinations of microphones. This represents the sound pressure distribution calculated from the m-th microphone. This represents the sound pressure distribution calculated at the nth microphone; Contains diagonal elements Therefore, iterative output is required. at this time,
[0048]
[0049] This indicates that the matrix resulting from multiplying the source component by its transpose is the conjugate complex number.
[0050] Step 5.2: Update the sound source map obtained in Step 4 after removing peak positions, and replace coherent sound sources with clean beams:
[0051]
[0052] in, This represents the maximum acoustic intensity value of the scanned region in the (i-1)th iteration, λ is a parameter that determines the bandwidth, and the safety factor is... The value ranges from 0 to 1, and the cross-spectral matrix becomes:
[0053]
[0054] The remaining acoustic diagrams are:
[0055]
[0056] Continue iterating until:
[0057]
[0058] The final sound source intensity distribution map A j It is the superposition of the clean beam and the intensity distribution map of the remaining sound sources, that is:
[0059]
[0060] According to sound source intensity distribution map A j The location of the sound source was obtained.
[0061] Compared to other sound source localization methods, the method of this invention first performs coarse localization and then fine localization, and uses the CLEAN-SC algorithm to optimize the beam, making full use of its low computational complexity and high resolution, thus achieving a sound source localization method with superior performance. Attached Figure Description
[0062] Figure 1 This is a flowchart of the sound source localization method.
[0063] Figure 2 This is a schematic diagram of the location of a distributed array sound source.
[0064] Figure 3 This is the sound field effect diagram obtained from the algorithm in step 3.
[0065] Figure 4 The beammaps are obtained by processing the same signal using different algorithms.
[0066] Figure 5 This is the sound field effect diagram obtained from the algorithm in step 4.
[0067] Figure 6 This is a diagram of the sound field effect obtained by the algorithm in step 4 under low-frequency signals.
[0068] Figure 7 This is a sound field effect diagram obtained by the CLEAN-SC algorithm under low-frequency signals. Detailed Implementation
[0069] The specific implementation process of this sound source localization method is as follows: Figure 1 Display. Step 1: Setting up the acquisition device. During the audio acquisition stage, a distributed microphone array needs to be arranged spatially, such as... Figure 2 As shown; Step 2 acquires audio and preprocesses it; Step 3 uses the algorithm to determine the approximate horizontal position of the preprocessed sound signal; Step 4 uses the algorithm to further locate the vertical position of the sound signal. If the audio frequency is low, Step 4 may not yield better results, and the improved CLEAN-SC algorithm needs to be used for processing and localization to finally obtain the accurate location of the sound source.
[0070] Step 3 is an algorithm for large-scale sound source localization using distributed arrays. It requires multiple microphone arrays distributed at different locations in space to collect sound signals within the space. The signal from the i-th sound source is used... The signal received by the m-th microphone element is represented as:
[0071]
[0072] in, Indicates sound signal from arrive Impulse response of the propagation path, n m (t) represents all noise sources. A Fourier transform of the above equation is performed:
[0073]
[0074] right PHAT-β weighting is performed, defined as follows:
[0075]
[0076] Here, β is a parameter controlling the degree of spectral whitening, with a value of [0,1]. PHAT-β mainly deconvolves the spectrum, making the contribution of each frequency region to the coherent sum of steering power more uniform. The β parameter limits the enhancement of the low-level spectrum region, which is dominated by noise. Thus, the weighted SRCP relative to position l can be expressed in the frequency domain as:
[0077]
[0078] Scan all grid points in the observation area and create an SRCP image for each grid point using the above formula. For example... Figure 3 As shown.
[0079] After obtaining the approximate horizontal position of the sound source signal using the algorithm in step 3, a small-scale localization can be performed. Using a single microphone array, the algorithm in step 4 can be used to precisely locate it.
[0080] Step 4 is a beamforming algorithm that searches for candidate locations or directions that maximize the directional delay and beamformer output. For a microphone array with Q microphones, each denoted by a sub-index q, q∈{1,…,Q}, each microphone has a discrete-time output signal s. m (n), at the point in space x = [x, y, z] T The SRP response can be represented as:
[0081]
[0082] Where Z represents s q (n) is the set of all points, τ q (x) represents the time delay of the signal from sound source x to the q-th microphone.
[0083] After weighting, we get:
[0084]
[0085] in, It is a frequency domain weighting function. The definition of is:
[0086]
[0087] Among them, f s The system's sampling frequency is represented by v, the speed of sound is x. q The position of the q-th microphone is represented by ||·||, where || represents the absolute value and "" represents the rounding operator.
[0088] if It is PHAT weighted, at this time
[0089]
[0090] PHAT weighting normalizes the amplitude, retaining only phase information. Step 4: The algorithm searches the spatial grid of the sound source, estimating P(x) at each candidate sound source grid point G. The grid point corresponding to the largest SRP value is the spatial location x of the sound source.s ,Right now:
[0091]
[0092] Compared to the MVDR and DSB algorithms, the algorithm in step 4 has better directional accuracy, is less affected by noise, and has higher reliability. The comparison results are as follows: Figure 4 The sound source effect diagram obtained by the algorithm in step 4 is as follows: Figure 5 .
[0093] Since the sound source may emit low-frequency audio, the localization result obtained by the algorithm in step 4 will have a lower resolution. Therefore, the CLEAN-SC algorithm is used for low-frequency signals to obtain higher-resolution localization. Figure 6 Figure 7 As shown.
[0094] The CLEAN-SC algorithm first scans the plane grid points of the sound source using traditional beamforming techniques. get The power of the sound source at that location for:
[0095]
[0096] in, Given the cross-power spectrum matrix, find the location of the global maximum value. Subtract from the global value To eliminate the influence of peak values and obtain sound power unaffected by peak values.
[0097]
[0098] in, It is by The cross-spectral matrix induced by the sound source at a given location. The CLEAN-SC algorithm requires arbitrary scan points... Peak position The acoustic mutual power must be determined by Decision, i.e., arbitrary exist:
[0099]
[0100] in, Is with The relevant weighted vector. When the following conditions are met... And assume It consists of a single coherent sound source component The resulting solution can be uniquely obtained:
[0101]
[0102] in,
[0103]
[0104] Contains diagonal elements Therefore, iterative output is required. at this time,
[0105]
[0106] The original source map was updated after peak positions were removed, and coherent sound sources were replaced by clean beams:
[0107]
[0108] Where λ is a parameter that determines the bandwidth, and the security factor is... The value ranges from 0 to 1, and the cross-spectral matrix becomes:
[0109]
[0110] The remaining acoustic diagrams are
[0111]
[0112] Continue iterating until:
[0113]
[0114] The final sound source intensity distribution map is the superposition of the clean beam and the residual sound source intensity distribution maps, that is:
[0115]
[0116] The resulting sound source intensity distribution map has strong directivity and high accuracy.
Claims
1. A sound source localization method based on a distributed array, the method comprising: Step 1: Arrange multiple microphone sensors into a single microphone array, and place multiple microphone arrays in space; Step 2: Use multiple microphone arrays to collect the total sound signal in the space, and preprocess the collected sound data signal, including: (1) sampling filtering to ensure correct reconstruction of the sound signal; (2) audio validity detection to detect valid endpoints of the received data and remove invalid sound data to ensure the real-time performance of the algorithm; (3) noise reduction to perform spectral subtraction noise reduction on the collected data; (4) windowing and framing to decompose the long audio signal into audio frames with a duration of milliseconds for smooth approximation processing. Step 3: Perform large-scale sound source localization to estimate the horizontal position of the sound source; Step 4: Determine the vertical position of the sound source; check if the frequency of the collected sound signal is lower than the set threshold. If yes, proceed to Step 5; otherwise, output the position directly. Step 4.1: For a microphone array, there are Q microphone elements, each microphone is represented by a sub-index q, q∈{1,…,Q}, and each microphone outputs a discrete-time signal s. q (n), at the point in space x = [x, y, z] T The steering response power P0(x) is expressed as: Where Z represents s q (n) is the set of all points. This macro equals, meaning that P0(x) is equal to its definition. τ q (x) represents the time delay of the signal from sound source x to the qth microphone; Step 4.2: After weighting P0(x), we obtain the weighted result P(x): in, It is a frequency domain weighting function. For each microphone array, there are a total of Q microphones. For Sq1(e jω ) indicates that the q1th microphone receives the signal s. q1 The Fourier transform of (n), also S q2 (e jω ) represents the discrete signal s received by the q2th microphone. q2 The Fourier transform of (n), where the superscript "*" indicates the conjugate transpose. The definition of is: Among them, f s The system's sampling frequency is represented by v, the speed of sound is x. q This represents the position of the q-th microphone. ‖·‖ represents the absolute value. This represents the rounding operator; in, Estimate P(x) at each candidate sound source grid point G, and the grid point corresponding to the largest value is the spatial location of the sound source. Right now: Step 5: Use the CLEAN-SC algorithm for processing and localization to obtain the accurate location of the sound source.
2. The sound source localization method based on a distributed array as described in claim 1, characterized in that, The specific method for step 3 is as follows: Step 3.1: The signal from the i-th sound source is used Indicates that t represents time. This indicates that at position l Position in space is represented by a three-dimensional vector, s i It is a function of time and position, representing the sound signal of the i-th sound source at position l at time t; the signal received by the m-th microphone array is represented as: in, Indicates sound signal from arrive Impulse response of the propagation path, This represents the location of the m-th microphone array, which is represented in space by a three-dimensional vector. n m (t) represents all noise sources; Step 3.2: For x m Perform a Fourier transform on (t) to obtain in, This represents the Fourier transform of the i-th sound source signal. Indicates from arrive Fourier transform of the impulse response of the propagation path, ω represents the Fourier transform of the ambient noise signal at the m-th microphone array, where ω represents the frequency. Step 3.3: For Perform weighted processing to obtain the weighted result. in, Let represent the definition, where β is a parameter controlling the degree of spectral whitening, taking values [0,1], |.| β β represents the norm 2 raised to the power of β; weighting is a process of deconvolving the spectrum to make the contribution of each frequency region to the coherent sum of steering power more uniform, and the β parameter limits the enhancement of the low-level spectrum region dominated by noise. Step 3.4: Represent the weighted result from Step 3.3 in the frequency domain as follows: Among them, S l This represents the estimated total coherence at estimation point l, where M indicates that the distributed array has a total of M microphone subarrays. This represents the weighted result of the Fourier transform of the signal received by the m-th microphone divided by its own modulus raised to the power of beta. q represents the sequence number of the microphone array, and m also represents the sequence number of the microphone array. Assuming there are a total of M microphone arrays, denoted as microphone array 1, 2, ..., m, ... M; here, the correlation between the signals received from any two microphone arrays is calculated.
3. The sound source localization method based on a distributed array as described in claim 1, characterized in that, The specific method for step 5 is as follows: Step 5.1: Scan the grid points of the sound source plane get The power of the sound source at that location for: in, The cross-power spectrum matrix, This represents a vector, calculated by taking the guide vector as a unit vector, assuming that for the focal point... The guide vector is The vector is The superscript "T" indicates transpose; find the location of the global maximum value from it. Subtract from the global value To eliminate the influence of peak values and obtain sound power unaffected by peak values. in, It is by The cross-spectral matrix induced by the sound source at a given location; requiring arbitrary scan points With peak position The acoustic mutual power must be determined by Decision, i.e., arbitrary exist: in, yes The weighted vector, This represents the steering vector with the peak value as the sound source in the i-th iteration. Let represent the cross-spectral matrix of the microphone sound pressure signal obtained in the (i-1)th iteration. CLEAN-SC is a multi-iteration algorithm, and the superscript indicates the iteration number; when satisfying And assume It consists of a single coherent sound source component The resulting solution can be uniquely obtained: in, This represents the maximum sound intensity value of the scanned region in the (i-1)th iteration; (m,n) where m and n are the indices of the microphones, and S represents all possible combinations of microphones. This represents the sound pressure distribution calculated from the m-th microphone. This represents the sound pressure distribution calculated at the nth microphone; Contains diagonal elements Therefore, iterative output is required. at this time, This indicates that the matrix resulting from multiplying the source component by its transpose is the conjugate complex number. Step 5.2: Update the sound source map obtained in Step 4 after removing peak positions, and replace coherent sound sources with clean beams: in, This represents the maximum acoustic intensity value of the scanned region in the (i-1)th iteration, λ is a parameter that determines the bandwidth, and the safety factor is... The value ranges from 0 to 1, and the cross-spectral matrix becomes: The remaining acoustic diagrams are: Continue iterating until: The final sound source intensity distribution map A j It is the superposition of the clean beam and the intensity distribution map of the remaining sound sources, that is: According to sound source intensity distribution map A j The location of the sound source was obtained.
Citation Information
Patent Citations
Low-frequency beam forming sound source positioning method based on spherical microphone array
CN114527427A
Continuous mobile microphone array sound source positioning method and device
CN118191732A