Sound source positioning method and system based on multi-microphone array
By employing a multi-microphone array-based sound source localization method, utilizing time-frequency decomposition and acoustic coherence analysis, and combining sound wave multipath and diffraction characteristics, accurate localization of multiple sound sources in complex acoustic environments is achieved, solving the problem of insufficient localization accuracy in traditional methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HESHENGCHENG TECH CO LTD
- Filing Date
- 2026-02-26
- Publication Date
- 2026-04-24
AI Technical Summary
In complex acoustic environments, traditional methods struggle to accurately distinguish between multiple sound sources emitting sound simultaneously, resulting in insufficient localization accuracy.
A sound source localization method based on a multi-microphone array is adopted. The signal collected by the microphone array is preprocessed, time-frequency decomposition and acoustic coherence analysis are performed, a time-frequency coherence matrix is constructed, the time-frequency region of a single source is screened by using the multipath and diffraction characteristics of sound waves, and the direction of arrival information is generated by combining frequency domain image signal processing and cluster weighted fusion.
Precise sound source separation and localization were achieved in a multi-source aliasing environment, improving the accuracy of multi-source localization.
Smart Images

Figure CN121918063A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound source localization technology, and in particular to a sound source localization method and system based on a multi-microphone array. Background Technology
[0002] In the field of sound source localization, insufficient localization accuracy is a common problem when locating multiple sound sources emitting sound simultaneously. Especially in complex acoustic environments, due to the multipath propagation and diffraction effects of sound waves, signals from different sound sources alias in the time and frequency domains, making it difficult for traditional methods based on beamforming or time delay estimation to accurately distinguish the directions of each sound source. Summary of the Invention
[0003] This application provides a sound source localization method and system based on a multi-microphone array, which enables accurate sound source separation and localization in a multi-source aliasing environment.
[0004] In a first aspect, embodiments of this application provide a sound source localization method based on a multi-microphone array, the method comprising: The first sound source signal acquired by the microphone array is preprocessed, and the second sound source signal and acoustic physical parameters are output. The second sound source signal is decomposed into multiple first time-frequency regions by time-frequency decomposition. Acoustic coherence analysis is performed on the first time-frequency region based on the acoustic physical parameters to obtain the time-frequency coherence matrix; Based on the multipath and diffraction characteristics of acoustic waves, a second time-frequency region dominated by a single source is detected from the time-frequency coherence matrix, and an output single-source time-frequency point index is constructed based on the second time-frequency region. Multiple direction-of-arrival (DOA) estimates are generated based on the preset frequency domain diagram signal and the single-source time-frequency point index. Clustering and weighted fusion of the multiple direction-of-arrival (DOA) estimates are performed to obtain the DOA information of the sound source.
[0005] Secondly, embodiments of this application provide a sound source localization system based on a multi-microphone array. The sound source localization system based on a multi-microphone array is used to execute any of the sound source localization methods based on a multi-microphone array described in the embodiments of this application. The sound source localization system based on a multi-microphone array includes: The signal processing module is used to preprocess the first sound source signal acquired by the microphone array and output the second sound source signal and acoustic physical parameters. The time-frequency decomposition module is used to perform time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions; An acoustic analysis module is used to perform acoustic coherence analysis on the first time-frequency region based on the acoustic physical parameters to obtain a time-frequency coherence matrix; A single-source determination module is used to detect a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on acoustic multipath and diffraction characteristics, and to construct an output single-source time-frequency point index based on the second time-frequency region. The direction estimation module is used to generate multiple direction-of-arrival estimates based on a preset frequency domain diagram signal and the single-source time-frequency point index. The preset frequency domain diagram signal is obtained by processing the acoustic physical parameters using a preset enhancement model. The result output module is used to cluster and weightedly fuse the multiple direction-of-arrival estimates to obtain the direction-of-arrival information of the sound source.
[0006] This application provides a sound source localization method based on a multi-microphone array. The method includes: preprocessing a first sound source signal acquired by the microphone array to output a second sound source signal and acoustic physical parameters; performing time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions; performing acoustic coherence analysis on the first time-frequency regions based on the acoustic physical parameters to obtain a time-frequency coherence matrix; detecting a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on sound wave multipath and diffraction characteristics, and constructing an output single-source time-frequency point index based on the second time-frequency region; generating multiple direction-of-arrival (DOA) estimates based on a preset frequency domain map signal and the single-source DOA index; and clustering and weighted fusion of the multiple DOA estimates to obtain the DOA information of the sound source. In the above method, sound source separation in a multi-source aliasing environment is effectively achieved by transferring acoustic physical parameters to the time-frequency analysis stage and constructing a time-frequency coherence matrix. Then, based on the sound wave multipath and diffraction characteristics, the time-frequency region of a single source is screened. Combined with frequency domain map signal processing enhanced by physical model and cluster weighted fusion, the accuracy of multi-source localization is improved. Attached Figure Description
[0007] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 A schematic flowchart illustrating a sound source localization method based on a multi-microphone array, provided for an embodiment of this application; Figure 2 This is a schematic block diagram of a sound source localization system based on a multi-microphone array, provided for an embodiment of this application. Detailed Implementation
[0009] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described below with reference to the accompanying drawings.
[0010] The terms "first" and "second," etc., used in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0011] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0012] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0013] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating a sound source localization method based on a multi-microphone array, provided in an embodiment of this application. Figure 1 As shown, the specific steps of this sound source localization method based on a multi-microphone array include: S101-S106.
[0014] S101. Preprocess the first sound source signal acquired by the microphone array and output the second sound source signal and acoustic physical parameters.
[0015] For example, by calculating the phase difference of the received signals from the six array elements and performing calibration compensation based on the sound velocity parameters under standard atmospheric conditions, a generalized cross-correlation algorithm is used to calculate the time delay characteristics between each array element and establish a multipath propagation model. An adaptive filter based on the minimum mean square error criterion is applied to suppress multipath effects on the first sound source signal, outputting a second sound source signal with a significantly improved signal-to-noise ratio, along with acoustic physical parameters including a sound velocity correction factor and a multipath attenuation coefficient. This preprocessing is performed under environmental conditions of 20-25 degrees Celsius and 40-60% relative humidity, achieving a sound velocity calibration accuracy on the order of 0.5 m / s. The multipath attenuation coefficient calculation is based on the analysis of 3-5 main propagation paths, and the signal filtering reduces reflected sound wave interference by 6%. -10 dB.
[0016] S102. Perform time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions.
[0017] For example, the time-frequency decomposition of the second sound source signal adopts the windowed short-time Fourier transform method. Based on the sound velocity correction factor in the acoustic physical parameters, the window length of the Hanning window is dynamically adjusted to 256-512 sampling points to maintain a specific proportional relationship with the sound wave wavelength of the main frequency band of the speech signal. A frame shift parameter of 50%-75% is set to ensure the continuity of time-frequency analysis. The second sound source signal with a duration of 2-5 seconds is decomposed into 150-300 time-frequency regions. Each time-frequency region contains 32-64 frequency components and has a time resolution of 10-20ms and a frequency resolution of 30-60Hz. These time-frequency regions completely preserve the energy distribution characteristics of the original signal in the time and frequency domains.
[0018] S103. Perform acoustic coherence analysis on the first time-frequency region based on acoustic physical parameters to obtain the time-frequency coherence matrix.
[0019] For example, the cross-correlation coefficients among the six array element signals in each time-frequency region are calculated. The cross-correlation coefficients are then weighted and corrected based on the multipath attenuation coefficient. The principal components of the cross-correlation coefficient matrix are extracted using the eigenvalue decomposition method to form a six-dimensional coherence vector. The coherence vector is then linearly fused with the energy values of each time-frequency region to construct a time-frequency coherence matrix. The diagonal elements of this matrix represent the self-coherence coefficients, while the non-diagonal elements represent the cross-coherence coefficients. This processing method can fully characterize the spatial correlation features of the signal in the time-frequency domain.
[0020] S104. Based on the multipath and diffraction characteristics of acoustic waves, detect the second time-frequency region dominated by a single source from the time-frequency coherence matrix, and construct the output single-source time-frequency point index based on the second time-frequency region.
[0021] For example, by calculating the delay spread and angle spread values at each time and frequency point, a sound wave attenuation model is established by applying Kirchhoff diffraction theory. Multipath effect threshold and diffraction attenuation threshold are set to filter the time and frequency regions. Second time and frequency regions that meet specific conditions are selected from the original time and frequency regions. The time and frequency coordinates of these regions are extracted to form a binary index sequence. The single-source regions marked by this index sequence have high signal-to-noise ratio characteristics.
[0022] S105. Generate multiple direction-of-arrival estimates based on the preset frequency domain diagram signal and single-source time-frequency point index.
[0023] For example, the preset frequency domain map signal is obtained by processing acoustic physical parameters using a preset enhancement model. The process of generating direction-of-arrival (DOA) estimates based on the preset frequency domain map signal and single-source time-frequency point indices involves constructing a frequency domain map signal processing model based on acoustic physical parameters. The weighting coefficients of the adjacency matrix are adjusted according to the acoustic wavefront curvature. The corresponding frequency domain signal components are extracted from the single-source time-frequency point indices, and the eigenvalues and eigenvectors of the graph Laplacian matrix are calculated. A spatial spectrum estimation algorithm is used to calculate the spatial spectrum function with a fixed step size across the entire omnidirectional range, identifying significant spectral peaks and determining the corresponding DOA estimates.
[0024] S106. Cluster and weightedly fuse multiple direction-of-arrival estimates to obtain the direction-of-arrival information of the sound source.
[0025] For example, the process of clustering and weighted fusion of multiple direction-of-arrival (DOA) estimates uses a density clustering algorithm to divide the estimates into several clusters, calculates the sound intensity attenuation factor and physical confidence level of the estimates in each cluster, and uses the product result as the fusion weight to perform a weighted average calculation on each cluster. The weight variance is controlled within a specific range to ensure fusion stability. The output DOA information can achieve complete resolution when the distance between sound sources is greater than a specific angle, meeting the positioning requirements in complex acoustic environments.
[0026] This application provides a sound source localization method based on a multi-microphone array. The method includes: preprocessing a first sound source signal acquired by the microphone array to output a second sound source signal and acoustic physical parameters; performing time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions; performing acoustic coherence analysis on the first time-frequency regions based on the acoustic physical parameters to obtain a time-frequency coherence matrix; detecting a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on sound wave multipath and diffraction characteristics, and constructing an output single-source time-frequency point index based on the second time-frequency region; generating multiple direction-of-arrival (DOA) estimates based on a preset frequency domain map signal and the single-source DOA index; and clustering and weighted fusion of the multiple DOA estimates to obtain the DOA information of the sound source. In the above method, sound source separation in a multi-source aliasing environment is effectively achieved by transferring acoustic physical parameters to the time-frequency analysis stage and constructing a time-frequency coherence matrix. Then, based on the sound wave multipath and diffraction characteristics, the time-frequency region of a single source is screened. Combined with frequency domain map signal processing enhanced by physical model and cluster weighted fusion, the accuracy of multi-source localization is improved.
[0027] To more clearly illustrate the technical solution of this application, the technical solution of this application will be described below through specific embodiments. It should be noted that the specific embodiments are used to expand the description of the technical solution of this application, and are not intended to limit this application.
[0028] In one embodiment, preprocessing the first sound source signal acquired by the microphone array to output a second sound source signal and acoustic physical parameters includes: calculating a first phase difference of the signal received by each array element based on the spatial position of each array element in the microphone array; calibrating and compensating the first phase difference to obtain a second phase difference; performing phase calibration on the first sound source signal based on the second phase difference to obtain a third sound source signal; performing multipath effect suppression on the third sound source signal to obtain a second sound source signal; and determining acoustic physical parameters based on the second phase difference and the second sound source signal, the acoustic physical parameters including a sound velocity correction factor and a multipath attenuation coefficient.
[0029] For example, the first phase difference is calculated based on the geometric distance between array elements and the propagation time difference of the arriving sound waves. The time offset of the signal wavefront is captured through accurate time delay estimation, yielding the phase difference at the corresponding sampling points of each array element. This first phase difference may deviate due to environmental factors and equipment characteristics; therefore, it is calibrated and compensated to obtain the second phase difference. The calibration and compensation uses the sound velocity parameter under standard atmospheric conditions as a benchmark, combined with actual measured environmental factors such as temperature and humidity, and corrects the phase by numerical interpolation or fitting methods to improve phase matching accuracy. Based on the calibrated second phase difference, phase correction processing is applied to the first sound source signal to eliminate phase distortion caused by propagation delay differences and multipath effects, forming the third sound source signal. This stage ensures phase consistency and signal synchronization through staged filter design or time-domain phase adjustment algorithms. Multipath effect suppression processing of the third sound source signal is a crucial preprocessing step. This process uses a generalized cross-correlation algorithm to identify the sound wave delay characteristics of multiple main transmission paths, establishes a multipath propagation model, and uses an adaptive filter designed with the minimum mean square error criterion to effectively suppress interference from multipath reflected and scattered sound waves, significantly improving the signal-to-noise ratio. The result of multipath suppression is the second sound source signal. Based on the second phase difference and the newly generated second sound source signal, combined with the propagation physical characteristics of sound waves, environmental parameters, and array geometry, acoustic physical parameters are determined, including a sound velocity correction factor, used to characterize the actual influence of different temperature and humidity conditions on sound velocity in the environment, and a multipath attenuation coefficient, used to describe the attenuation ratio of sound wave energy on multiple propagation paths.
[0030] In one embodiment, time-frequency decomposition of the second sound source signal to obtain multiple first time-frequency regions includes: dynamically adjusting the window length of a preset window function based on a sound velocity correction factor, wherein the window length is proportional to the wavelength of the sound wave, and the preset window function includes: windowed short-time Fourier transform and Hanning window; loading the overlap rate of the time-frequency decomposition, wherein the overlap rate of the time-frequency decomposition is 50%-75%; and performing time-frequency decomposition of the second sound source signal according to the adjusted window length and the overlap rate of the time-frequency decomposition to obtain multiple first time-frequency regions.
[0031] For example, based on the spatial relationship of each element in the microphone array, the first phase difference of the received signal of each element is obtained by calculating the difference in the propagation path of the sound wave to different elements. This calculation process needs to consider the geometric configuration of the array and the physical characteristics of sound wave propagation. When calibrating and compensating for the first phase difference, it is necessary to combine environmental atmospheric condition parameters and use acoustic propagation theory to correct the initially calculated phase difference to eliminate measurement errors caused by environmental factors, thereby obtaining a more accurate second phase difference. During the phase calibration of the first sound source signal based on the second phase difference, advanced signal processing algorithms are used to time-align and phase-synchronize the signals received by each element, so that the signals of different elements can be processed on the same time reference. The third sound source signal generated by this step has better consistency characteristics. When performing multipath suppression of the third sound source signal, it is necessary to establish an accurate multipath propagation model, analyze the reflection and scattering characteristics of sound waves in the environment, and use specific filtering techniques to weaken the impact of multipath interference on signal quality. This processing can significantly improve the purity of the signal. The process of determining acoustic physical parameters based on the second phase difference and the second sound source signal requires the comprehensive application of multiple signal analysis methods. Through statistical processing and feature extraction, a set of parameters that can accurately describe the characteristics of the acoustic environment is obtained. These parameters provide important physical basis for subsequent sound source localization.
[0032] In one embodiment, acoustic coherence analysis is performed on a first time-frequency region based on acoustic physical parameters to obtain a time-frequency coherence matrix, including: calculating a first cross-correlation coefficient between any two array element signals in each first time-frequency region; correcting the first cross-correlation coefficient based on the multipath attenuation coefficient in the acoustic physical parameters to obtain a second cross-correlation coefficient; forming a coherence vector based on the second cross-correlation coefficient; and fusing the coherence vector with the energy information of each time-frequency region to generate a time-frequency coherence matrix.
[0033] For example, the time-frequency decomposition process for the second sound source signal needs to fully consider the signal's time-varying characteristics and frequency domain features. The window length of the preset window function is dynamically adjusted based on the sound velocity correction factor. This adjustment must follow a specific proportional relationship between the sound wave wavelength and the window length to ensure that the time-frequency analysis achieves an appropriate resolution balance in both time and frequency dimensions. The selection of the preset window function needs to consider the impact of its spectral characteristics on the analysis results. Windowed short-time Fourier transform combined with a specific type of window function can achieve a good balance between suppressing spectral leakage and maintaining frequency resolution. When loading the overlap rate parameter for time-frequency decomposition, a suitable overlap ratio needs to be determined based on the signal characteristics and analysis requirements. This parameter directly affects the continuity of the time-frequency representation and the computational complexity of the analysis. Based on the adjusted window length and the overlap rate of the time-frequency decomposition, the time-frequency decomposition process for the second sound source signal requires segmenting the signal. Each segment undergoes a specific mathematical transformation to convert it into a time-frequency domain representation. These time-frequency representations constitute the first time-frequency region, which contains the energy distribution information of the signal at different times and frequencies. The quality of the time-frequency decomposition directly affects the accuracy of subsequent acoustic feature analysis; therefore, it is necessary to strictly control the selection of each processing parameter and the precision of the calculation process.
[0034] In one embodiment, detecting a second time-frequency region dominated by a single source from a time-frequency coherence matrix based on acoustic multipath and diffraction characteristics includes: calculating the delay spread value of each time-frequency point in the time-frequency coherence matrix based on the time delay difference of acoustic wave propagation; calculating the angular spread value of each time-frequency point using a preset array manifold matrix; establishing an attenuation model based on Kirchhoff's theory of acoustic diffraction, the delay spread value, and the angular spread value, setting a multipath effect threshold of 0.3-0.5 and a diffraction attenuation threshold of 3-6 dB in the attenuation model; marking the time-frequency region that simultaneously satisfies the multipath effect threshold and the diffraction attenuation threshold as the second time-frequency region; and generating a single-source time-frequency point index based on the time-frequency coordinates of the second time-frequency region.
[0035] For example, based on the multipath and diffraction characteristics of acoustic waves, a second time-frequency region dominated by a single source is detected from the time-frequency coherence matrix. The time-frequency region contributing to the single source is filtered by calculating the delay spread and angular spread values at the time-frequency points. The delay spread value is determined by the time delay difference of acoustic wave propagation, characterizing the distribution range of propagation delay between different paths. The angular spread value is determined by eigenvalue decomposition based on a preset array manifold matrix and signal covariance matrix, and the spatial directional spread range of the signal is estimated using the angular power spectrum. These parameters collectively reflect the spatial time-frequency characteristics of signal propagation. Based on Kirchhoff diffraction theory, an acoustic wave attenuation model incorporating multipath attenuation and diffraction attenuation is constructed. Threshold screening is performed on the energy and coherence of each time-frequency region within the time-frequency coherence matrix. The multipath effect threshold is set between 0.3 and 0.5, characterizing the allowable range of multipath reflection intensity, and the diffraction attenuation threshold is set between 3 and 6 dB to limit the acceptable magnitude of acoustic wave diffraction attenuation. Only time-frequency regions that simultaneously meet the above two thresholds are marked as the second time-frequency region, thereby ensuring the single-source dominance characteristic of the selected region and possessing a high signal-to-noise ratio. Single-source time-frequency point indices are generated based on the time-frequency coordinates of the second time-frequency region, and these regions are marked using binary encoding.
[0036] In one embodiment, generating multiple direction-of-arrival (DOA) estimates based on a preset frequency domain graph signal and a single-source time-frequency point index includes: constructing a first processing model of the preset frequency domain graph signal based on acoustic physical parameters; in the first processing model, adjusting the adjacency matrix weights of the frequency domain graph signal according to a preset acoustic wavefront curvature and a preset acoustic impedance to generate a second processing model; based on the second processing model, extracting the corresponding frequency domain signal components from the single-source time-frequency point index; calculating the eigenvectors of the frequency domain signal components corresponding to the graph Laplacian matrix; calculating the spatial spectrum function of the eigenvectors using the MUSIC algorithm; and extracting the angles corresponding to the peak values of the spatial spectrum function as multiple DOA estimates.
[0037] For example, in sound source localization technology, the pre-defined frequency domain graph signal construction process involves preprocessing the original time-domain signal acquired by the microphone array, converting it into a frequency domain representation through short-time Fourier transform, and further enhancing it by combining the array's geometric topology and environmental acoustic characteristics, aiming to improve the signal's spatial resolution. For example, in an indoor acoustic environment, such pre-defined processing can cover decorrelation operations for signal aliasing caused by multipath effects, simulating the impact of reflection paths on signals by introducing a priori sound field models (such as the mirror source method), thereby forming a more robust frequency domain graph structure, where the nodes of the graph correspond to each microphone element in the array, and the edge weights reflect the spatial correlation between the elements. The core of the first processing model constructed based on acoustic physical parameters lies in using the basic laws of sound wave propagation (such as the wave equation and boundary conditions) to initialize the adjacency matrix of the frequency domain graph signal. This matrix defines the intensity of signal interaction between elements, and the acoustic physical parameters involved include sound velocity, medium density, and array spacing, etc. These parameters are usually obtained through experimental calibration or numerical simulation to ensure that the model can accurately capture the actual sound field characteristics. In this first processing model, the weights of the adjacency matrix are adjusted based on the preset wavefront curvature and acoustic impedance, which can adapt to the differences in wavefront shape caused by changes in the actual sound source location. For example, in near-field sound source scenarios, the wavefront exhibits spherical curvature. In this case, the weight calculation needs to incorporate distance attenuation factors and phase compensation. Meanwhile, the acoustic impedance affects the reflection and transmission behavior of sound waves at the medium interface. By adjusting the weight coefficients, such impedance mismatch effects can be effectively simulated, thereby generating a more accurate second processing model. This model can more realistically reflect the spatial distribution and energy transfer characteristics of the wavefront in the sound field. Based on this second processing model, the corresponding frequency domain signal components are extracted from the single-source time-frequency point index. It is necessary to identify those regions in the time-frequency domain where energy is concentrated and phase is consistent. These index points are generally provided by the previous sound source detection algorithm to ensure that each point corresponds to a potential single sound source. The extraction process includes local sampling of the frequency domain graph signal and using graph filters to separate the signal components associated with the index points, thereby obtaining a relatively pure frequency domain representation. These frequency domain signal components correspond to the calculation of eigenvectors of the graph Laplacian matrix. Essentially, this involves using graph signal processing theory to capture the global patterns of signals on the array graph structure. The graph Laplacian matrix is derived from the adjacency matrix, and its eigenvectors can reveal the smooth changes and dominant oscillation behavior of the signal in space. For example, in a uniform linear array, the eigenvectors may correspond to plane wave modes in different directions. The main components in the signal subspace can be identified through eigenvalue decomposition.A multi-signal classification algorithm is employed to calculate the spatial spectral function of the feature vectors. The principle is based on the orthogonality between the signal subspace and the noise subspace to estimate the direction of arrival (DOA). The specific steps include projecting the feature vectors onto a steering vector space defined by an array manifold matrix and calculating its distance metric from the noise subspace, thereby generating a spatial spectral function. This function traverses the entire omnidirectional range at fixed angular steps, for example, scanning from 0 degrees to 360 degrees at 1-degree intervals, thus accurately identifying the spectral peak positions. The extraction of the angles corresponding to the peaks of the spatial spectral function, serving as multiple DOA estimates, is accomplished by detecting local maxima within the spectral function. Each significant peak represents a potential sound source direction, highly reflecting the signal intensity and correlation probability. For example, in scenarios with multiple sound sources, the spectral function may exhibit multiple peaks, each corresponding to the incident direction of a different sound source. This outputs a set of candidate DOA data for subsequent fusion and analysis.
[0038] In one embodiment, clustering and weighted fusion of multiple direction-of-arrival (DOA) estimates to obtain the DOA information of a sound source includes: performing cluster analysis on the multiple DOA estimates to determine multiple clusters; calculating the sound intensity attenuation factor corresponding to the DOA estimates contained in each cluster; determining the physical confidence level of each DOA estimate based on the multipath attenuation coefficient in the acoustic physical parameters; and using the product of the sound intensity attenuation factor and the physical confidence level as the fusion weight to perform a weighted average calculation on the DOA estimates contained in each cluster to obtain the DOA information.
[0039] For example, cluster analysis can be performed on multiple direction-of-arrival (DOA) estimates. Unsupervised learning algorithms can be used to group spatially similar DOA values into several clusters to eliminate random errors and noise interference. For instance, the density-based DBSCAN clustering method can be applied. This method automatically divides clusters based on the distribution density of the estimated values in angular space and their neighborhood relationships, effectively handling outliers and ensuring the robustness of the clustering results. During the clustering process, relevant parameters such as neighborhood radius and minimum number of points need to be adjusted according to the actual acoustic environment. For example, in a conference room sound source localization application, setting the neighborhood radius to 5 degrees helps to distinguish the directions of different speakers, thereby forming multiple clusters representing the directions of potential sound sources. The calculation of the sound intensity attenuation factor corresponding to the direction-of-arrival (DOA) estimates within each cluster requires analyzing the energy loss of the sound wave along its propagation path. This attenuation factor depends on the distance from the sound source to the array, the medium absorption coefficient, and the environmental reflection characteristics. For example, under free-field conditions, the attenuation factor can be calculated based on the inverse square law. However, in indoor multipath environments, a ray tracing model is needed to estimate the additional attenuation caused by multiple reflections, thereby assigning an index reflecting the energy reliability of each DOA estimate. The physical confidence of each DOA estimate is determined based on the multipath attenuation coefficient in the acoustic physics parameters. The multipath attenuation coefficient is derived from the measurement or simulation of the signal attenuation after the sound wave undergoes reflection, diffraction, and scattering in the environment. A higher coefficient means that the estimate is more severely affected by multipath interference, and therefore its confidence is lower. For example, in scenarios with strong reflective surfaces, the multipath attenuation coefficient may increase significantly, leading to a decrease in the confidence of estimates in some directions. The calculation of physical confidence usually combines the signal-to-noise ratio (SNR) and coherence index to quantify the reliability of the estimate in complex sound fields. Using the product of sound intensity attenuation factor and physical confidence level as the fusion weight, a weighted average is calculated for the direction of arrival (DOA) estimates within each cluster. This aims to balance the contributions of different estimates. The weight design ensures that estimates with high energy and low multipath impact receive higher priority. For example, during the weighting process, the weights are normalized and the weight variance is limited to avoid the excessive influence of individual outliers, thereby generating a stable DOA information that represents the weighted center of the direction within the cluster. This effectively integrates multiple estimates and improves positioning accuracy.
[0040] In one embodiment, calculating the first cross-correlation coefficient between any two array element signals within each first time-frequency region includes: obtaining a signal pair between any two array elements within the first time-frequency region, calculating a generalized cross-correlation function between the signal pairs; performing PHAT weighting on the generalized cross-correlation function to obtain a third cross-correlation function; extracting the peak position from the third cross-correlation function, and calculating the first cross-correlation coefficient based on the peak position and the third cross-correlation function.
[0041] For example, acquiring signal pairs of any two array elements within the first time-frequency region is accomplished by selecting a specific combination of array elements in the microphone array and extracting corresponding signal segments in the time-frequency domain. The time-frequency region is defined by the short-time Fourier transform, for example, using a Hamming window and a 50% overlap rate to balance time-frequency resolution. The selection of signal pairs typically covers all possible array element combinations, thereby comprehensively capturing spatial correlation. The calculation of the generalized cross-correlation function between signal pairs involves multiplying the frequency domain representations of the two signals by their complex conjugates and applying a weighting function to estimate the time delay. The generalized cross-correlation function enhances robustness to noise and reverberation. Its calculation process involves standardizing the signal spectrum to highlight phase information. For example, in environments with background noise, the time delay estimation sequence can be obtained through the integral form of this function, providing basic data for subsequent analysis. The generalized cross-correlation function is weighted by phase transform to obtain the third cross-correlation function. Phase transform weighting, as a frequency domain amplitude normalization technique, focuses on the phase difference by eliminating the influence of signal amplitude variations, thereby improving the accuracy of time delay estimation. For example, in indoor acoustic applications, this weighting can effectively suppress spurious peaks caused by multipath effects, making the third cross-correlation function more clearly indicate the true time delay location. The peak position is extracted from the third cross-correlation function, and the first cross-correlation coefficient is calculated based on this position and the function value. This is achieved by detecting global or local maxima in the function curve. The peak position corresponds to the time delay between signals, while the cross-correlation coefficient is calculated by combining the peak amplitude and the function shape. For example, the ratio of peak height to the total energy of the function can be used to quantify the correlation strength. This coefficient not only reflects the similarity of the signal pairs but also provides crucial time delay information for subsequent source localization and direction-of-arrival estimation.
[0042] In one embodiment, calculating the angular spread value at each time-frequency point using a preset array manifold matrix includes: constructing an array manifold matrix corresponding to the geometry of the microphone array; calculating the signal covariance matrix corresponding to each time-frequency point; performing eigenvalue decomposition on the signal covariance matrix to obtain signal eigenvalues and signal eigenvectors; calculating the signal subspace based on the signal eigenvalues and signal eigenvectors; calculating the angular power spectrum based on the array manifold matrix and the signal subspace; and calculating the angular spread value at each time-frequency point based on the variance of the angular power spectrum.
[0043] For example, constructing the array manifold matrix corresponding to the microphone array geometry requires defining the response vector of each element in the array at different incident angles. This matrix contains information such as the spatial position, frequency response, and steering vector of the elements. For instance, in a uniform circular array, this matrix can be generated based on the assumption of spherical waves or plane waves, ensuring that it strictly reflects the array's spatial sampling capability for the direction of sound waves. Pre-setting steps may involve calibrating the phase differences and gain variations between elements to ensure the accuracy of the matrix. The calculation of the signal covariance matrix corresponding to each time-frequency point is achieved by statistically averaging the multi-channel signals of the array at that time-frequency point. The covariance matrix captures the spatial correlation and energy distribution between signals. For example, in broadband sound source processing, it may be necessary to calculate the covariance matrix separately on multiple frequency sub-bands to comprehensively characterize the spatiotemporal characteristics of the signal. The signal covariance matrix is decomposed into eigenvalues and eigenvectors. Eigenvalue decomposition divides the matrix into a signal subspace and a noise subspace. The eigenvectors corresponding to larger eigenvalues represent the dominant signal pattern, while smaller eigenvalues are related to noise components. For example, in the presence of multiple sound sources, the distribution of eigenvalues can reveal the number and directionality of the sources. The signal subspace is calculated based on the signal eigenvalues and their eigenvectors, constructed by selecting the eigenvectors corresponding to the k largest eigenvalues. This subspace effectively represents the main spatial characteristics of the sound source and is used for subsequent angle estimation. The angular power spectrum is calculated based on the array manifold matrix and the signal subspace. It is achieved by projecting the signal subspace onto the steering vector space defined by the array manifold matrix. The angular power spectrum displays the relative energy distribution of the signal at different incident angles. For example, such spectra can be generated using multi-signal classification algorithms or beamforming techniques, thereby identifying the direction of the sound source. The angular spread value at each time frequency point is calculated based on the variance of the angular power spectrum. The variance calculation quantifies the degree of diffusion of the power spectrum in the angular domain. A higher angular spread value indicates that the sound wave direction is more dispersed, which may be due to multipath effect or diffuse sound field. For example, in an indoor environment, a higher angular spread value may indicate the spatial scattering characteristics of the sound wave after multiple reflections. This value can be used to distinguish between a single-direction sound source and a complex acoustic scene.
[0044] Please see Figure 2 , Figure 2 This is a schematic block diagram of a sound source localization system based on a multi-microphone array, provided in an embodiment of this application. The sound source localization system 200 based on a multi-microphone array is used to execute the aforementioned sound source localization method based on a multi-microphone array. The sound source localization system 200 based on a multi-microphone array can be configured in a server.
[0045] The server can be a standalone server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0046] like Figure 2 As shown, the sound source localization system 200 based on a multi-microphone array includes: a signal processing module 201, a time-frequency decomposition module 202, an acoustic analysis module 203, a single source determination module 204, and direction estimation modules 205 and 206.
[0047] The signal processing module 201 is used to preprocess the first sound source signal acquired by the microphone array and output the second sound source signal and acoustic physical parameters.
[0048] The time-frequency decomposition module 202 is used to perform time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions.
[0049] The acoustic analysis module 203 is used to perform acoustic coherence analysis on the first time-frequency region based on acoustic physical parameters to obtain the time-frequency coherence matrix.
[0050] The single-source determination module 204 is used to detect a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on the multipath and diffraction characteristics of acoustic waves, and to construct an output single-source time-frequency point index based on the second time-frequency region.
[0051] The direction estimation module 205 is used to generate multiple direction of arrival estimates based on a preset frequency domain diagram signal and a single-source time-frequency point index. The preset frequency domain diagram signal is obtained by processing acoustic physical parameters using a preset enhancement model.
[0052] The output module 206 is used to cluster and weighted fuse multiple direction-of-arrival estimates to obtain the direction-of-arrival information of the sound source.
[0053] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A sound source localization method based on a multi-microphone array, characterized in that, The method includes: The first sound source signal acquired by the microphone array is preprocessed, and the second sound source signal and acoustic physical parameters are output. The second sound source signal is decomposed into multiple first time-frequency regions by time-frequency decomposition. Acoustic coherence analysis is performed on the first time-frequency region based on the acoustic physical parameters to obtain the time-frequency coherence matrix; Based on the multipath and diffraction characteristics of acoustic waves, a second time-frequency region dominated by a single source is detected from the time-frequency coherence matrix, and an output single-source time-frequency point index is constructed based on the second time-frequency region. Multiple direction-of-arrival (DOA) estimates are generated based on the preset frequency domain diagram signal and the single-source time-frequency point index. Clustering and weighted fusion of the multiple direction-of-arrival (DOA) estimates are performed to obtain the DOA information of the sound source.
2. The method as described in claim 1, characterized in that, The preprocessing of the first sound source signal acquired by the microphone array to output the second sound source signal and acoustic physical parameters includes: Calculate the first phase difference of the received signal of each element based on the spatial position of each element in the microphone array; The first phase difference is calibrated and compensated to obtain the second phase difference; The first sound source signal is phase-calibrated based on the second phase difference to obtain the third sound source signal; Multipath suppression is performed on the third sound source signal to obtain the second sound source signal; The acoustic physical parameters are determined based on the second phase difference and the second sound source signal. The acoustic physical parameters include: sound velocity correction factor and multipath attenuation coefficient.
3. The method as described in claim 2, characterized in that, The time-frequency decomposition of the second sound source signal yields multiple first time-frequency regions, including: The window length of the preset window function is dynamically adjusted based on the sound speed correction factor. The window length is proportional to the sound wave wavelength. The preset window function includes: windowed short-time Fourier transform and Hanning window. Load the overlap rate of the time-frequency decomposition, wherein the overlap rate of the time-frequency decomposition is 50%-75%; Based on the adjusted window length and the overlap rate of the time-frequency decomposition, the second sound source signal is decomposed into multiple first time-frequency regions.
4. The method as described in claim 2, characterized in that, The step of performing acoustic coherence analysis on the first time-frequency region based on the acoustic physical parameters to obtain the time-frequency coherence matrix includes: Calculate the first cross-correlation coefficient between any two array element signals in each of the first time-frequency regions; The first cross-correlation coefficient is corrected based on the multipath attenuation coefficient in the acoustic physical parameters to obtain the second cross-correlation coefficient; A coherence vector is formed based on the second cross-correlation coefficient; The coherence vector is fused with the energy information of each time-frequency region to generate the time-frequency coherence matrix.
5. The method as described in claim 1, characterized in that, The method of detecting a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on acoustic multipath and diffraction characteristics includes: The delay spread value at each time frequency point in the time-frequency coherence matrix is calculated based on the time delay difference of sound wave propagation. The angular spread value of each time-frequency point is calculated using a preset array manifold matrix; An attenuation model is established based on Kirchhoff's theory of acoustic diffraction, the delay spread value, and the angular spread value. In the attenuation model, the multipath effect threshold is set to 0.3-0.5, and the diffraction attenuation threshold is set to 3-6 dB. The time-frequency region that simultaneously satisfies the multipath effect threshold and the diffraction attenuation threshold is marked as the second time-frequency region; The single-source time-frequency point index is generated based on the time-frequency coordinates of the second time-frequency region.
6. The method as described in claim 1, characterized in that, The process of generating multiple direction-of-arrival (DOA) estimates based on a preset frequency domain signal and the single-source time-frequency index includes: A first processing model for the preset frequency domain signal is constructed based on the acoustic physical parameters; In the first processing model, the adjacency matrix weights of the frequency domain graph signal are adjusted according to the preset acoustic wavefront curvature and the preset acoustic impedance to generate the second processing model. Based on the second processing model, the corresponding frequency domain signal components are extracted from the single-source time-frequency point index; Calculate the eigenvectors of the frequency domain signal components corresponding to the Graph Laplacian matrix; The spatial spectral function of the eigenvectors is calculated using the MUSIC algorithm; The angle corresponding to the peak value of the spatial spectral function is extracted as the estimated value of the multiple directions of arrival.
7. The method as described in claim 1, characterized in that, The process of clustering and weighted fusion of the multiple direction-of-arrival (DOA) estimates to obtain the DOA information of the sound source includes: Cluster analysis is performed on multiple directions of arrival estimates to determine multiple clusters; Calculate the acoustic intensity attenuation factor corresponding to the estimated direction of arrival contained in each of the clusters; The physical confidence level of each of the direction-of-arrival estimates is determined based on the multipath attenuation coefficient in the acoustic physical parameters. The product of the sound intensity attenuation factor and the physical confidence level is used as the fusion weight to calculate the weighted average of the direction of arrival estimates contained in each cluster, thereby obtaining the direction of arrival information.
8. The method as described in claim 4, characterized in that, The calculation of the first cross-correlation coefficient between any two array element signals within each of the first time-frequency regions includes: Obtain signal pairs of any two array elements within the first time-frequency region, and calculate the generalized cross-correlation function between the signal pairs; The generalized cross-correlation function is subjected to PHAT weighting to obtain the third cross-correlation function; The peak position is extracted from the third cross-correlation function, and the first cross-correlation coefficient is calculated based on the peak position and the third cross-correlation function.
9. The method as described in claim 5, characterized in that, The calculation of the angle spread value for each time-frequency point using a preset array manifold matrix includes: Construct an array manifold matrix corresponding to the geometry of the microphone array; Calculate the signal covariance matrix corresponding to each of the aforementioned time-frequency points; The signal covariance matrix is decomposed into eigenvalues to obtain signal eigenvalues and signal eigenvectors. Calculate the signal subspace based on the signal feature values and the signal feature vectors; Calculate the angular power spectrum based on the array manifold matrix and the signal subspace; The angular spread value at each time-frequency point is calculated based on the variance of the angular power spectrum.
10. A sound source localization system based on a multi-microphone array, characterized in that, The sound source localization system based on a multi-microphone array is used to execute the sound source localization method based on a multi-microphone array as described in any one of claims 1-9, wherein the sound source localization system based on a multi-microphone array comprises: The signal processing module is used to preprocess the first sound source signal acquired by the microphone array and output the second sound source signal and acoustic physical parameters. The time-frequency decomposition module is used to perform time-frequency decomposition on the second sound source signal to obtain multiple first time-frequency regions; An acoustic analysis module is used to perform acoustic coherence analysis on the first time-frequency region based on the acoustic physical parameters to obtain a time-frequency coherence matrix; A single-source determination module is used to detect a second time-frequency region dominated by a single source from the time-frequency coherence matrix based on acoustic multipath and diffraction characteristics, and to construct an output single-source time-frequency point index based on the second time-frequency region. The direction estimation module is used to generate multiple direction-of-arrival estimates based on the preset frequency domain diagram signal and the single-source time-frequency point index. The result output module is used to cluster and weightedly fuse the multiple direction-of-arrival estimates to obtain the direction-of-arrival information of the sound source.