Sound source localization method and device based on acoustics and vibration heterogeneous perception
By combining asymmetric heterogeneous arrays with acoustic and vibration sensors to form a feature-level sparse fusion localization method, the problem of poor localization robustness of multi-microphone arrays in reverberant environments is solved, achieving low-cost, high-precision sound source localization, which is applicable to a variety of complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PIONEER TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing multi-microphone array sound source localization technology has poor robustness in reverberant environments, is susceptible to interference from reflected waves, has high hardware costs, requires high synchronization accuracy, and the algorithm relies on the time difference or phase difference of sound wave propagation for calculation, resulting in insufficient localization accuracy and anti-interference capability.
A sparse fusion localization method is constructed by using an asymmetric heterogeneous array, combined with acoustic and vibration sensors, and extracting Mel frequency cepstral coefficients and wavelet packet energy features. The sound source location is calculated by L1 norm regularized least squares optimization, and environmental adaptive calibration is used to update the reference matrix.
It achieves high-precision sound source localization in low synchronization accuracy and complex environments, has strong anti-interference capabilities, reduces hardware costs, and is suitable for scenarios such as indoor robots, vehicle-mounted voice interaction, and equipment fault localization, with a positioning error of less than 3°.
Smart Images

Figure CN121899751A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization technology, and in particular to a sound source localization method and apparatus based on acoustic and vibration heterogeneous sensing. Background Technology
[0002] With the development of technologies and devices such as intelligent sensing and intelligent terminals, sound source localization is being used more and more widely. Using microphone arrays to locate sound sources has become a hot research area, with significant applications in intelligent human-computer interaction and industrial control. Microphone array technology changes the traditional use of a single microphone by arranging multiple microphones in a specific array shape to simultaneously collect speech. By analyzing and processing multiple signals and combining them with the known spatial geometry of the microphone array, the planar or spatial coordinates of one or more sound sources can be determined in the spatial domain, thereby calculating the location of the sound source.
[0003] Existing multi-microphone arrays typically employ linear or circular array arrangements. Their positioning relies on algorithms such as TDOA (Time Difference of Arrival), PHAT (Phase Transform), and SRP-PHAT (Synchronous Power-Responsive Phase Transform), requiring strict synchronization of multi-channel signals, sometimes even at the microsecond level, resulting in high hardware costs. Furthermore, these algorithms inherently depend on calculating the time or phase difference of sound wave propagation, leading to poor positioning robustness and susceptibility to interference from reflected waves in reverberant environments, such as indoors or enclosed spaces.
[0004] Therefore, the existing technology of using microphone arrays for sound source localization has significant limitations. Summary of the Invention
[0005] To address the above problems, this invention provides a sound source localization method based on acoustic and vibration heterogeneous sensing, the method comprising:
[0006] S1: Arrange M acoustic sensors and N vibration sensors into an asymmetric heterogeneous array and fix them to the surface of the carrier;
[0007] S2: Synchronous acquisition of acoustic signals and vibration signals ,in, , ;
[0008] S3: Acoustic signal Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. And the K-dimensional feature matrix of M acoustic sensors The signals are spliced together to form a large M×K dimensional acoustic signal matrix A;
[0009] Vibration signals Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. And the Q-dimensional feature matrix of N vibration sensors A large matrix V of N×Q dimensional vibration signals is spliced together;
[0010] S4: Concatenate the large acoustic signal matrix A and the large vibration signal matrix V to form the observation vector X.
[0011] ;
[0012] S5: Compare the observation vector X with the reference matrix D of the reference sound source to obtain the azimuth angle of the observation vector X. .
[0013] As an optional technical solution, a preprocessing step S0 is included before step S1, and the preprocessing step S0 includes:
[0014] Acoustic reference feature matrices of the reference sound sources were collected and obtained respectively. and vibration reference characteristic matrix ,in The range is from 0° to 359°, with a step size of 1°, and a reference matrix D of the reference sound source is constructed.
[0015] .
[0016] As an optional technical solution, in step S3, "the acoustic signal..." Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. "include:
[0017] acoustic signals After pre-emphasis, framing, Mel filtering, and DCT transformation, the K-dimensional feature matrix of the M acoustic sensors is obtained. .
[0018] As an optional technical solution, K∈{12,13,14,15,16}.
[0019] As an optional technical solution, step S3 "on the vibration signal" Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. "middle, , .
[0020] As an optional technical solution, in step S3, "the K-dimensional feature matrix of the M acoustic sensors is..." The splicing method in "A" which is spliced into an M×K dimensional acoustic signal matrix is as follows:
[0021] ;
[0022] "The Q-dimensional feature matrix of N vibration sensors" The splicing method in the splicing of the N×Q dimensional vibration signal large matrix V is as follows:
[0023] .
[0024] As an optional technical solution, step S5 includes:
[0025] S51: Substitute the observation vector X and the reference matrix D of the reference sound source into the L1 norm regularized least squares optimization formula to calculate the sparse coefficient vector. ,
[0026] ,
[0027] in, The regularization parameter and ;
[0028] S52: For sparse coefficient vectors Indexing to find the sparse coefficient vector The angle corresponding to the element with the largest median value Let X be the azimuth angle of the observation vector X. ,
[0029] .
[0030] As an optional technical solution, step S5 is followed by an environmental calibration step S6, which includes:
[0031] S61: Collect ambient background sound features and ambient background vibration features at intervals t, and obtain the background acoustic feature matrix. and background vibration characteristic matrix ;
[0032] S62: Background acoustic feature matrix and background vibration characteristic matrix Concatenate into background observation vector ,
[0033] ;
[0034] S63: Calculate the background sparse coefficient vector ;
[0035] S64: Update the reference matrix D and obtain the updated reference matrix D';
[0036] , It is a constant;
[0037] S65: Substitute the updated reference matrix D' into step S5 for calculation.
[0038] To address the above problems, this invention provides a sound source localization device based on acoustic and vibration heterogeneous sensing, the localization device comprising:
[0039] A heterogeneous sensing array, comprising M acoustic sensors and N vibration sensors, wherein the heterogeneous sensing array is arranged in an asymmetric topological layout;
[0040] The signal acquisition module is used to synchronously acquire acoustic signals. and vibration signals ,in, , ;
[0041] The feature extraction module is used for acoustic signal processing. Extract Mel-frequency cepstral coefficient features and obtain K-dimensional features for each of the M acoustic sensors. For vibration signals Extract wavelet packet energy features and obtain Q-dimensional features from N vibration sensors. ;
[0042] The feature stitching module is used to stitch together the K-dimensional features of M acoustic sensors. The acoustic signals are spliced into a large M×K dimensional matrix A, which incorporates the Q-dimensional features of N vibration sensors. A large N×Q dimensional vibration signal matrix V is concatenated, and the large acoustic signal matrix A and the large vibration signal matrix V are concatenated to form the observation vector X. ;
[0043] The sparse fusion localization module is used to construct the reference matrix D of the reference sound source, and to retrieve and obtain the azimuth angle θ of the observation vector X by substituting the observation vector X into the reference matrix D.
[0044] As an optional technical solution, the sound source localization device also includes an environment adaptive calibration module, which is used to collect environmental background sound characteristics and environmental background vibration characteristics, and update the reference matrix D in real time.
[0045] Compared with existing technologies, this invention avoids traditional algorithms such as TDOA (Time Difference of Arrival) or PHAT (Phase Transformation) and adopts feature-level sparse fusion, which eliminates the need to calculate the signal arrival time difference and achieves a synchronization accuracy of ≤1ms, greatly reducing the requirements for multi-channel synchronization.
[0046] Furthermore, this technical solution exhibits strong anti-interference capabilities. While the acoustic characteristics may be affected by ambient air noise, the vibration characteristics of the carrier, directly collected from the carrier surface, are almost unaffected by air noise and can still stably reflect the location of the sound source. When the carrier itself experiences vibration interference, the vibration characteristics will introduce noise, but the acoustic characteristics can serve as a reference, filtering out false peaks caused by vibration interference. It can achieve a positioning error of ≤3° in environments with a signal-to-noise ratio ≥5dB and a reverberation time ≤0.5s.
[0047] Furthermore, this technical solution employs an asymmetric heterogeneous array, enhancing directional resolution through spatial complementarity. Therefore, this solution is applicable to a wide range of scenarios, including indoor robots, in-vehicle voice interaction, and equipment fault location, eliminating the need for complex hardware synchronization modules and significantly reducing costs. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0050] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0051] like Figure 1 As shown, this invention provides a sound source localization method based on acoustic and vibration heterogeneous sensing, the method comprising:
[0052] S1: Arrange M acoustic sensors and N vibration sensors into an asymmetric heterogeneous array and fix them to the surface of the carrier;
[0053] S2: Synchronous acquisition of acoustic signals and vibration signals ,in, , ;
[0054] S3: Acoustic signal Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. And the K-dimensional feature matrix of M acoustic sensors The signals are spliced together to form a large M×K dimensional acoustic signal matrix A;
[0055] Vibration signals Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. And the Q-dimensional feature matrix of N vibration sensors A large matrix V of N×Q dimensional vibration signals is spliced together;
[0056] S4: Concatenate the large acoustic signal matrix A and the large vibration signal matrix V to form the observation vector X.
[0057] ;
[0058] S5: Compare the observation vector X with the reference matrix D of the reference sound source to obtain the azimuth angle of the observation vector X. .
[0059] In this technical solution, traditional algorithms such as TDOA (Time Difference of Arrival) or PHAT (Phase Transformation) are avoided. Instead, feature-level sparse fusion is used, which eliminates the need to calculate the signal arrival time difference and achieves a synchronization accuracy of ≤1ms, greatly reducing the requirements for multi-channel synchronization.
[0060] Furthermore, this technical solution exhibits strong anti-interference capabilities. While the acoustic characteristics may be affected by ambient air noise, the vibration characteristics of the carrier, directly collected from the carrier surface, are almost unaffected by air noise and can still stably reflect the location of the sound source. When the carrier itself experiences vibration interference, the vibration characteristics will introduce noise, but the acoustic characteristics can serve as a reference, filtering out false peaks caused by vibration interference. It can achieve a positioning error of ≤3° in environments with a signal-to-noise ratio ≥5dB and a reverberation time ≤0.5s.
[0061] Furthermore, this technical solution employs an asymmetric heterogeneous array, enhancing directional resolution through spatial complementarity. Therefore, this solution is applicable to a wide range of scenarios, including indoor robots, in-vehicle voice interaction, and equipment fault location, eliminating the need for complex hardware synchronization modules and significantly reducing costs.
[0062] This technical solution employs heterogeneous sensing, which refers to multiple types of sensing signals generated from the same sound source, despite differences in sensing origin, physical properties, and propagation characteristics. This solution uses acoustic and vibration sensors to generate acoustic and vibration signals respectively, and then fuses these two signals with different properties in the subsequent algorithm. The anti-interference capability of the vibration signal compensates for the deficiencies of the acoustic signal, while the distance sensing capability of the acoustic signal compensates for the limitations of the vibration signal, greatly improving the accuracy of positioning.
[0063] Specifically, a preprocessing step S0 is included before step S1, and the preprocessing step S0 includes:
[0064] Acoustic reference feature matrices of the reference sound sources were collected and obtained respectively. and vibration reference characteristic matrix ,in The range is from 0° to 359°, with a step size of 1°, and a reference matrix D of the reference sound source is constructed.
[0065] .
[0066] This process involves constructing a joint sparse dictionary, placing a reference sound source in the experimental environment, and sequentially acquiring the acoustic reference feature matrix of the reference sound source from 0° to 359° with a step size of 1°. and vibration reference characteristic matrix The acquisition and calculation process is similar to steps S2 and S3 above. Specifically, the acoustic reference signal is acquired synchronously first. and vibration reference signal Then, the acoustic reference signal Extracting Mel frequency cepstral coefficient features to obtain the acoustic reference feature matrix. Simultaneously, the vibration reference signal Extract wavelet packet energy features to obtain the vibration reference feature matrix. . This represents vector concatenation, using the formula above to combine the acoustic reference feature matrix from 0 to 359°. and vibration reference signal The components are spliced and integrated separately to form a reference matrix D for the reference sound source, which is to construct a joint sparse dictionary to serve as a reference for the sound source in specific situations, so as to identify the specific angle of the sound source in specific situations.
[0067] The S0 step and the subsequent S1 to S5 steps are not in the same environment. The S0 step is located in the experimental environment and is used to construct a joint sparse dictionary from different angles. The subsequent S1 to S5 steps are located in a specific detection environment and are used to detect and locate the actual sound source.
[0068] Specifically, in step S1, M acoustic sensors and N vibration sensors are arranged into an asymmetric heterogeneous array and fixed to the surface of a carrier. The carrier can be planar or non-planar, without affecting the actual positioning effect. The acoustic sensors are MEMS microphones, and the vibration sensors are piezoelectric accelerometers, using an asymmetric topological layout. The planes on which all sensors are arranged are not collinear or circular, avoiding regular array topologies. Randomly distributed, irregular polygonal derivative, or hybrid distributed layouts can be used to improve the positioning robustness in complex scenarios. M≥2, N≥1, and the spatial distance between the acoustic and vibration sensors is 5-50mm, resulting in a relatively compact distribution. This technical solution provides a specific embodiment where, when M=3 and N=1, the acoustic sensor coordinates are (0,0), (30mm,10mm), and (15mm,40mm), and the vibration sensor coordinates are (20mm,25mm), which meets the above requirements.
[0069] In step S2, acoustic signals are acquired synchronously. and vibration signals ,in, , This step requires synchronous signal acquisition. However, due to the compact array and the use of feature-level sparse fusion of heterogeneous signals in this scheme, there is no need to calculate the signal arrival time difference, thus the synchronization accuracy can be greatly reduced. The required synchronization accuracy is ≤1ms, which is 10 times lower than that of traditional arrays. Specifically, the sampling frequency in this scheme is 16-48kHz.
[0070] In step S3, "the acoustic signal" Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. "include:
[0071] acoustic signals After pre-emphasis, framing, Mel filtering, and DCT transformation, the K-dimensional feature matrix of the M acoustic sensors is obtained. .
[0072] because Therefore, the K-dimensional characteristics of acoustic sensors There are M in total. In this specific implementation plan, the first to the Mth acoustic signals will be... After pre-emphasis, framing, Mel filtering, and DCT transform, K-dimensional MFCC (Mel frequency cepstral coefficients) features are extracted. In this scheme, K-dimensional MFCCs are extracted for each acoustic sensor, and finally, they are spliced into an M×K dimensional acoustic signal matrix A.
[0073] In this specific implementation, the pre-emphasis process enhances the energy of high-frequency signals, highlighting the high-frequency differences of sound sources in different directions, while reducing noise interference in subsequent processing. The framing process divides the signal into multiple frames according to framing parameters, facilitating the extraction of stable features from each frame. Mel filtering filters the framed signal, retaining only the frequency components useful for subsequent localization to obtain the Mel spectrum. The DCT (Discrete Cosine Transform) process performs a discrete cosine transform on the Mel spectrum, extracting K coefficient features, retaining key information, and further filtering noise.
[0074] Therefore, the feature matrix of each acoustic sensor Each of them has K-dimensional features, and thus M acoustic sensors can be spliced together to form an M×K-dimensional acoustic signal matrix A.
[0075] In this specific embodiment, If the value of K is too small, the collected feature information will be insufficient, making it impossible to distinguish similar locations. If the value of K is too large, there will be redundancy, increasing computational complexity. Of course, if the value of K is in other ranges, as long as the purpose of this technical solution is achieved, it is within the protection scope of this invention.
[0076] In step S3, "for vibration signals" Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. "middle , .
[0077] In this specific embodiment, the vibration signal Wavelet packet decomposition is used, and the vibration signal of each vibration sensor is... The signal is decomposed into Q-dimensional signals. Wavelet packet decomposition is a multi-scale time-frequency analysis method for non-stationary signals, such as vibration signals. Its core function is to decompose the signal into multiple non-overlapping frequency sub-bands, simultaneously preserving the signal's frequency and time characteristics. This embodiment uses two-band wavelet packet decomposition, which is also an inherent frequency band number rule of wavelet packet decomposition. In the first level of decomposition, the original vibration signal is... The first level of decomposition divides the frequency bands into two; in the second level, these two bands are further divided into two, resulting in four frequency bands; in the third level, these four bands are again divided into two, resulting in eight frequency bands. Therefore, in the L-level decomposition, a total of [number missing] frequency bands can be obtained. Each frequency band corresponds to a fixed, non-overlapping frequency range, and energy calculations are performed on each frequency band to obtain the Q-dimensional characteristics of each vibration sensor. .
[0078] In this specific embodiment, Therefore, the value of Q ranges from 8 to 32. If the value of L is larger, the frequency band division is finer and the characterization of signal frequency details is more accurate, but the amount of calculation will also increase accordingly.
[0079] Furthermore, in step S3, "the K-dimensional feature matrix of the M acoustic sensors is..." The splicing method in "A" which is spliced into an M×K dimensional acoustic signal matrix is as follows:
[0080] ;
[0081] "The Q-dimensional feature matrix of N vibration sensors" The splicing method in the splicing of the N×Q dimensional vibration signal large matrix V is as follows:
[0082] .
[0083] As mentioned above, in obtaining the K-dimensional feature matrix of each acoustic sensor Then, use the formula. M submatrices , , ..., After transposing each matrix separately, they are concatenated in the column direction to form an intermediate matrix. Then, the intermediate matrix is transposed as a whole to obtain the large acoustic signal matrix A. Essentially, this involves stacking M sub-matrices in the row direction.
[0084] Similarly, in obtaining the Q-dimensional feature matrix of each vibration sensor Then, use the formula. , divide N submatrices After transposing each matrix separately, they are concatenated in the column direction to form an intermediate matrix. Then, the intermediate matrix is transposed as a whole to obtain the large vibration signal matrix V, which is essentially a stack of N sub-matrices in the row direction.
[0085] Therefore, after obtaining the large acoustic signal matrix A and the large vibration signal matrix V, the subsequent steps require further fusion of the acoustic and vibration signals, thereby achieving dual-modal feature fusion.
[0086] In step S4, the acoustic signal matrix A and the vibration signal matrix V are concatenated to form the observation vector X. The concatenation process is essentially a column-wise concatenation of two row vectors of the same dimension. The observation vector X is a cross-modal fusion observation carrier that contains both acoustic (MFCC) and vibration (wavelet packet energy) feature information, solving the problem of incomplete feature information in a single mode and providing comprehensive feature input for subsequent sparse coefficient solving.
[0087] In this step S4, the dimensions of the acoustic signal matrix A and the vibration signal matrix V can be unified, enabling dual-modal feature complementarity, improving positioning accuracy and anti-interference capability. At the same time, there is no need for complex weight allocation; it can be achieved directly by transposing, flattening, and splicing the matrices.
[0088] Additionally, it should be noted that during the process of constructing the reference matrix D of the reference sound source, the acoustic reference feature matrix is detected and calculated. and vibration reference characteristic matrix The process is similar to the steps described above, and the reference matrix D of the sound source is used for each angle. Acoustic reference feature matrix and vibration reference characteristic matrix The fusion process is similar to the process described above, namely... This ensures consistency between the reference and detection stages, avoids feature distortion caused by differences in fusion methods, and ensures the matching degree between the joint sparse dictionary, i.e., the reference matrix D, and the real-time detection features.
[0089] After obtaining the observation vector X, it is necessary to input the observation vector X into the reference matrix D and then retrieve the azimuth angle of the observation vector. Step S5 includes:
[0090] S51: Substitute the observation vector X and the reference matrix D of the reference sound source into the L1 norm regularized least squares optimization formula to calculate the sparse coefficient vector. ,
[0091] ,
[0092] in, The regularization parameter and ;
[0093] S52: For sparse coefficient vectors Indexing to find the sparse coefficient vector The angle corresponding to the element with the largest median value Let X be the azimuth angle of the observation vector X. ,
[0094] .
[0095] Specifically, in step S51, Using the L1 norm forces x to be sparse, thus promoting a sparse coefficient vector. It has sparsity, that is, it yields a sparse coefficient vector. Most of the elements are 0, except for one element which is significantly larger than the actual sound source location. The norm squared of L2 is used as a fitting error constraint to measure the difference between the "product of the reference matrix D and the coefficients x" and the observed data X. The regularization parameter and λ is used to balance the weight between "sparseness requirements" and "data fitting accuracy". The larger λ is, the more importance is placed on "fitting accuracy", but x may have lower sparsity and multiple high-weight orientations may appear. The smaller λ is, the more importance is placed on "sparseness", but the fitting error may be large, resulting in inaccurate positioning. In essence, the goal is to find the minimum value of the sparse coefficients x, resulting in a 360×1 column vector x that minimizes the L2 norm (weighted by L1 norm and λ). This column vector x is the final sparse coefficient vector used for localization. .
[0096] Furthermore, in finding the sparse coefficient vector Afterwards, although the sparse vector x has 360 columns, most of the values are 0 after sparsification, with only one maximum value. Therefore, the next step is to find the maximum value in the sparse vector x and the angle corresponding to that maximum value. That is, the actual azimuth angle of the sound source. Specifically, the function argmax is a function that finds the value of the independent variable corresponding to the maximum value, in order to find the θ that maximizes x(θ) in the formula.
[0097] To ensure the long-term usability of the reference matrix D and the stability of the positioning, this specific embodiment also requires that the reference matrix D be updated in real time according to the detection background, so that the reference matrix D can perform environmental adaptive calibration, preventing the original reference matrix D from misjudging the background interference as sound source features when the background interference at the detection site changes, resulting in positioning drift or errors.
[0098] Therefore, in this embodiment, step S5 is followed by an environmental calibration step S6, which includes:
[0099] S61: Collect ambient background sound features and ambient background vibration features at intervals t, and obtain the background acoustic feature matrix. and background vibration characteristic matrix ;
[0100] S62: Background acoustic feature matrix and background vibration characteristic matrix Concatenate into background observation vector ,
[0101] ;
[0102] S63: Calculate the background sparse coefficient vector ;
[0103] S64: Update the reference matrix D and obtain the updated reference matrix D';
[0104] ε is a constant;
[0105] S65: Substitute the updated reference matrix D' into step S5 for calculation.
[0106] In step S61, the time t is 5-10s. In this embodiment, the current environmental background features are collected every 5-10s to correct the reference matrix D.
[0107] Of course, in order to match the format, dimensions, etc. of the reference matrix D, the background acoustic feature matrix is collected and acquired. and background vibration characteristic matrix Background observation vectors are obtained by splicing. Obtain the background sparse coefficient vector The process is similar to the steps described above, and will not be repeated here.
[0108] In step S64, the formula for obtaining the updated reference matrix D' is as follows: The eigencomponents representing the background interference in the original reference matrix D; Background sparsity coefficient The squared L2 norm serves as a normalization parameter, preventing excessively large update amplitudes and ensuring smoother changes to the reference matrix. ε is a very small regularization parameter, preventing the denominator from being zero and improving computational stability. In this embodiment... Step S64 is used to subtract the components in the reference matrix D corresponding to the background interference from the original reference matrix D to obtain the updated reference matrix D', which is used to filter the background interference of the current environment.
[0109] Of course, the present invention also provides a sound source localization device based on acoustic and vibration heterogeneous sensing, the localization device comprising:
[0110] A heterogeneous sensing array, comprising M acoustic sensors and N vibration sensors, wherein the heterogeneous sensing array is arranged in an asymmetric topological layout;
[0111] The signal acquisition module is used to synchronously acquire acoustic signals. and vibration signals ,in, , ;
[0112] The feature extraction module is used for acoustic signal processing. Extract Mel-frequency cepstral coefficient features and obtain K-dimensional features for each of the M acoustic sensors. For vibration signals Extract wavelet packet energy features and obtain Q-dimensional features from N vibration sensors. ;
[0113] The feature stitching module is used to stitch together the K-dimensional features of M acoustic sensors. The acoustic signals are spliced into a large M×K dimensional matrix A, which incorporates the Q-dimensional features of N vibration sensors. A large N×Q dimensional vibration signal matrix V is concatenated, and the large acoustic signal matrix A and the large vibration signal matrix V are concatenated to form the observation vector X. ;
[0114] The sparse fusion localization module is used to construct a reference matrix D for the reference sound source, and to retrieve and obtain the azimuth angle of the observation vector X by substituting it into the reference matrix D. .
[0115] The sound source localization device also includes an environment adaptive calibration module, which is used to collect environmental background sound characteristics and environmental background vibration characteristics, and update the reference matrix D in real time.
[0116] The working method of the sound source localization device corresponds to the sound source localization method described above, and will not be repeated here.
[0117] In summary, this technical solution avoids traditional algorithms such as TDOA (Time Difference of Arrival) or PHAT (Phase Transformation) and adopts feature-level sparse fusion, which eliminates the need to calculate the signal arrival time difference and achieves a synchronization accuracy of ≤1ms, greatly reducing the requirements for multi-channel synchronization.
[0118] Furthermore, this technical solution exhibits strong anti-interference capabilities. While acoustic characteristics may be affected by ambient air noise, the carrier vibration characteristics, directly collected from the carrier surface, are virtually unaffected by air noise and can still stably reflect the sound source location. When the carrier itself experiences vibration interference, the vibration characteristics introduce noise, but the acoustic characteristics can serve as a reference, filtering out false peaks caused by vibration interference. It can achieve a positioning error of ≤3° in environments with a signal-to-noise ratio ≥5dB and a reverberation time ≤0.5s. Moreover, this technical solution achieves full-angle coverage, clearly defining angles as "0°-359° (1 / step)," solving the problems of large angle intervals (e.g., ≥5°) or coverage of only partial angles (e.g., only 0°-180°) in existing technologies.
[0119] Furthermore, this technical solution employs an asymmetric heterogeneous array, enhancing directional resolution through spatial complementarity. Therefore, this technical solution is applicable to a wide range of scenarios, including indoor robots, in-vehicle voice interaction, and equipment fault location, eliminating the need for complex hardware synchronization modules and significantly reducing costs.
[0120] Furthermore, this technical solution includes an environmental adaptive calibration module, which updates the reference matrix D in real time based on the ambient sound of the detection environment, allowing the reference matrix D to undergo environmental adaptive calibration. This prevents the original reference matrix D from misinterpreting background interference as sound source features when the background interference changes at the detection site, leading to positioning drift or errors. Therefore, the positioning in this technical solution is more accurate and has less error.
[0121] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0122] The detailed descriptions listed above are merely specific descriptions of feasible implementations of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent implementations or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.
Claims
1. A sound source localization method based on acoustic and vibration heterogeneous sensing, characterized in that, The method includes: S1: Arrange M acoustic sensors and N vibration sensors into an asymmetric heterogeneous array and fix them to the surface of the carrier; S2: Synchronous acquisition of acoustic signals and vibration signals ,in, , ; S3: Acoustic signal Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. And the K-dimensional feature matrix of M acoustic sensors The signals are spliced together to form a large M×K dimensional acoustic signal matrix A; Vibration signals Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. And the Q-dimensional feature matrix of N vibration sensors A large matrix V of N×Q dimensional vibration signals is spliced together; S4: Concatenate the large acoustic signal matrix A and the large vibration signal matrix V to form the observation vector X. ; S5: Compare the observation vector X with the reference matrix D of the reference sound source to obtain the azimuth angle of the observation vector X. .
2. The sound source localization method according to claim 1, characterized in that, The S1 step is preceded by a preprocessing step S0, which includes: Acoustic reference feature matrices of the reference sound sources were collected and obtained respectively. and vibration reference characteristic matrix ,in The range is from 0° to 359°, with a step size of 1°, and a reference matrix D of the reference sound source is constructed. 。 3. The sound source localization method according to claim 1, characterized in that, In step S3, "the acoustic signal" Mel-frequency cepstral coefficients are extracted to obtain the K-dimensional feature matrices of the M acoustic sensors. "include: acoustic signals After pre-emphasis, framing, Mel filtering, and DCT transformation, the K-dimensional feature matrix of the M acoustic sensors is obtained. .
4. The sound source localization method according to claim 3, characterized in that, K∈{12,13,14,15,16}。 5. The sound source localization method according to claim 1, characterized in that, In step S3, "for the vibration signal" Wavelet packet energy features are extracted to obtain the Q-dimensional feature matrices of N vibration sensors. "middle, , .
6. The sound source localization method according to claim 1, characterized in that, In step S3, "the K-dimensional feature matrix of the M acoustic sensors is..." The splicing method in "A" which is spliced into an M×K dimensional acoustic signal matrix is as follows: ; "The Q-dimensional feature matrix of N vibration sensors" The splicing method in the splicing of the N×Q dimensional vibration signal large matrix V is as follows: 。 7. The sound source localization method according to claim 1, characterized in that, Step S5 includes: S51: Substitute the observation vector X and the reference matrix D of the reference sound source into the L1 norm regularized least squares optimization formula to calculate the sparse coefficient vector. , , in, The regularization parameter is and ; S52: For sparse coefficient vectors Indexing to find the sparse coefficient vector The angle corresponding to the element with the largest median value Let X be the azimuth angle of the observation vector X. , 。 8. The sound source localization method according to claim 1, characterized in that, Following step S5, an environmental calibration step S6 is further included, which includes: S61: Collect ambient background sound features and ambient background vibration features at intervals t, and obtain the background acoustic feature matrix. and background vibration characteristic matrix ; S62: Background acoustic feature matrix and background vibration characteristic matrix Concatenate into background observation vector , ; S63: Calculate the background sparse coefficient vector ; S64: Update the reference matrix D and obtain the updated reference matrix D'; , It is a constant; S65: Substitute the updated reference matrix D' into step S5 for calculation.
9. A sound source localization device based on acoustic and vibration heterogeneous sensing, characterized in that, The positioning device includes: A heterogeneous sensing array, comprising M acoustic sensors and N vibration sensors, wherein the heterogeneous sensing array is arranged in an asymmetric topological layout; The signal acquisition module is used to synchronously acquire acoustic signals. and vibration signals ,in, , ; The feature extraction module is used for acoustic signal processing. Extract Mel-frequency cepstral coefficient features and obtain K-dimensional features for each of the M acoustic sensors. For vibration signals Extract wavelet packet energy features and obtain Q-dimensional features from N vibration sensors. ; The feature stitching module is used to stitch together the K-dimensional features of M acoustic sensors. The acoustic signals are spliced into a large M×K dimensional matrix A, which incorporates the Q-dimensional features of N vibration sensors. A large N×Q dimensional vibration signal matrix V is concatenated, and the large acoustic signal matrix A and the large vibration signal matrix V are concatenated to form the observation vector X. ; The sparse fusion localization module is used to construct the reference matrix D of the reference sound source, and to retrieve and obtain the azimuth angle θ of the observation vector X by substituting the observation vector X into the reference matrix D.
10. The sound source localization device according to claim 9, characterized in that, The sound source localization device also includes an environment adaptive calibration module, which is used to collect environmental background sound characteristics and environmental background vibration characteristics, and update the reference matrix D in real time.