A sound source localization method based on collaborative networking of multi-microphone arrays

Through the coordinated networking of multi-microphone arrays, a two-dimensional coordinate system is established and azimuth angle and time difference is calculated, which solves the problem of low accuracy of traditional sound source positioning systems in far-field scenarios, and realizes accurate positioning of sound sources at different locations.

CN120065115BActive Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510542196.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-08
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Traditional sound source positioning systems have dual constraints on the complexity of the acoustic environment and the physical characteristics of the array in far-field scenarios, resulting in low positioning accuracy, especially in large conference rooms that cannot effectively distinguish sound sources at different locations.

Method used

The method of collaborative networking of multiple microphone arrays is adopted to establish a two-dimensional coordinate system of microphone arrays, calculate the azimuth angle of the sound source relative to each microphone array, and calculate the position of the sound source in the two-dimensional coordinate system by combining the time difference and trigonometric functions.

Benefits of technology

It improves the accuracy of sound source positioning, can effectively distinguish sound sources in different locations, and is suitable for rectangular conference scenes such as large conference rooms and lecture halls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065115B_ABST
    Figure CN120065115B_ABST
Patent Text Reader

Abstract

The present invention discloses a sound source localization method based on collaborative networking of multiple microphone arrays, which relates to the technical field of sound source localization. Aiming at the defect that it is difficult to accurately locate the position of the sound source at present, the adopted solution includes: establishing a two-dimensional coordinate system of the microphone arrays, deploying multiple microphone arrays with the same orientation at equal intervals along the x-axis, and the midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system; the multiple microphone arrays synchronously collect the sound source signals, obtain multi-channel audio data, calculate the time difference of the sound source signals arriving at each microphone array, and estimate the azimuth angle of the sound source relative to each microphone array; calculate the absolute value of the difference between the azimuth angle and the positive or negative half-axis of the y-axis, select the microphone arrays closest to and farthest from the sound source, and calculate the distance between the two selected microphone arrays; calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system based on the distance value and the azimuth angle. The present invention is used to improve the accuracy of sound source localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound source localization, and specifically to a sound source localization method based on collaborative networking of multi-microphone arrays. Background Art

[0002] The performance bottlenecks faced by traditional sound source localization systems in far-field scenarios are mainly reflected in the dual constraints of the complexity of the acoustic environment and the physical characteristics of the array. In a typical rectangular meeting scenario, when the distance between the sound source and the microphone array exceeds 2 meters, the error caused by the approximation of the spherical wavefront will significantly reduce the spatial resolution ability of the system. At this time, the air absorption during the sound wave propagation process (attenuation proportional to the square of the frequency) and the reverberation field formed by boundary reflections (RT60 often exceeds 300 ms) will cause a sharp drop in the signal-to-noise ratio of the received signal. Especially for high-frequency components (above 4 kHz), the attenuation can reach 15 dB / distance doubling, which poses a severe challenge to localization algorithms based on phase difference or spectral estimation.

[0003] From the perspective of array processing, the aperture of a conventional uniform linear array or circular array is limited by the size constraints of the meeting room (usually not exceeding 3 meters). According to the Rayleigh resolution criterion, the azimuth resolution limit is about λ / (2L) (λ is the wavelength, and L is the effective aperture of the array). When the sound source is in the far field, the difference in the incident angles of adjacent speakers (spacing more than 1.5 meters) reaching the array may be less than 0.3°, which is close to the theoretical resolution limit in the 8 kHz frequency band. More seriously, the problem of spatial ambiguity introduced by the multi-row seat layout - the sound sources in the front row and the back row may have the same horizontal azimuth angle but there are significant differences in the vertical pitch angle, while traditional two-dimensional DOA estimation algorithms can only obtain the horizontal plane projection angle, resulting in topological overlap of sound sources at different positions in the localization results.

[0004] Although the existing technical improvement schemes can locally improve the performance, they have significant application limitations. Constructing a superarray by increasing the number of microphones can improve the spatial sampling density, but the resulting computational complexity increases exponentially (when the number of arrays is N, the dimension of the covariance matrix reaches O(N²)), and the problem of coherent sources caused by indoor multipath effects still cannot be fundamentally solved. Using FPGA / DSP hardware acceleration can achieve real-time processing, but the equipment cost and maintenance complexity make it difficult to popularize in ordinary meeting rooms. More importantly, these technical improvements do not touch on the inherent defects of the array geometry: the sensitivity of the linear array to the end-fire direction decays by 6 dB / octave, the circular array has phase ambiguity in the axial direction, and the planar array is limited by the height of the meeting room, resulting in the lack of vertical dimension observation. This perspective blind spot caused by physical space constraints makes it difficult to break through the ceiling of localization accuracy simply by algorithm optimization or hardware upgrade. Summary of the Invention

[0005] In view of the defect that sound sources at different positions may correspond to the same azimuth angle estimate, making it impossible to effectively distinguish sound sources at different positions, the present invention provides a sound source localization method based on collaborative networking of multiple microphone arrays to improve the localization accuracy of sound sources.

[0006] The technical solution adopted by a sound source localization method based on collaborative networking of multiple microphone arrays of the present invention to solve the above technical problems is as follows:

[0007] A sound source localization method based on collaborative networking of multiple microphone arrays includes the following steps:

[0008] S1. Establish a two-dimensional coordinate system for the microphone arrays, deploy multiple microphone arrays with the same orientation at equal intervals along the x-axis, and the midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system;

[0009] S2. Multiple microphone arrays synchronously collect sound source signals to obtain multi-channel audio data, calculate the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data, and combine the two-dimensional coordinate system of the microphone arrays to estimate the azimuth angle of the sound source relative to each microphone array;

[0010] S3. Calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the positive or negative half-axis of the y-axis, select the one with the smallest absolute value as the microphone array closest to the sound source, select the one with the largest absolute value as the microphone array farthest from the sound source, and calculate the distance between the two microphone arrays closest to and farthest from the sound source;

[0011] S4. Based on the distance between the two microphone arrays closest to and farthest from the sound source, the azimuth angle of the sound source relative to the closest microphone array, and the azimuth angle of the sound source relative to the farthest microphone array, calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system through trigonometric functions.

[0012] Optionally, the positive half-axis of the x-axis of the two-dimensional coordinate system of the microphone arrays is 0°, the negative half-axis of the x-axis is 180°, the positive half-axis of the y-axis is 90°, the negative half-axis of the y-axis is 270°, the azimuth angle is the angle starting from the positive half-axis of the x-axis and rotating counterclockwise to the sound source direction, and the azimuth angle value is greater than or equal to 0° and less than 360°.

[0013] Optionally, when performing step S2, after multiple microphone arrays obtain multi-channel audio data, they send the multi-channel audio data to the cloud server in real time through their own WiFi function;

[0014] The cloud server calculates the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone arrays to estimate the azimuth angle of the sound source relative to each microphone array by using the beamforming algorithm.

[0015] Further optionally, the cloud server calculates the time difference of the sound source signal arriving at each microphone array through multi-channel audio data, combines the two-dimensional coordinate system of the microphone array, and uses the beamforming algorithm to estimate the azimuth angle of the sound source relative to each microphone array. This process specifically includes:

[0016] S2.1. Synchronously calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use the adaptive filtering algorithm to remove background noise, and perform signal enhancement through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing;

[0017] S2.2. Calculate the time difference of the sound source signal arriving at any two microphone arrays through the generalized cross-correlation-phase transform algorithm, convert the time difference into a distance difference by combining the speed of sound, and generate multiple sets of TDOA data to characterize the relative distance difference between the sound source and each microphone array;

[0018] S2.3. Based on the two-dimensional coordinate system of the microphone array, determine the positional relationship of each microphone array to provide a geometric reference for azimuth angle calculation;

[0019] S2.4. Perform time delay compensation on the signals of each microphone array, use the delay-sum beamforming algorithm to weight and sum the compensated signals, and generate a 0°-360° omnidirectional response spectrum. The peak of the spectrum corresponds to the potential azimuth of the sound source;

[0020] S2.5. Identify the main lobe peak direction in the beamforming response spectrum as the initial azimuth angle estimate value, and eliminate the side lobe false peaks through TDOA distance difference constraints; use the extended Kalman filter algorithm to iteratively optimize the sequential azimuth angle estimate value to suppress the estimation errors caused by environmental noise and multipath effects, and finally output the accurate azimuth angles corresponding to each microphone array.

[0021] Preferably, the microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module, and a CPU; multiple microphone arrays synchronously collect the sound source signal, and after obtaining multi-channel audio data, send the multi-channel audio data to the cloud server in real time through the WiFi module;

[0022] Introduce a portable forwarding device, which includes a camera, a Bluetooth module, a WiFi module, a 4G module, and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

[0023] Further optionally, perform step S3 to calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the y-axis. In this process:

[0024] When the sound source is on the positive half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90°, that is, calculate |θ - 90°|. The smaller |θ - 90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 90°| is, the farther the corresponding microphone array is from the sound source;

[0025] When the sound source is on the negative half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270°, that is, calculate |θ - 270°|. The smaller |θ - 270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 270°| is, the farther the corresponding microphone array is from the sound source.

[0026] Further optionally, the trigonometric function expressions are as follows:

[0027] y / x = tan(θ1 - 180°),

[0028] -y / (2L - x) = tan(θ2 - 180°),

[0029] In the formula, θ1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the distance between adjacent two microphone arrays, and x and y represent the abscissa and ordinate of the sound source in the two-dimensional microphone array coordinate system.

[0030] Preferably, at least three microphone arrays with the same orientation are deployed equidistantly along the x-axis direction, and the distance between adjacent microphone arrays does not exceed 3 meters.

[0031] A sound source localization method based on multi-microphone array collaborative networking according to the present invention has the beneficial effects compared with the prior art as follows:

[0032] The present invention can accurately locate the sound source position, make up for the defect that sound sources at different positions may correspond to the same azimuth angle estimation value, and thus cannot effectively distinguish sound sources at different positions, and is particularly suitable for accurate sound source localization in rectangular meeting scenarios such as large conference rooms and lecture halls. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Attached Figure 1 is the flowchart of the method according to an embodiment of the present invention;

[0034] Attached Figure 2 is the structural connection diagram related to the method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] To make the technical solution, the technical problems solved and the technical effects of the present invention clearer and more understandable, the following describes the technical solution of the present invention clearly and completely in combination with specific embodiments.

[0036] Embodiment:

[0037] Refer to the attached Figure 1 , this embodiment proposes a sound source localization method based on collaborative networking of multiple microphone arrays, which includes the following steps:

[0038] S1. Establish a two-dimensional coordinate system for the microphone array, deploy multiple microphone arrays with the same orientation at equal intervals along the x-axis, and the midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system.

[0039] Specifically, the positive half-axis of the x-axis of the two-dimensional coordinate system of the microphone array is 0°, the negative half-axis of the x-axis is 180°, the positive half-axis of the y-axis is 90°, the negative half-axis of the y-axis is 270°, and the azimuth angle is the angle starting from the positive half-axis of the x-axis and rotating counterclockwise to the sound source direction. The azimuth angle value is greater than or equal to 0° and less than 360°.

[0040] S2. Multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through their own WiFi functions.

[0041] Refer to the attached Figure 2 , the microphone array in this step includes a multi-channel microphone, a Bluetooth module, a WiFi module, and a CPU; multiple microphone arrays synchronously collect sound source signals, and after obtaining multi-channel audio data, send the multi-channel audio data to the cloud server in real time through the WiFi module; a portable forwarding device is introduced, and the portable forwarding device includes a camera, a Bluetooth module, a WiFi module, a 4G module, and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

[0042] The cloud server calculates the time difference of the sound source signal reaching each microphone array through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array by using the beamforming algorithm. This process specifically includes:

[0043] S2.1. Synchronously calibrate the audio signals transmitted by each microphone array based on timestamp alignment, remove background noise by using an adaptive filtering algorithm, and enhance the signal through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing;

[0044] S2.2. Calculate the time difference of arrival of the sound source signal at any two microphone arrays through the generalized cross-correlation-phase transformation algorithm, convert the time difference into a distance difference in combination with the speed of sound, and generate multiple sets of TDOA data to characterize the relative distance differences between the sound source and each microphone array;

[0045] S2.3. Based on the two-dimensional coordinate system of the microphone arrays, determine the positional relationship of each microphone array to provide a geometric reference for azimuth calculation;

[0046] S2.4. Perform time-delay compensation on the signals of each microphone array, and use the delay-and-sum beamforming algorithm to weighted sum the compensated signals to generate a full-direction response map from 0° to 360°. The peak of the map corresponds to the potential azimuth of the sound source;

[0047] S2.5. Identify the main lobe peak direction in the beamforming response map as the initial azimuth estimate value, and eliminate the sidelobe false peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the sequential azimuth estimate value to suppress the estimation errors caused by environmental noise and multipath effects, and finally output the accurate azimuth angles corresponding to each microphone array.

[0048] S3. Calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the positive or negative half-axis of the y-axis, select the one with the smallest absolute value as the microphone array closest to the sound source, select the one with the largest absolute value as the microphone array farthest from the sound source, and calculate the distance between the two microphone arrays closest and farthest from the sound source; in this process:

[0049] When the sound source is on one side of the positive half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90°, that is, calculate |θ - 90°|. The smaller |θ - 90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 90°| is, the farther the corresponding microphone array is from the sound source;

[0050] When the sound source is on one side of the negative half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270°, that is, calculate |θ - 270°|. The smaller |θ - 270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 270°| is, the farther the corresponding microphone array is from the sound source.

[0051] S4. Based on the distance between the two microphone arrays closest and farthest from the sound source, the azimuth angle of the sound source relative to the closest microphone array, and the azimuth angle of the sound source relative to the farthest microphone array, calculate the coordinate values of the sound source in the two-dimensional microphone array coordinate system through trigonometric functions.

[0052] The trigonometric function expressions are as follows:

[0053] y / x = tan(θ1 - 180°),

[0054] -y / (2L - x) = tan(θ2 - 180°),

[0055] In the formula, θ1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between two adjacent microphone arrays, and x and y represent the abscissa and ordinate of the sound source in the two-dimensional microphone array coordinate system.

[0056] Taking "deploying five microphone arrays with the same orientation at equal intervals along the x-axis direction, and the spacing between adjacent microphone arrays is 2 meters" as an example, the implementation process of the above steps is specifically described as follows:

[0057] (1) The five microphone arrays from left to right are sequentially marked as A, B, C, D, and E. Taking microphone array C as the origin of the coordinate system, then A(-4, 0), B(-2, 0), C(0, 0), D(2, 0), E(4, 0).

[0058] (i) Assume that the sound source S1 is located on the positive y-axis, that is, in the first quadrant or the second quadrant of the coordinate system; microphone arrays A, B, C, D, and E respectively collect the sound source S1, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through their own WiFi functions. The cloud server calculates the time difference between the sound source S1 reaching microphone arrays A, B, C, D, and E through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array. Assume that the azimuth angles obtained are θ A , θ B , θ C , θ D and θ E .

[0059] Further calculate the absolute value of the difference between the azimuth angle and the positive y-axis, that is, |θ - 90°|, where θ takes values of θ A , θ B , θ C , θ D and θ E。 Assume that the value of |θ A - 90°| is the largest, and the value of |θ C - 90°| is the smallest. Then microphone array A is used as the microphone array farthest from the sound source, microphone array C is used as the microphone array closest to the sound source, and the distance between microphone array A and microphone array C is 4 meters.

[0060] Assume that θ A specifically takes a value of 80°, θ CIf the specific value is 45°, then the coordinate values of the sound source S in the two-dimensional microphone array coordinate system are obtained through trigonometric calculations.

[0061] y / x = tan(θ C -180°),

[0062] -y / (2L - x) = tan(θ A -180°),

[0063] In the formula, θ1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between two adjacent microphone arrays, and x and y represent the abscissa and ordinate of the sound source S1 in the two-dimensional microphone array coordinate system. The specific calculated coordinates are (0.86, 4.86).

[0064] (ii) Assume that the sound source S2 is located on the negative half-axis of the y-axis, that is, in the third or fourth quadrant of the coordinate system; the microphone arrays A, B, C, D, and E respectively collect the sound source S2, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through its own WiFi function. The cloud server calculates the time difference between the sound source S2 reaching the microphone arrays A, B, C, D, and E through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array. Assume that the azimuth angles θ A , θ B , θ C , θ D and θ E are obtained.

[0065] Further calculate the absolute value of the difference between the azimuth angle and the positive half-axis of the y-axis, that is, |θ – 270°|, where θ takes the values of θ A , θ B , θ C , θ D and θ E。 Assume that the value of |θ A – 90°| is the largest, and the value of |θ C – 270°| is the smallest. Then the microphone array E is used as the microphone array farthest from the sound source, the microphone array C is used as the microphone array closest to the sound source, and the distance between the microphone array E and the microphone array C is 4 meters.

[0066] Assume that θ E takes the specific value of 225°, and θ C takes the specific value of 260°. Then the coordinate values of the sound source S2 in the two-dimensional microphone array coordinate system are obtained through trigonometric calculations.

[0067] y / x = tan(θ C -180°),

[0068] -y / (2L - x)=tan(θ E -180°),

[0069] In the formula, θ1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between two adjacent microphone arrays, and x and y represent the abscissa and ordinate of the sound source S2 in the two-dimensional microphone array coordinate system, and the specific calculated coordinates are (-0.86, -4.86).

[0070] In summary, by using a sound source localization method based on multi-microphone array collaborative networking of the present invention, by establishing a two-dimensional coordinate system of the microphone array, calculating the azimuth angles of the sound source and different microphone arrays, calculating the absolute values of the differences between the azimuth angles and the nearest and farthest microphone arrays, and then calculating the coordinate values of the sound source in the two-dimensional microphone array coordinate system, this process can improve the sound source localization accuracy and make up for the defect that sound sources at different positions may correspond to the same azimuth angle estimation value, so that sound sources at different positions cannot be effectively distinguished.

[0071] The above application of specific examples has elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art in the technical field of the present invention without departing from the principle of the present invention shall fall within the scope of the patent protection of the present invention.

Claims

1. A sound source localization method based on collaborative networking of multi-microphone arrays, characterized in that, It includes the following steps: S1. Establish a two-dimensional coordinate system for the microphone array. Deploy multiple microphone arrays with the same orientation at equal intervals along the x-axis. The midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system; S2. The multiple microphone arrays synchronously collect the sound source signals to obtain multi-channel audio data. Calculate the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data. Combine the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array; S3. Calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the positive or negative half-axis of the y-axis. Select the one with the smallest absolute value as the microphone array closest to the sound source, and select the one with the largest absolute value as the microphone array farthest from the sound source. Calculate the distance between the two microphone arrays closest and farthest from the sound source; S4. Based on the distance between the two microphone arrays closest and farthest from the sound source, the azimuth angle of the sound source relative to the closest microphone array, and the azimuth angle of the sound source relative to the farthest microphone array, calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system through trigonometric functions.

2. The method for sound source localization based on collaborative networking of multi-microphone arrays according to claim 1, wherein, The positive half-axis of the x-axis of the two-dimensional coordinate system of the microphone array is 0°, the negative half-axis of the x-axis is 180°, the positive half-axis of the y-axis is 90°, and the negative half-axis of the y-axis is 270°. The azimuth angle is the angle starting from the positive half-axis of the x-axis and rotating counterclockwise to the sound source direction. The azimuth angle value is greater than or equal to 0° and less than 360°.

3. The method for sound source localization based on multi-microphone array collaborative networking according to claim 2, wherein Execute step S2. After the multiple microphone arrays obtain the multi-channel audio data, they send the multi-channel audio data to the cloud server in real time through their own WiFi function; The cloud server calculates the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data. Combine the two-dimensional coordinate system of the microphone array and use the beamforming algorithm to estimate the azimuth angle of the sound source relative to each microphone array.

4. A sound source localization method based on multi-microphone array collaborative networking according to claim 3, characterized in that The cloud server calculates the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data. Combine the two-dimensional coordinate system of the microphone array and use the beamforming algorithm to estimate the azimuth angle of the sound source relative to each microphone array. This process specifically includes: S2.

1. Perform synchronous calibration on the audio signals transmitted by each microphone array based on timestamp alignment. Use the adaptive filtering algorithm to remove background noise, and perform signal enhancement through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing; S2.

2. Through the generalized cross-correlation-phase transform algorithm, calculate the time difference of the sound source signal arriving at any two microphone arrays. Convert the time difference into a distance difference in combination with the speed of sound to generate multiple groups of TDOA data, which characterize the relative distance difference between the sound source and each microphone array; S2.

3. Based on the two-dimensional coordinate system of the microphone array, determine the positional relationship of each microphone array to provide a geometric reference for azimuth angle calculation; S2.

4. Perform time delay compensation on the signal of each microphone array. Use the delay-sum beamforming algorithm to weight and sum the compensated signals to generate a full-direction response spectrum from 0° to 360°. The peak of the spectrum corresponds to the potential azimuth of the sound source. S2.

5. Identify the main lobe peak direction in the beamforming response pattern as the initial azimuth angle estimate value, and eliminate the sidelobe false peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time-series azimuth angle estimate value, suppress the estimation errors caused by environmental noise and multipath effects, and finally output the accurate azimuth angles corresponding to each microphone array.

5. A sound source localization method based on multi-microphone array collaborative networking according to claim 3, characterized in that, The microphone array includes multi-channel microphones, a Bluetooth module, a WiFi module, and a CPU; multiple microphone arrays synchronously collect sound source signals, and after obtaining multi-channel audio data, the multi-channel audio data is sent to the cloud server in real time through the WiFi module; Introduce a portable forwarding device, which includes a camera, a Bluetooth module, a WiFi module, a 4G module, and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

6. The method for sound source localization based on collaborative networking of multiple microphone arrays according to claim 3, wherein Execute step S3, calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the y-axis. In this process: When the sound source is on one side of the positive half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90°, that is, calculate |θ - 90°|. The smaller |θ - 90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 90°| is, the farther the corresponding microphone array is from the sound source; When the sound source is on one side of the negative half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270°, that is, calculate |θ - 270°|. The smaller |θ - 270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 270°| is, the farther the corresponding microphone array is from the sound source.

7. A sound source localization method based on multi-microphone array collaborative networking according to claim 6, characterized in that The trigonometric function expression is as follows: y / x = tan(θ1 - 180°), -y / (2L - x) = tan(θ2 - 180°), In the formula, θ1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the distance between adjacent two microphone arrays, and x and y represent the abscissa and ordinate of the sound source in the two-dimensional microphone array coordinate system.

8. A sound source localization method based on multi-microphone array collaborative networking according to claim 7, characterized in that, Deploy at least three microphone arrays with the same orientation at equal intervals along the x-axis direction, and the distance between adjacent microphone arrays does not exceed 3 meters.

Citation Information

Patent Citations

  • Method for positioning sound source by using microphone array

    CN102033223A

  • Hybrid deployment method for heterogeneous multi-core chip application program

    CN119883650A