Sound source localization method based on multi-microphone array collaborative networking

By using the method of collaborative networking of multi-microphone arrays in the sound source positioning system, a two-dimensional coordinate system is established and the sound source coordinates are calculated, which solves the problems of low positioning accuracy and difficulty in distinguishing sound source in traditional systems, and achieves higher accuracy sound source positioning.

CN120065115AActive Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510542196.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In far-field scenarios, traditional sound source positioning systems have low positioning accuracy due to the complexity of the acoustic environment and the physical characteristics of the array. In particular, sound sources at different locations may correspond to the same azimuth estimation value and cannot be effectively distinguished.

Method used

The sound source positioning method based on the coordinated networking of multiple microphone arrays is adopted. By establishing a two-dimensional coordinate system of the microphone array, the azimuth angle between the sound source and different microphone arrays are calculated, the two microphone arrays closest and farthest away from the sound source are selected, and the coordinate value of the sound source in the two-dimensional microphone array coordinate system is calculated using trigonometric function.

Benefits of technology

It improves the accuracy of sound source positioning, can effectively distinguish sound sources at different locations, and is suitable for rectangular conference scenes such as large conference rooms and lecture halls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065115A_ABST
    Figure CN120065115A_ABST
Patent Text Reader

Abstract

The invention discloses a sound source localization method based on multi-microphone array collaborative networking, relates to the technical field of sound source localization, and aims to overcome the defect that the position of a sound source is difficult to locate accurately at present, and adopts the scheme that the method comprises the following steps: establishing a microphone array two-dimensional coordinate system, deploying a plurality of microphone arrays with the same orientation along an x axis at equal intervals, the midpoint of the connecting line of the plurality of microphone arrays is the origin of the coordinate system; the method comprises the following steps: synchronously acquiring sound source signals by a plurality of microphone arrays, acquiring multichannel audio data, calculating the time difference of the sound source signals reaching each microphone array, and estimating the azimuth angle of a sound source relative to each microphone array; calculating the absolute value of the difference between the azimuth angle and the positive half axis or the negative half axis of the y axis, selecting the microphone arrays closest to and farthest from the sound source, and calculating the distance between the two microphone arrays; and calculating the coordinate value of the sound source in the two-dimensional microphone array coordinate system based on the distance value and the azimuth angle. The method is used for improving the sound source positioning precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound source localization, and specifically to a sound source localization method based on multi-microphone array collaborative networking. Background Art

[0002] The performance bottlenecks faced by traditional sound source localization systems in far-field scenarios are mainly reflected in the dual constraints of the complexity of the acoustic environment and the physical characteristics of the array. In a typical rectangular meeting scenario, when the distance between the sound source and the microphone array exceeds 2 meters, the error caused by the spherical wavefront approximation will significantly reduce the spatial resolution ability of the system. At this time, the air absorption during the sound wave propagation process (attenuation proportional to the square of the frequency) and the reverberation field formed by the boundary reflection (RT60 often exceeds 300 ms) will cause the signal-to-noise ratio of the received signal to drop sharply. In particular, the attenuation of high-frequency components (above 4 kHz) can reach 15 dB / decade, which poses a severe challenge to localization algorithms based on phase difference or spectral estimation.

[0003] From the perspective of array processing, the aperture of a conventional uniform linear array or circular array is limited by the size constraints of the meeting room (usually not exceeding 3 meters). According to the Rayleigh resolution criterion, the azimuth resolution limit is approximately λ / (2L) (λ is the wavelength, and L is the effective aperture of the array). When the sound source is in the far field, the difference in the incident angles of adjacent speakers (spacing more than 1.5 meters) arriving at the array may be less than 0.3°, which is close to the theoretical resolution limit in the 8 kHz frequency band. More seriously, the problem of spatial ambiguity introduced by the multi-row seating layout - the sound sources in the front row and the back row may have the same horizontal azimuth angle but there are significant differences in the vertical pitch angle, while traditional two-dimensional DOA estimation algorithms can only obtain the horizontal plane projection angle, resulting in topological overlap of sound sources at different positions in the localization results.

[0004] Although the existing technical improvement schemes can locally improve the performance, they have significant application limitations. Constructing a superarray by increasing the number of microphones can improve the spatial sampling density, but the resulting computational complexity increases exponentially (when the number of arrays is N, the dimension of the covariance matrix reaches O(N²)), and the problem of coherent sources caused by indoor multipath effects still cannot be fundamentally solved. Using FPGA / DSP hardware acceleration can achieve real-time processing, but the equipment cost and maintenance complexity make it difficult to popularize in ordinary meeting rooms. More importantly, these technical improvements do not touch on the inherent defects of the array geometry: the sensitivity of the linear array to the end-fire direction decays by 6 dB / octave, the circular array has phase ambiguity in the axial direction, and the planar array is limited by the height of the meeting room, resulting in a lack of vertical dimension observation. This perspective blind spot caused by physical space constraints makes it difficult to break through the ceiling of localization accuracy simply by algorithm optimization or hardware upgrade. Summary of the Invention

[0005] In view of the defect that sound sources at different positions may correspond to the same azimuth estimation value, making it impossible to effectively distinguish sound sources at different positions, the present invention provides a sound source localization method based on collaborative networking of multiple microphone arrays to improve the localization accuracy of sound sources.

[0006] For the sound source localization method based on collaborative networking of multiple microphone arrays of the present invention, the technical solution adopted to solve the above technical problems is as follows: A sound source localization method based on collaborative networking of multiple microphone arrays includes the following steps: S1. Establish a two-dimensional coordinate system of the microphone array, deploy multiple microphone arrays with the same orientation at equal intervals along the x-axis, and the midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system; S2. Multiple microphone arrays synchronously collect sound source signals to obtain multi-channel audio data, calculate the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data, and combine the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array; S3. Calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the positive or negative half-axis of the y-axis, select the one with the smallest absolute value as the microphone array closest to the sound source, select the one with the largest absolute value as the microphone array farthest from the sound source, and calculate the distance between the two microphone arrays closest and farthest from the sound source; S4. Based on the distance between the two microphone arrays closest and farthest from the sound source, the azimuth angle of the sound source relative to the closest microphone array, and the azimuth angle of the sound source relative to the farthest microphone array, calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system through trigonometric functions.

[0007] Optionally, the positive half-axis of the x-axis of the two-dimensional coordinate system of the microphone array is 0°, the negative half-axis of the x-axis is 180°, the positive half-axis of the y-axis is 90°, the negative half-axis of the y-axis is 270°, the azimuth angle is the angle starting from the positive half-axis of the x-axis and rotating counterclockwise to the sound source direction, and the azimuth angle value is greater than or equal to 0° and less than 360°.

[0008] Optionally, when performing step S2, after multiple microphone arrays obtain multi-channel audio data, they send the multi-channel audio data to the cloud server in real time through their own WiFi function; The cloud server calculates the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone array to estimate the azimuth angle of the sound source relative to each microphone array by using the beamforming algorithm.

[0009] Further optionally, the cloud server calculates the time difference of the sound source signal arriving at each microphone array through multi-channel audio data, combines the two-dimensional coordinate system of the microphone array, and uses the beamforming algorithm to estimate the azimuth angle of the sound source relative to each microphone array. This process specifically includes: S2.1. Synchronously calibrate the audio signals transmitted by each microphone array based on timestamp alignment, remove background noise using an adaptive filtering algorithm, and enhance the signals through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing; S2.2. Calculate the time difference of the sound source signal arriving at any two microphone arrays through the generalized cross-correlation-phase transform algorithm, convert the time difference into a distance difference by combining the speed of sound, and generate multiple sets of TDOA data to characterize the relative distance differences from the sound source to each microphone array; S2.3. Based on the two-dimensional coordinate system of the microphone array, determine the positional relationship of each microphone array to provide a geometric reference for azimuth angle calculation; S2.4. Perform time delay compensation on the signals of each microphone array, use the delay-and-sum beamforming algorithm to weighted sum the compensated signals, and generate a full-direction response map from 0° to 360°. The peak of the map corresponds to the potential azimuth of the sound source; S2.5. Identify the main lobe peak direction in the beamforming response map as the initial azimuth angle estimate value, and eliminate the side lobe false peaks through TDOA distance difference constraints; use the extended Kalman filter algorithm to iteratively optimize the sequential azimuth angle estimate value to suppress the estimation errors caused by environmental noise and multipath effects, and finally output the accurate azimuth angles corresponding to each microphone array.

[0010] Preferably, the microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module, and a CPU; multiple microphone arrays synchronously collect the sound source signal, and after obtaining the multi-channel audio data, send the multi-channel audio data to the cloud server in real time through the WiFi module; Introduce a portable forwarding device, which includes a camera, a Bluetooth module, a WiFi module, a 4G module, and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

[0011] Further optionally, perform step S3 to calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the y-axis. In this process: When the sound source is on one side of the positive half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90°, that is, calculate |θ – 90°|. The smaller |θ – 90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ – 90°| is, the farther the corresponding microphone array is from the sound source; When the sound source is located on the negative half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270°, that is, calculate |θ – 270°|. The smaller |θ – 270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ – 270°| is, the farther the corresponding microphone array is from the sound source.

[0012] Further optionally, the trigonometric function expressions are as follows: y / x = tan(θ 1 -180°), -y / (2L - x) = tan(θ 2 -180°), In the formula, θ 1 represents the azimuth angle of the sound source relative to the nearest microphone array, θ 2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between two adjacent microphone arrays, and x and y represent the abscissa and ordinate of the sound source in the two-dimensional microphone array coordinate system.

[0013] Preferably, at least three microphone arrays with the same orientation are deployed at equal intervals along the x-axis direction, and the spacing between adjacent microphone arrays does not exceed 3 meters.

[0014] A sound source localization method based on multi-microphone array collaborative networking according to the present invention has the beneficial effects compared with the prior art as follows: The present invention can accurately locate the position of the sound source, making up for the defect that sound sources at different positions may correspond to the same azimuth angle estimation value, thus making it impossible to effectively distinguish sound sources at different positions, and is particularly suitable for accurate sound source localization in rectangular meeting scenarios such as large conference rooms and lecture halls. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Attached Figure 1 is the flowchart of the method of the embodiment of the present invention; Attached Figure 2 is the structural connection diagram related to the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] In order to make the technical solutions, technical problems to be solved, and technical effects of the present invention clearer and more understandable, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.

[0017] Embodiment: Referring to Attached Figure 1 , this embodiment proposes a sound source localization method based on multi-microphone array collaborative networking, which includes the following steps: S1. Establish a two-dimensional coordinate system for the microphone array. Deploy multiple microphone arrays with the same orientation at equal intervals along the x-axis. The midpoint of the connection line of the multiple microphone arrays is the origin of the coordinate system.

[0018] Specifically, the positive half-axis of the x-axis of the two-dimensional coordinate system of the microphone array is 0°, the negative half-axis of the x-axis is 180°, the positive half-axis of the y-axis is 90°, and the negative half-axis of the y-axis is 270°. The azimuth angle is the angle starting from the positive half-axis of the x-axis and rotating counterclockwise to the sound source direction. The azimuth angle value is greater than or equal to 0° and less than 360°.

[0019] S2. Multiple microphone arrays synchronously collect the sound source signals to obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through their own WiFi functions.

[0020] Reference appendix Figure 2 In this step, the microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module, and a CPU. Multiple microphone arrays synchronously collect the sound source signals. After obtaining the multi-channel audio data, the multi-channel audio data is sent to the cloud server in real time through the WiFi module. A portable forwarding device is introduced. The portable forwarding device includes a camera, a Bluetooth module, a WiFi module, a 4G module, and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

[0021] The cloud server calculates the time difference of the sound source signal arriving at each microphone array through the multi-channel audio data. Combining with the two-dimensional coordinate system of the microphone array, the beamforming algorithm is used to estimate the azimuth angle of the sound source relative to each microphone array. This process specifically includes: S2.1. Synchronously calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use the adaptive filtering algorithm to remove background noise, and perform signal enhancement through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing. S2.2. Through the generalized cross-correlation-phase transformation algorithm, calculate the time difference of the sound source signal arriving at any two microphone arrays, convert the time difference into a distance difference in combination with the speed of sound, and generate multiple groups of TDOA data to characterize the relative distance difference between the sound source and each microphone array. S2.3. Based on the two-dimensional coordinate system of the microphone array, determine the positional relationship of each microphone array to provide a geometric reference for azimuth angle calculation. S2.4. Perform time delay compensation on the signals of each microphone array, and use the delay-sum beamforming algorithm to weight and sum the compensated signals to generate a 0°-360° omnidirectional response map. The peak of the map corresponds to the potential azimuth of the sound source. S2.5. Identify the main lobe peak direction in the beamforming response pattern as the initial azimuth angle estimate, and eliminate the side lobe false peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time-series azimuth angle estimate, suppress the estimation errors caused by environmental noise and multipath effects, and finally output the accurate azimuth angles corresponding to each microphone array.

[0022] S3. Calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the positive or negative half-axis of the y-axis, select the one with the smallest absolute value as the microphone array closest to the sound source, select the one with the largest absolute value as the microphone array farthest from the sound source, and calculate the distance between the two microphone arrays closest and farthest from the sound source; in this process: When the sound source is on one side of the positive half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90°, that is, calculate |θ - 90°|. The smaller |θ - 90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 90°| is, the farther the corresponding microphone array is from the sound source. When the sound source is on one side of the negative half-axis of the y-axis, calculate the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270°, that is, calculate |θ - 270°|. The smaller |θ - 270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ - 270°| is, the farther the corresponding microphone array is from the sound source.

[0023] S4. Based on the distance between the two microphone arrays closest and farthest from the sound source, the azimuth angle of the sound source relative to the closest microphone array, and the azimuth angle of the sound source relative to the farthest microphone array, calculate the coordinate values of the sound source in the two-dimensional microphone array coordinate system through trigonometric functions.

[0024] The trigonometric function expressions are as follows: y / x = tan(θ 1 - 180°), -y / (2L - x) = tan(θ 2 - 180°), where θ 1 represents the azimuth angle of the sound source relative to the closest microphone array, θ 2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between adjacent two microphone arrays, and x, y represent the abscissa and ordinate of the sound source in the two-dimensional microphone array coordinate system.

[0025] Taking "deploying five microphone arrays with the same orientation at equal intervals along the x-axis direction, and the spacing between adjacent microphone arrays is 2 meters" as an example, specifically describe the implementation process of the above steps: (1) The five microphone arrays from left to right are sequentially labeled as A, B, C, D, and E. Taking microphone array C as the origin of the coordinate system, then A(-4, 0), B(-2, 0), C(0, 0), D(2, 0), and E(4, 0).

[0026] (i) Assume that the sound source S1 is located on the positive y-axis, that is, in the first quadrant or the second quadrant of the coordinate system; the microphone arrays A, B, C, D, and E respectively collect the sound source S1, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through their own WiFi functions. The cloud server calculates the time difference between the sound source S1 reaching the microphone arrays A, B, C, D, and E through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone arrays. The beamforming algorithm is used to estimate the azimuth angle of the sound source relative to each microphone array, and assume that the azimuth angles obtained are θ A , θ B , θ C , θ D and θ E .

[0027] Further calculate the absolute value of the difference between the azimuth angle and the positive y-axis, that is, |θ – 90°|, where θ takes the values of θ A , θ B , θ C , θ D and θ E。 Assume that the value of |θ A – 90°| is the largest, and the value of |θ C – 90°| is the smallest. Then microphone array A is used as the microphone array farthest from the sound source, microphone array C is used as the microphone array closest to the sound source, and the distance between microphone array A and microphone array C is 4 meters.

[0028] Assume that θ A specifically takes the value of 80°, and θ C specifically takes the value of 45°. Then the coordinate values of the sound source S in the two-dimensional microphone array coordinate system are calculated through trigonometric functions, y / x = tan(θ C - 180°), -y / (2L - x) = tan(θ A - 180°), In the formula, θ 1 represents the azimuth angle of the sound source relative to the closest microphone array, and θ 2 represents the azimuth angle of the sound source relative to the farthest microphone array. L represents the distance between adjacent two microphone arrays, and x and y represent the abscissa and ordinate of the sound source S1 in the two-dimensional microphone array coordinate system. The specific calculated coordinates are (0.86, 4.86).

[0029] (ii) Assume that the sound source S2 is located on the negative y-axis, i.e., in the third or fourth quadrant of the coordinate system; the microphone arrays A, B, C, D, and E respectively collect the sound source S2, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through their own WiFi functions. The cloud server calculates the time difference between the sound source S2 reaching the microphone arrays A, B, C, D, and E through the multi-channel audio data, and combines the two-dimensional coordinate system of the microphone arrays. The beamforming algorithm is used to estimate the azimuth angle of the sound source relative to each microphone array. Assume that the obtained azimuth angles are θ A , θ B , θ C , θ D and θ E .

[0030] Further calculate the absolute value of the difference between the azimuth angle and the positive y-axis, i.e., |θ – 270°|, where θ takes the values of θ A , θ B , θ C , θ D and θ E。 Assume that the value of |θ A – 90°| is the largest, and the value of |θ C – 270°| is the smallest. Then the microphone array E is used as the microphone array farthest from the sound source, the microphone array C is used as the microphone array closest to the sound source, and the distance between the microphone array E and the microphone array C is 4 meters.

[0031] Assume that θ E specifically takes the value of 225°, and θ C specifically takes the value of 260°. Then, through trigonometric functions, the coordinate values of the sound source S2 in the two-dimensional microphone array coordinate system are calculated. y / x = tan(θ C - 180°), -y / (2L - x) = tan(θ E - 180°), where θ 1 represents the azimuth angle of the sound source relative to the closest microphone array, θ 2 represents the azimuth angle of the sound source relative to the farthest microphone array, L represents the spacing between two adjacent microphone arrays, and x and y represent the abscissa and ordinate of the sound source S2 in the two-dimensional microphone array coordinate system. The specific calculated coordinates are (-0.86, -4.86).

[0032] In summary, by using a sound source localization method based on multi-microphone array collaborative networking according to the present invention, the coordinate value of the sound source in the two-dimensional microphone array coordinate system can be calculated by establishing a two-dimensional coordinate system of the microphone array, calculating the azimuth angles between the sound source and different microphone arrays, calculating the absolute values of the differences between the azimuth angles and the nearest and farthest microphone arrays. This process can improve the sound source localization accuracy and make up for the defect that sound sources at different positions may correspond to the same azimuth angle estimation value, thus making it impossible to effectively distinguish sound sources at different positions.

[0033] The above applications have elaborated in detail the principle and implementation manner of the present invention with specific examples. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art in this technical field without departing from the principle of the present invention shall fall within the patent protection scope of the present invention.

Claims

1. A sound source localization method based on multi-microphone array collaborative networking, characterized in that: The steps include: S1. Establish a two-dimensional coordinate system for the microphone array, and deploy multiple microphone arrays with the same orientation at equal distances along the x-axis, with the midpoint of the line connecting the multiple microphone arrays as the origin of the coordinate system; S2, multiple microphone arrays synchronously collect sound source signals to obtain multi-channel audio data, calculate the time difference between the sound source signal reaching each microphone array through the multi-channel audio data, and estimate the azimuth of the sound source relative to each microphone array in combination with the two-dimensional coordinate system of the microphone array; S3, calculating the absolute value of the difference between the azimuth of the sound source relative to each microphone array and the positive semi-axis or negative semi-axis of the y-axis, selecting the microphone array with the smallest absolute value as the microphone array closest to the sound source, selecting the microphone array with the largest absolute value as the microphone array farthest from the sound source, and calculating the distance between the two microphone arrays closest to and farthest from the sound source; S4. Based on the distance between the two microphone arrays closest and farthest from the sound source, the azimuth of the sound source relative to the nearest microphone array, and the azimuth of the sound source relative to the farthest microphone array, the coordinate value of the sound source in the two-dimensional microphone array coordinate system is obtained by trigonometric calculation.

2. A method for sound source localization based on multi-microphone array collaborative networking according to claim 1, characterized in that: The positive x-axis of the microphone array two-dimensional coordinate system is 0°, the negative x-axis is 180°, the positive y-axis is 90°, and the negative y-axis is 270°. The azimuth angle is the angle starting from the positive x-axis and rotated counterclockwise to the direction of the sound source. The azimuth angle is greater than or equal to 0° and less than 360°.

3. The method for sound source localization based on multi-microphone array collaborative networking according to claim 2, characterized in that: Executing step S2, after the multiple microphone arrays acquire the multi-channel audio data, the multiple microphone arrays send the multi-channel audio data to the cloud server in real time through their own WiFi function; The cloud server calculates the time difference between the sound source signal reaching each microphone array through multi-channel audio data, and uses a beamforming algorithm to estimate the azimuth of the sound source relative to each microphone array in combination with the two-dimensional coordinate system of the microphone array.

4. The method for sound source localization based on multi-microphone array collaborative networking according to claim 3, characterized in that: The cloud server calculates the time difference between the sound source signal and each microphone array through multi-channel audio data, and uses the beamforming algorithm to estimate the azimuth of the sound source relative to each microphone array in combination with the two-dimensional coordinate system of the microphone array. This process specifically includes: S2.

1. Synchronize and calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use an adaptive filtering algorithm to remove background noise, and enhance the signal through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing; S2.2, calculate the time difference between the sound source signal reaching any two microphone arrays through the generalized cross-correlation-phase transformation algorithm, convert the time difference into a distance difference based on the speed of sound, and generate multiple sets of TDOA data to characterize the relative distance difference from the sound source to each microphone array; S2.3, based on the microphone array two-dimensional coordinate system, determine the positional relationship of each microphone array to provide a geometric reference for azimuth calculation; S2.4, delay compensation is performed on the signal of each microphone array, and the compensated signal is weighted and summed using a delay-sum beamforming algorithm to generate a 0°-360° omnidirectional response spectrum, where the spectrum peak corresponds to the potential direction of the sound source; S2.

5. Identify the main lobe peak direction in the beamforming response spectrum as the initial azimuth angle estimate, and eliminate the sidelobe pseudo peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time series azimuth angle estimate to suppress the estimation error caused by environmental noise and multipath effects, and finally output the precise azimuth angle corresponding to each microphone array.

5. The method for sound source localization based on multi-microphone array collaborative networking according to claim 3, characterized in that: The microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module and a CPU; the multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and send the multi-channel audio data to the cloud server in real time through the WiFi module; A portable forwarding device is introduced, which includes a camera, a Bluetooth module, a WiFi module, a 4G module and a CPU. The portable forwarding device pre-configures the WiFi module of the microphone array through the Bluetooth module.

6. The method for sound source localization based on multi-microphone array collaborative networking according to claim 3, characterized in that: Step S3 is executed to calculate the absolute value of the difference between the azimuth angle of the sound source relative to each microphone array and the y-axis. In this process: When the sound source is located on the positive half axis of the y-axis, the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 90° is calculated, that is, |θ–90°| is calculated. The smaller |θ–90°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ–90°| is, the farther the corresponding microphone array is from the sound source. When the sound source is located on the negative half-axis of the y-axis, the absolute value of the difference between the azimuth angle θ of the sound source relative to each microphone array and 270° is calculated, that is, |θ–270°| is calculated. The smaller |θ–270°| is, the closer the corresponding microphone array is to the sound source, and the larger |θ–270°| is, the farther the corresponding microphone array is from the sound source.

7. The method for sound source localization based on multi-microphone array collaborative networking according to claim 6, characterized in that: The trigonometric function expressions are as follows: y / x=tan(θ1-180°), -y / (2L-x)=tan(θ2-180°), Where θ1 represents the azimuth of the sound source relative to the nearest microphone array, θ2 represents the azimuth of the sound source relative to the farthest microphone array, L represents the distance between two adjacent microphone arrays, and x and y represent the horizontal and vertical coordinates of the sound source in the two-dimensional microphone array coordinate system.

8. The method for sound source localization based on multi-microphone array collaborative networking according to claim 7, characterized in that: At least three microphone arrays with the same orientation are deployed equidistantly along the x-axis direction, and the distance between adjacent microphone arrays does not exceed 3 meters.

Citation Information

Patent Citations

  • Method for positioning sound source by using microphone array

    CN102033223A

  • SRP-PHAT multi-source spatial positioning method

    CN104142492A

  • Sound source three-dimensional positioning method and device

    CN116299182A

  • Sound positioning system, sound pickup method, sound positioning method and equipment

    CN117202014A

  • Hybrid deployment method for heterogeneous multi-core chip application program

    CN119883650A