Sound source localization optimization method based on distributed multi-microphone array

Through infrared pulse triggering and cloud computing, the synchronization and coordination mechanism of microphone arrays is optimized, and the synchronization and coordination mechanism of traditional microphone arrays is solved, and high-precision sound source positioning is achieved.

CN120428167APending Publication Date: 2025-08-05SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510538478.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Traditional distributed microphone arrays have shortcomings in synchronization accuracy and coordinated positioning accuracy, especially in complex acoustic environments, and are affected by reverb and have outliers.

Method used

By establishing a two-dimensional coordinate system, using infrared pulses to trigger a synchronous microphone array, combining cloud servers to calculate the time difference and azimuth angle, eliminating the multipath effect, iteratively compute the sound source coordinates, and optimizing the synchronization and coordination mechanism.

Benefits of technology

It improves the accuracy and robustness of sound source positioning, is suitable for high-precision positioning in complex acoustic environments, and reduces time synchronization errors and multipath interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428167A_ABST
    Figure CN120428167A_ABST
Patent Text Reader

Abstract

The invention discloses a sound source localization optimization method based on a distributed multi-microphone array, which relates to the technical field of sound source localization, and comprises the following steps: establishing a two-dimensional coordinate system, and deploying a plurality of microphone arrays with the same orientation in the coordinate system in a linear or rectangular manner; the portable forwarding device triggers and starts the arrays through infrared pulses, the multiple arrays synchronously collect sound source signals, acquire multichannel audio data and upload the multichannel audio data to the cloud server, and the cloud server calculates the time difference of the sound source signals reaching the arrays and estimates the azimuth angle of a sound source relative to each array. Calculating a reference point according to the azimuth angle rays of the array without the multipath effect, and eliminating false azimuth angle rays of the array generated by the multipath effect; the reference point is calculated again, then iterative calculation is carried out, the azimuth angle rays which deviate from the new reference point to the maximum degree are removed one by one till two azimuth angle rays are left, the intersection point of the two azimuth angle rays is calculated, and then the sound source coordinate value is obtained. According to the method, time difference calculation errors and multipath interference can be reduced, and the sound source positioning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound source localization, and in particular to a sound source localization optimization method based on a distributed multi-microphone array. Background Art

[0002] Traditional distributed microphone arrays achieve synchronization through GPS hardware clocks or internet time servers, but this suffers from poor indoor signal quality and large synchronization errors. This low synchronization accuracy prevents effective coordination between microphone arrays, hindering positioning accuracy. Due to reverberation, some microphone arrays may generate abnormal azimuth estimates at certain times, significantly reducing positioning accuracy. Summary of the Invention

[0003] In response to the needs and shortcomings of current technological development, the present invention provides a sound source localization optimization method based on a distributed multi-microphone array. By optimizing the synchronization mechanism and coordination mechanism between microphone arrays, the positioning accuracy is improved. It is particularly suitable for accurate sound source localization in rectangular venues such as large conference rooms and lecture halls.

[0004] The present invention provides a sound source localization optimization method based on a distributed multi-microphone array, which solves the above technical problems using the following technical solutions:

[0005] A sound source localization optimization method based on a distributed multi-microphone array comprises the following steps:

[0006] S1. Establish a two-dimensional coordinate system for the microphone array. Deploy multiple microphone arrays equidistantly along the x-axis with the same orientation, and use the midpoint of the line connecting the arrays as the origin of the coordinate system. Alternatively, deploy multiple microphone arrays in a rectangular structure, with the microphones on the same side of the rectangle facing the same direction, and the geometric center of the rectangle coinciding with the origin of the coordinate system.

[0007] S2. The portable forwarding device triggers the microphone array through infrared pulses. Multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and upload it to the cloud server;

[0008] S3. The cloud server calculates the time difference between the sound source signal and each microphone array using the multi-channel audio data, and estimates the azimuth angle of the sound source relative to each microphone array based on the two-dimensional coordinate system of the microphone array.

[0009] S4. Calculate the reference point ρ based on the azimuth ray of the microphone array without multipath effect. * , according to the reference point ρ * Eliminate false azimuth rays of the microphone array caused by multipath effects;

[0010] S5. Recalculate the reference point ρ based on the azimuth rays corresponding to the multiple microphone arrays. final , then iteratively calculate and eliminate the reference point ρ one by one final The azimuth ray with the largest deviation is continued until only two azimuth rays remain. The intersection of the two azimuth rays is calculated to obtain the coordinate value of the sound source in the two-dimensional microphone array coordinate system.

[0011] Optionally, the positive x-axis of the two-dimensional coordinate system of the microphone array involved is 0°, the negative x-axis is 180°, the positive y-axis is 90°, and the negative y-axis is 270°. The azimuth angle is the angle starting from the positive x-axis and rotated counterclockwise to the direction of the sound source. The azimuth angle is greater than or equal to 0° and less than 360°.

[0012] Further optionally, the microphone array involved includes a multi-channel microphone, a Bluetooth module, a WiFi module, a CPU, an infrared receiving module and a power management unit;

[0013] The portable forwarding devices involved include cameras, Bluetooth modules, WiFi modules, 4G modules, CPUs, infrared transmitter modules and synchronous clock units;

[0014] The following operations are performed between the portable forwarding device and the multiple microphone arrays:

[0015] S2.1. The portable forwarding device establishes a temporary connection with the microphone array via the Bluetooth module. The CPU of the portable forwarding device then sends the WiFi network parameters to the CPU of the microphone array via the Bluetooth module. The CPU of the microphone array writes the parameters into the WiFi module register. After the configuration is complete, the Bluetooth connection is disconnected and the RF function of the Bluetooth module is disabled.

[0016] S2.2. The portable forwarding device connects the infrared transmitter module to the infrared receiver modules of all microphone arrays one by one through an optical fiber splitter. Use equal-length optical fibers or delay lines for calibration to ensure consistent optical signal transmission delay.

[0017] S2.3. The CPU of the portable forwarding device generates a trigger instruction with a global timestamp, and sends a modulated infrared pulse through the infrared transmitter module. The infrared pulse reaches the infrared receiving modules of all microphone arrays synchronously through the optical fiber splitter. The infrared receiving modules parse the pulse code and send an interrupt signal to the CPU after confirming a valid trigger.

[0018] S2.4. Wake up the CPU of the microphone array, start the microphone array through the power management unit, start synchronously collecting audio data at the set sampling rate, and record the local trigger time;

[0019] S2.5. The CPU of the microphone array starts the radio frequency function of the WiFi module and establishes a TCP / IP connection with the cloud server. The microphone array uploads the collected audio data to the server through the WiFi module.

[0020] Optionally, step S3 is performed, where the cloud server calculates the time difference between the sound source signal and each microphone array using the multi-channel audio data, and uses a beamforming algorithm to estimate the azimuth of the sound source relative to each microphone array in combination with the two-dimensional coordinate system of the microphone array. This process specifically includes:

[0021] S3.1. Synchronize and calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use an adaptive filtering algorithm to remove background noise, and enhance the signal through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing.

[0022] S3.2. Calculate the time difference between the sound source signal reaching any two microphone arrays using a generalized cross-correlation-phase transformation algorithm. Combined with the speed of sound, this time difference is converted into a distance difference, generating multiple sets of TDOA data to represent the relative distance differences between the sound source and each microphone array.

[0023] S3.3. Determine the positional relationship of each microphone array based on the microphone array two-dimensional coordinate system to provide a geometric reference for azimuth angle calculation;

[0024] S3.4. Delay compensation is performed on the signal from each microphone array. The compensated signals are weighted and summed using a delay-sum beamforming algorithm to generate a 0°-360° omnidirectional response spectrum, where the peak of the spectrum corresponds to the potential direction of the sound source.

[0025] S3.5. Identify the mainlobe peak direction in the beamforming response spectrum as the initial azimuth angle estimate, and eliminate the sidelobe pseudo-peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time-series azimuth angle estimate to suppress the estimation error caused by environmental noise and multipath effects, and finally output the precise azimuth angle corresponding to each microphone array.

[0026] Further optionally, step S4 specifically includes:

[0027] S4.1. Calculate the time difference between the arrival of the sound source signal at each microphone array using multi-channel audio data to select microphone arrays that may receive direct sound. The microphone arrays that are excluded from the selection have at least two azimuth rays.

[0028] S4.2. Based on the selected microphone array, calculate the intersection point set ρ generated by the intersection of its azimuth rays core , for ρ coreApply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point ρ * , and determine the reference point ρ when the number of points and the standard deviation of the coordinates in the main cluster meet the preset conditions * efficient;

[0029] S4.3. Using a valid reference point ρ * , at least two azimuth rays of the microphone array screened out in step S4.1 are screened in sequence, and false azimuth rays introduced by the multipath effect are eliminated, so that only one azimuth ray is retained by the microphone array screened out in step S4.1, that is, azimuth rays corresponding to multiple microphone arrays are obtained.

[0030] Further optionally, step S5 specifically includes:

[0031] S5.1. Calculate the intersection point set ρ generated by the intersection of the azimuth rays obtained in step S4.3. final , for ρ final Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point Elimination and The azimuth ray with the largest deviation distance;

[0032] S5.2, then iterate and eliminate the The ray with the largest deviation distance is deviated until only two azimuth rays remain. The intersection of these two azimuth rays is the coordinate value of the sound source in the two-dimensional microphone array coordinate system.

[0033] Preferably, the following formula is used to calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system:

[0034]

[0035] y = tanθ a *(xx a )+y a ,

[0036] Where a and b are the microphone arrays corresponding to the remaining two azimuth rays, θ a represents the azimuth angle of microphone array a, θ b represents the azimuth angle of microphone array b, x a 、y a represents the coordinate value of microphone array a, x b 、y b Represents the coordinate value of microphone array b.

[0037] Optionally, at least five microphone arrays with the same orientation are deployed equidistantly along the x-axis, and the distance between adjacent microphone arrays does not exceed 3 meters.

[0038] Optionally, at least eight microphone arrays are deployed in a rectangular structure, wherein four microphone arrays are located at four end points of the rectangle, and the remaining four microphone arrays are located at midpoints of four sides of the rectangle.

[0039] The present invention provides a sound source localization optimization method based on a distributed multi-microphone array, which has the following beneficial effects compared with the prior art:

[0040] The present invention optimizes the synchronization mechanism of the microphone array through infrared synchronous triggering and regularized layout, ensuring the accuracy of the time and space references; improves the robustness of the collaborative mechanism through multi-array data collaborative filtering and iterative optimization, effectively suppressing multipath effects and outliers; the two work together to improve the sound source localization accuracy from three dimensions: time synchronization, spatial calibration, and data filtering, and is particularly suitable for high-precision positioning scenarios in complex acoustic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Attachment Figure 1 is a flow chart of a method according to embodiment 1 of the present invention;

[0042] Attachment Figure 2 This is a structural diagram of the connection between the microphone array, portable forwarding device and cloud server involved in the first embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the technical solution, the technical problems solved and the technical effects of the present invention more clear, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.

[0044] Example 1:

[0045] Reference Attachment Figure 1 This embodiment proposes a sound source localization optimization method based on a distributed multi-microphone array, which includes the following steps:

[0046] S1. Establish a two-dimensional coordinate system for the microphone array. Deploy multiple microphone arrays with the same orientation at equal distances along the x-axis, and use the midpoint of the line connecting the arrays as the origin of the coordinate system. Alternatively, deploy multiple microphone arrays in a rectangular structure, with the microphones on the same side of the rectangle facing the same direction, and the geometric center point of the rectangle coinciding with the origin of the coordinate system.

[0047] The positive x-axis of the established microphone array two-dimensional coordinate system is 0°, the negative x-axis is 180°, the positive y-axis is 90°, and the negative y-axis is 270°. The azimuth angle is the angle starting from the positive x-axis and rotated counterclockwise to the direction of the sound source. The azimuth angle value is greater than or equal to 0° and less than 360°.

[0048] S2. The portable forwarding device triggers the microphone array through infrared pulses. Multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and upload it to the cloud server.

[0049] Reference Attachment Figure 2 The microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module, a CPU, an infrared receiver module, and a power management unit. The portable forwarding device includes a camera, a Bluetooth module, a WiFi module, a 4G module, a CPU, an infrared transmitter module, and a synchronous clock unit.

[0050] The following operations are performed between the portable forwarding device and the multiple microphone arrays:

[0051] S2.1. The portable forwarding device establishes a temporary connection with the microphone array via the Bluetooth module. The CPU of the portable forwarding device then sends the WiFi network parameters to the CPU of the microphone array via the Bluetooth module. The CPU of the microphone array writes the parameters into the WiFi module register. After the configuration is complete, the Bluetooth connection is disconnected and the RF function of the Bluetooth module is disabled.

[0052] S2.2. The portable forwarding device connects the infrared transmitter module to the infrared receiver modules of all microphone arrays one by one through an optical fiber splitter. Use equal-length optical fibers or delay lines for calibration to ensure consistent optical signal transmission delay.

[0053] S2.3. The CPU of the portable forwarding device generates a trigger instruction with a global timestamp, and sends a modulated infrared pulse through the infrared transmitter module. The infrared pulse reaches the infrared receiving modules of all microphone arrays synchronously through the optical fiber splitter. The infrared receiving modules parse the pulse code and send an interrupt signal to the CPU after confirming a valid trigger.

[0054] S2.4. Wake up the CPU of the microphone array, start the microphone array through the power management unit, start synchronously collecting audio data at the set sampling rate, and record the local trigger time;

[0055] S2.5. The CPU of the microphone array starts the radio frequency function of the WiFi module and establishes a TCP / IP connection with the cloud server. The microphone array uploads the collected audio data to the server through the WiFi module.

[0056] S3. The cloud server calculates the time difference between the sound source signal reaching each microphone array using multi-channel audio data. Combining the two-dimensional coordinate system of the microphone array, it uses a beamforming algorithm to estimate the azimuth of the sound source relative to each microphone array. This process specifically includes:

[0057] S3.1. Synchronize and calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use an adaptive filtering algorithm to remove background noise, and enhance the signal through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing.

[0058] S3.2. Calculate the time difference between the sound source signal reaching any two microphone arrays using a generalized cross-correlation-phase transformation algorithm. Combined with the speed of sound, this time difference is converted into a distance difference, generating multiple sets of TDOA data to represent the relative distance differences between the sound source and each microphone array.

[0059] S3.3. Determine the positional relationship of each microphone array based on the microphone array two-dimensional coordinate system to provide a geometric reference for azimuth angle calculation;

[0060] S3.4. Delay compensation is performed on the signal from each microphone array. The compensated signals are weighted and summed using a delay-sum beamforming algorithm to generate a 0°-360° omnidirectional response spectrum, where the peak of the spectrum corresponds to the potential direction of the sound source.

[0061] S3.5. Identify the mainlobe peak direction in the beamforming response spectrum as the initial azimuth angle estimate, and eliminate the sidelobe pseudo-peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time-series azimuth angle estimate to suppress the estimation error caused by environmental noise and multipath effects, and finally output the precise azimuth angle corresponding to each microphone array.

[0062] S4. Calculate the reference point ρ based on the azimuth ray of the microphone array without multipath effect. * , according to the reference point ρ * Eliminate the false azimuth rays of the microphone array caused by multipath effects, including:

[0063] S4.1. Calculate the time difference between the arrival of the sound source signal at each microphone array using multi-channel audio data to select microphone arrays that may receive direct sound. The microphone arrays that are excluded from the selection have at least two azimuth rays.

[0064] S4.2. Based on the selected microphone array, calculate the intersection point set ρ generated by the intersection of its azimuth rays core , for ρ core Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point ρ * , and determine the reference point ρ when the number of points and the standard deviation of the coordinates in the main cluster meet the preset conditions * efficient;

[0065] S4.3. Using a valid reference point ρ *, at least two azimuth rays of the microphone array screened out in step S4.1 are screened in sequence, and false azimuth rays introduced by the multipath effect are eliminated, so that only one azimuth ray is retained by the microphone array screened out in step S4.1, that is, azimuth rays corresponding to multiple microphone arrays are obtained.

[0066] S5. Recalculate the reference point ρ based on the azimuth rays corresponding to the multiple microphone arrays. final , then iteratively calculate and eliminate the reference point ρ one by one final Deviate from the largest azimuth ray until only two azimuth rays remain. Calculate the intersection of these two azimuth rays to obtain the coordinates of the sound source in the two-dimensional microphone array coordinate system. This includes:

[0067] S5.1. Calculate the intersection point set ρ generated by the intersection of the azimuth rays obtained in step S4.3. final , for ρ final Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point Elimination and The azimuth ray with the largest deviation distance;

[0068] S5.2, then iterate and eliminate the The ray with the largest deviation distance is deviated until only two azimuth rays remain. The intersection of these two azimuth rays is the coordinate value of the sound source in the two-dimensional microphone array coordinate system.

[0069] Specifically, the following formula is used to calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system:

[0070]

[0071] y = tanθ a *(xx a )+y a ,

[0072] Where a and b are the microphone arrays corresponding to the remaining two azimuth rays, θ a represents the azimuth angle of microphone array a, θ b represents the azimuth angle of microphone array b, x a 、y a represents the coordinate value of microphone array a, x b 、y b Represents the coordinate value of microphone array b.

[0073] Example 2:

[0074] Based on the sound source localization optimization method of Example 1, this example takes "8 microphone arrays are deployed equidistantly along the x-axis with the midpoint of the array connection line as the origin of the coordinate system" as an example to describe the specific process of sound source localization optimization:

[0075] (1) Establish a two-dimensional coordinate system for the microphone array. Deploy eight microphone arrays with the same orientation at equal distances along the x-axis, and use the midpoint of the line connecting the arrays as the origin of the coordinate system.

[0076] Among them, the positive x-axis of the established microphone array two-dimensional coordinate system is 0°, the negative x-axis is 180°, the positive y-axis is 90°, and the negative y-axis is 270°. The azimuth angle is the angle starting from the positive x-axis and rotated counterclockwise to the direction of the sound source. The azimuth angle value is greater than or equal to 0° and less than 360°.

[0077] (2) The portable forwarding device triggers the microphone array through infrared pulses. Multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and upload it to the cloud server.

[0078] (3) The cloud server calculates the time difference between the sound source signal and the arrival of each microphone array through multi-channel audio data, and uses a beamforming algorithm to estimate the azimuth angle of the sound source relative to each microphone array in combination with the two-dimensional coordinate system of the microphone array.

[0079] (4) Calculate the reference point ρ based on the azimuth ray of the microphone array without multipath effect * , according to the reference point ρ * Eliminate the false azimuth rays of the microphone array caused by multipath effects, including:

[0080] (4.1) The time difference between the sound source signal reaching each microphone array (specifically referred to as microphone array 1, microphone array 2, microphone array 3, microphone array 4, microphone array 5, microphone array 6, microphone array 7, and microphone array 8) is calculated through multi-channel audio data to select the microphone array that may receive the direct sound. Assume that only one azimuth ray is obtained from microphone array 1 to microphone array 6, which is recorded as {v1, v2, v3, v4, v5, v6}, and microphone arrays 7 and 8 obtain two or more azimuth rays, which are recorded as {v 7a ,v 7b ,v 8a ,v 8b}.

[0081] (4.2) Based on the screened microphone arrays 1 to 6, calculate the set of 15 intersection points ρ generated by the intersection of their azimuth rays {v1, v2, v3, v4, v5, v6} core , denoted as ρ core ={ρij |ρ ij =v i ∩v j ; i, j = 1,…, 6; i ≠ j};

[0082] ρ core Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point ρ * (x ρ ,y ρ ), and the reference point ρ is determined when the number of points in the main cluster and the standard deviation of the coordinates meet the preset conditions (assuming that the preset conditions are that the number of points in the main cluster is ≥8 and the standard deviation of the coordinates is ≤0.05m). * efficient.

[0083] (4.3) Using the effective reference point ρ * , the azimuth rays {v 7a ,v 7b ,v 8a ,v 8b}Screen in sequence to eliminate the false azimuth rays introduced by the multipath effect, so that only one azimuth ray is retained by the microphone array screened out in step S4.1, that is, azimuth rays corresponding to multiple microphone arrays are obtained.

[0084] The specific screening process is as follows:

[0085] Starting from the microphone array 7, move towards the reference point ρ * Draw reference ray v7 for microphone array 7 * , Reference Ray v7 * The azimuth angle relative to the microphone array 7 is

[0086] Starting from the microphone array 8, * Draw reference ray v8 for microphone array 8 * , the azimuth angle of the reference ray relative to the microphone array 8 is

[0087] Where x7 and y7 represent the coordinates of microphone array 7, x ρ 、y ρ represents the reference point ρ * The coordinate value of .

[0088] Calculate {v 7a ,v 7b} corresponding azimuth angle θ 7a ,θ 7b and the calculated azimuth The absolute difference of:

[0089]

[0090] The rays with larger Δθ are the false azimuth rays produced by the multipath effect. These azimuth rays are discarded and the rays with smaller Δθ are retained.

[0091] Similarly, calculate {v 8a ,v 8b}The corresponding azimuth and the calculated azimuth and eliminates the false azimuth rays generated by the microphone array 8 due to the multipath effect.

[0092] If v 7b and v 8b It is a false azimuth ray generated by the multipath effect. After it is removed, only one azimuth ray is left for microphone arrays 7 and 8, which together with the azimuth rays from microphone array 1 to microphone array 6 form a new set {v1,v2,v3,v4,v5,v6,v 7a ,v 8a}.

[0093] (5) The set of azimuth rays from microphone array 1 to microphone array 8 is {v1,v2,v3,v4,v5,v6,v 7a ,v 8a}, generally speaking, these 8 azimuth rays will produce 28 intersection points ρ final ,

[0094] ρ final Apply DBSCAN clustering to obtain the main cluster C final , take the main cluster C final The geometric median of Elimination and The azimuth ray with the largest deviation is selected, and then the azimuth rays with the largest deviation are removed one by one through iterative calculation until only two azimuth rays remain. The intersection of the remaining two azimuth rays is the precise location of the sound source.

[0095] Specifically, the following formula is used to calculate the coordinate value of the sound source in the two-dimensional microphone array coordinate system:

[0096]

[0097] y = tanθ a *(xx a )+y a ,

[0098] Where a and b are the microphone arrays corresponding to the remaining two azimuth rays. a and b are taken from microphone arrays 1-8 and are different. a represents the azimuth angle of microphone array a, θb represents the azimuth angle of microphone array b, x a 、y a represents the coordinate value of microphone array a, x b 、y b Represents the coordinate value of microphone array b.

[0099] It should be added that, taking "8 microphone arrays are deployed in a rectangular structure, where four microphone arrays are located at the four endpoints of the rectangle, and the remaining four microphone arrays are located at the midpoints of the four sides of the rectangle, the microphones on the same side of the rectangle are oriented in the same direction, and the geometric center point of the rectangle coincides with the origin of the coordinate system" as an example, the implementation process is the same as the steps described in Example 2, and will not be repeated here.

[0100] In summary, the sound source localization optimization method based on a distributed multi-microphone array of the present invention ensures that each microphone array starts collecting data at the same time through infrared synchronous triggering, thereby reducing time synchronization errors, which belongs to the optimization of the synchronization mechanism; at the same time, through the collaborative processing of multi-array azimuth data, false rays caused by the multipath effect are eliminated, and the reference point is optimized through iterative calculation, which belongs to the optimization of the collaborative mechanism; the combined effect of the above two reduces the time difference calculation error and multipath interference, thereby improving the accuracy of sound source localization.

[0101] The above specific examples are used to illustrate the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by those skilled in the art without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.

Claims

1. A sound source localization optimization method based on a distributed multi-microphone array, characterized in that: The steps include: S1. Establish a two-dimensional coordinate system for the microphone array. Deploy multiple microphone arrays equidistantly along the x-axis with the same orientation, and use the midpoint of the line connecting the arrays as the origin of the coordinate system. Alternatively, deploy multiple microphone arrays in a rectangular structure, with the microphones on the same side of the rectangle facing the same direction, and the geometric center of the rectangle coinciding with the origin of the coordinate system. S2. The portable forwarding device triggers the microphone array through infrared pulses. Multiple microphone arrays synchronously collect sound source signals, obtain multi-channel audio data, and upload it to the cloud server; S3. The cloud server calculates the time difference between the sound source signal and each microphone array using the multi-channel audio data, and estimates the azimuth angle of the sound source relative to each microphone array based on the two-dimensional coordinate system of the microphone array. S4. Calculate the reference point ρ based on the azimuth ray of the microphone array without multipath effect. * , according to the reference point ρ * Eliminate false azimuth rays of the microphone array caused by multipath effects; S5. Recalculate the reference point ρ based on the azimuth rays corresponding to the multiple microphone arrays. final , then iteratively calculate and eliminate the reference point ρ one by one final The azimuth ray with the largest deviation is continued until only two azimuth rays remain. The intersection of the two azimuth rays is calculated to obtain the coordinate value of the sound source in the two-dimensional microphone array coordinate system.

2. The sound source localization optimization method based on a distributed multi-microphone array according to claim 1, characterized in that: The positive x-axis of the microphone array's two-dimensional coordinate system is 0°, the negative x-axis is 180°, the positive y-axis is 90°, and the negative y-axis is 270°. The azimuth angle is the angle measured counterclockwise from the positive x-axis to the direction of the sound source. The azimuth angle is greater than or equal to 0° and less than 360°.

3. The sound source localization optimization method based on a distributed multi-microphone array according to claim 2, characterized in that: The microphone array includes a multi-channel microphone, a Bluetooth module, a WiFi module, a CPU, an infrared receiving module and a power management unit; The portable forwarding device includes a camera, a Bluetooth module, a WiFi module, a 4G module, a CPU, an infrared transmission module and a synchronous clock unit; The portable forwarding device and the plurality of microphone arrays specifically perform the following operations: S2.

1. The portable forwarding device establishes a temporary connection with the microphone array via the Bluetooth module. The CPU of the portable forwarding device then sends the WiFi network parameters to the CPU of the microphone array via the Bluetooth module. The CPU of the microphone array writes the parameters into the WiFi module register. After the configuration is complete, the Bluetooth connection is disconnected and the RF function of the Bluetooth module is disabled. S2.

2. The portable forwarding device connects the infrared transmitter module to the infrared receiver modules of all microphone arrays one by one through an optical fiber splitter. Use equal-length optical fibers or delay lines for calibration to ensure consistent optical signal transmission delay. S2.

3. The CPU of the portable forwarding device generates a trigger instruction with a global timestamp, and sends a modulated infrared pulse through the infrared transmitter module. The infrared pulse reaches the infrared receiving modules of all microphone arrays synchronously through the optical fiber splitter. The infrared receiving modules parse the pulse code and send an interrupt signal to the CPU after confirming a valid trigger. S2.

4. Wake up the CPU of the microphone array, start the microphone array through the power management unit, start synchronously collecting audio data at the set sampling rate, and record the local trigger time; S2.

5. The CPU of the microphone array starts the radio frequency function of the WiFi module and establishes a TCP / IP connection with the cloud server. The microphone array uploads the collected audio data to the server through the WiFi module.

4. The sound source localization optimization method based on a distributed multi-microphone array according to claim 3, characterized in that: In step S3, the cloud server calculates the time difference between the sound source signal and each microphone array using the multi-channel audio data. Combined with the two-dimensional coordinate system of the microphone array, a beamforming algorithm is used to estimate the azimuth of the sound source relative to each microphone array. This process specifically includes: S3.

1. Synchronize and calibrate the audio signals transmitted by each microphone array based on timestamp alignment, use an adaptive filtering algorithm to remove background noise, and enhance the signal through signal normalization and dynamic range compression to ensure the time accuracy and signal-to-noise ratio of subsequent processing. S3.

2. Calculate the time difference between the sound source signal reaching any two microphone arrays using a generalized cross-correlation-phase transformation algorithm. Combined with the speed of sound, this time difference is converted into a distance difference, generating multiple sets of TDOA data to represent the relative distance differences between the sound source and each microphone array. S3.

3. Determine the positional relationship of each microphone array based on the microphone array two-dimensional coordinate system to provide a geometric reference for azimuth angle calculation; S3.

4. Delay compensation is performed on the signal from each microphone array. The compensated signals are weighted and summed using a delay-sum beamforming algorithm to generate a 0°-360° omnidirectional response spectrum, where the peak of the spectrum corresponds to the potential direction of the sound source. S3.

5. Identify the mainlobe peak direction in the beamforming response spectrum as the initial azimuth angle estimate, and eliminate the sidelobe pseudo-peaks through the TDOA distance difference constraint; use the extended Kalman filter algorithm to iteratively optimize the time-series azimuth angle estimate to suppress the estimation error caused by environmental noise and multipath effects, and finally output the precise azimuth angle corresponding to each microphone array.

5. The sound source localization optimization method based on a distributed multi-microphone array according to claim 3, characterized in that: The step S4 specifically includes: S4.

1. Calculate the time difference between the arrival of the sound source signal at each microphone array using multi-channel audio data to select microphone arrays that may receive direct sound. The microphone arrays that are excluded from the selection have at least two azimuth rays. S4.

2. Based on the selected microphone array, calculate the intersection point set ρ generated by the intersection of its azimuth rays core , for ρ core Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point ρ * , and determine the reference point ρ when the number of points and the standard deviation of the coordinates in the main cluster meet the preset conditions * efficient; S4.

3. Using a valid reference point ρ * , at least two azimuth rays of the microphone array screened out in step S4.1 are screened in sequence, and false azimuth rays introduced by the multipath effect are eliminated, so that only one azimuth ray is retained by the microphone array screened out in step S4.1, that is, azimuth rays corresponding to multiple microphone arrays are obtained.

6. The sound source localization optimization method based on a distributed multi-microphone array according to claim 5, characterized in that: The step S5 specifically includes: S5.

1. Calculate the intersection point set ρ generated by the intersection of the azimuth rays obtained in step S4.

3. final , for ρ final Apply DBSCAN clustering to identify the main cluster and take the geometric median of the main cluster as the reference point Elimination and The azimuth ray with the largest deviation distance; S5.2, then iterate and eliminate the The ray with the largest deviation distance is deviated until only two azimuth rays are left. The intersection of these two azimuth rays is the coordinate value of the sound source in the two-dimensional microphone array coordinate system.

7. The sound source localization optimization method based on a distributed multi-microphone array according to claim 6, characterized in that: Use the following formula to calculate the coordinates of the sound source in the two-dimensional microphone array coordinate system: y=tanθ a *(x-x a )+y a , Where a and b are the microphone arrays corresponding to the remaining two azimuth rays, θ a represents the azimuth angle of microphone array a, θ b represents the azimuth angle of microphone array b, x a 、y a represents the coordinate value of microphone array a, x b 、y b Represents the coordinate value of microphone array b.

8. The sound source localization optimization method based on a distributed multi-microphone array according to claim 1, characterized in that: At least five microphone arrays with the same orientation are deployed equidistantly along the x-axis, and the distance between adjacent microphone arrays does not exceed 3 meters.

9. The sound source localization optimization method based on a distributed multi-microphone array according to claim 1, characterized in that: At least eight microphone arrays are deployed in a rectangular structure, where four microphone arrays are located at the four end points of the rectangle and the remaining four microphone arrays are located at the midpoints of the four sides of the rectangle.

Citation Information

Patent Citations

  • Sound source localization method under strong multipath interference condition

    CN114624652A

  • Sound source positioning method and device, equipment and storage medium

    CN119024270A

  • Sound source localization method based on double-microphone rotating array

    CN119828076A

  • Method and system for acoustic sound localization based on microphone array and coordinate transform method

    KR101645135B1

  • Sound source positioning method and apparatus

    US20230333205A1

Cited By

  • Three-dimensional-TDOA positioning method based on distributed microphone array cooperative positioning

    CN121186707A