A sound source localization method and system based on four microphones
Through the calculation of the three-dimensional right-angle placement and time difference of four microphones, combined with the generalized cross-correlation function, the accuracy and calculation complexity of microphone array sound source positioning in noise and reverberation environments are solved, and low-cost and high-precision sound source positioning is achieved.
Patent Information
- Application Number
- CN202510780589.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing microphone array sound source positioning method has low accuracy in noise and reverberation environments, high computational complexity, difficult to achieve real-time measurement, and requires a large amount of sound source signal information, resulting in large errors.
Four microphones are placed at three-dimensional right angles to form a triangular shape. Through time difference calculation and generalized cross-correlation function, combined with finding the extreme points of the sound signal, the first envelope signal is extracted, the calculation amount is reduced and the reverberation effect is eliminated, and the sound source positioning is achieved.
It realizes low-cost and high-precision sound source positioning, and can accurately detect the sound source position in complex environments, with small calculation amount and strong real-time performance, reducing the impact of reverb on positioning.
Smart Images

Figure CN120294677B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sound testing, and in particular relates to a sound source localization method and system based on four microphones. Background Art
[0002] With the advancement of technologies such as human-computer interaction and intelligent control, intelligent robots and smart driving are gradually becoming part of people's lives. Accurate sound localization technology is becoming increasingly important in these applications. In bionic engineering and its applications, simulating the human ability to discern subtle sounds and determine the location of objects through binaural hearing, the development of a low-cost, high-speed, and high-precision sound localization system is crucial for the development and application of bionic robots. Compared to devices based on optical signals for localization, sound localization devices offer unique advantages. Because the speed of sound is much slower than the speed of light, the timing requirements for sound localization systems can be significantly reduced. Localization can be accomplished without the need for external infrastructure, reducing initial system investment. Furthermore, sound localization offers advantages over optical methods in light-polluted environments. Sound localization has a wide range of applications, including determining the speaker's location, measuring the position of surrounding objects in autonomous vehicles, and locating vehicles that honked illegally in urban noise monitoring systems.
[0003] Microphone array-based sound source localization is currently the mainstream method for sound localization. It involves using a number of microphones to collect sound signals, then analyzing and processing them using relevant algorithms to determine the location of the sound source. Using arrays for sound source localization involves knowledge from multiple fields, including array signal processing, speech processing, compressed sensing, and artificial intelligence. Microphone array sound source localization primarily studies the direction and distance of a sound source—that is, direction estimation and distance estimation—to uniquely determine the sound source's position in space.
[0004] Currently, three mainstream microphone array sound source localization methods have been developed: the first is the steerable beamforming localization method, the second is the high-resolution spectrum estimation localization method, and the third is the localization method based on time delay estimation.
[0005] Steerable beamforming methods require prior knowledge of the statistical characteristics of the sound source, noise, and reverberation, and use maximum likelihood estimation to determine the source information. Since these characteristics are often estimated, they are prone to large errors. Furthermore, due to the influence of initial values, the objective function may have more than one maximum, resulting in a local optimal solution that fails to meet the requirements of global optimization. Achieving a global optimal solution would significantly increase the algorithm's time complexity. In summary, this method requires extensive knowledge of the sound source signal, and it is difficult to achieve a satisfactory result with both accuracy and time complexity.
[0006] High-resolution spectral estimation localization methods record the correlation matrix of the same sound source signal from different microphones and solve the matrix to obtain the sound signal's location. However, this method is only suitable for cases where the sound source signal is stationary. However, most sound signals in real life do not meet these conditions, making estimation difficult.
[0007] Positioning methods based on time delay estimation have advantages such as low waveform requirements, small amount of computation, low computational complexity, and high accuracy. The basic principle is that due to different microphone positions, the arrival time of the sound source signal will also be different. This time difference and the geometric relationship between the microphones can be used to determine the location of the sound source. Currently, microphone array positioning technology based on time delay is widely used, but this method requires high accuracy in time delay estimation. In noisy and reverberant environments, this method has poor adaptability, resulting in large measurement errors. In addition, the calculation of the entire signal is computationally intensive and time-consuming, making real-time measurement difficult to achieve. Summary of the Invention
[0008] Purpose of the invention: In order to solve the problems existing in the above-mentioned prior art, the present invention provides a sound source localization method and system based on four microphones.
[0009] Technical solution: The present invention provides a sound source localization method based on four microphones, which specifically includes the following steps:
[0010] Step 1: Use four microphones and place them at right angles in three dimensions to form a triangular pyramid. Specifically, select one microphone and place it at the origin of the three-dimensional coordinate system. This microphone is denoted as A. The microphone placed on the X-axis of the three-dimensional coordinate system is denoted as B. The microphone placed on the Z-axis of the three-dimensional coordinate system is denoted as D. The microphone placed on the Y-axis of the three-dimensional coordinate system is denoted as C. The distances between microphones B, C, and D and microphone A are all d.
[0011] Step 2: Collect the sounds from the four microphones and save them in a text file;
[0012] Step 3: Convert the data saved in the text document into an array in MATLAB, and then preprocess the converted data;
[0013] Step 4: Extract the first envelope sound signal of each channel based on the preprocessed data;
[0014] Step 5: Using the intercepted sound signal from microphone A as the reference signal, calculate the time difference between the sound signal received by microphone A and the three microphones B, C, and D.
[0015] Step 6: Locate the sound source based on the time difference calculated in step 5.
[0016] Furthermore, in step 2, a virtual oscilloscope 6000EU is used to simultaneously collect sound signals from four microphones.
[0017] Furthermore, the extraction of the first envelope sound signal of each channel is specifically as follows:
[0018] Step 4.1: For the sound signal of any channel, use the findpeaks function to find all the maximum extreme points of the sound signal of the channel;
[0019] Step 4.2: Take the first maximum extreme point as the base point M and determine whether M is greater than or equal to 2N0. If not, go to step 4.3. If so, intercept the sound signal in the range [M-N0, M+N0]. This signal is the extracted first envelope signal. N0 is the reference parameter, N0 = d × Fs / v, Fs is the sampling frequency, and v is the speed of sound.
[0020] Step 4.3: Return the sound signal within the range [M,5N0] to zero, then go to step 4.1 and search for the base point M again.
[0021] Furthermore, in step 5, a generalized cross-correlation function calculation method is adopted, and the sound signal of microphone A is used as the reference signal to obtain a curve of the cross-correlation function of microphones B, C, D and microphone A, and the peak on the curve is used as the time difference between the corresponding microphones and microphone A receiving the sound signal.
[0022] Furthermore, the step 6 is specifically as follows:
[0023] Step 6.1: Establish the following equation:
[0024] ;
[0025] in, Indicates the time required for the sound to travel from the sound source to microphone A. , as well as are coefficients, , , The expression is as follows:
[0026] ;
[0027] in, Indicates the time difference between microphone B and microphone A receiving the sound signal. Indicates the time difference between microphone C and microphone A receiving the sound signal. It represents the time difference between microphone D and microphone A receiving the sound signal; v represents the speed of sound;
[0028] Step 6.2: Set the following constraints:
[0029] ;
[0030] Step 6.3: Calculate the two roots of the equation in step 6.1 and select the root that is greater than zero as ,based on , get the coordinates of the sound source The expression is as follows:
[0031] .
[0032] A system for a sound source localization method based on four microphones, comprising:
[0033] The sound collection module is used to collect the sounds emitted by the four microphones;
[0034] Data conversion module, used to convert data into arrays in Matlab;
[0035] Preprocessing module, used to preprocess the data in the array in matlab;
[0036] A first envelope sound signal extraction module, configured to extract a first envelope sound signal of each channel;
[0037] The time difference calculation module is used to calculate the time difference between the sound signal received by microphone A and the three microphones B, C, and D;
[0038] The sound source localization module is used to calculate the sound source position based on the time difference.
[0039] Beneficial effect: The present invention uses four microphones placed in a triangular pyramid to collect sound signals. This microphone placement method can solve a unique The value is eliminated, eliminating the influence of the pseudo-roots of the equation on the time calculation. At the same time, compared with other microphone placement methods, this method has a small amount of calculation, and can effectively collect three-dimensional information of the sound, and can detect the spatial position of the sound source at a large angle without blind spots. At the same time, the present invention adopts the method of finding the extreme value to intercept the sound signal in the interval to reduce the amount of calculation. The selection of N0 is constrained by the microphone coordinate d. The selection of the N0 value can ensure that the intercepted sound is not too short to capture the required sound, and at the same time it is not too long to increase the amount of calculation. This is determined by the four microphones placed in the triangular pyramid in the present invention. The distance from the other three microphones to the central microphone is d. To ensure that the same extreme point of different sounds can be captured in the time difference measurement and to leave a certain amount of redundancy, considering extreme cases, the interception time must be at least more than twice its time. Therefore, the method of intercepting sound in the present invention ensures the accuracy and speed of the four-microphone algorithm to calculate the spatial position. The two are indispensable for matching.
[0040] The present invention automatically detects the location of sound sources in real time, achieving low-cost, high-precision measurements. Four microphones are placed at right angles in three dimensions to form a triangular pyramid. The four microphones receive sound signals and process the four collected sound signals. The system searches for extreme values, finding the first extreme point of the data and using it as a base point. Using a method of intercepting the front and back fixed value ranges, the first sound wave envelope information of the sound source signal is extracted, effectively removing the effects of reverberation. The extracted signals from the four microphones are more accurate, and the calculated time difference is more precise, resulting in more accurate sound source positioning. Using Labview to display a three-dimensional image, the sound source position can be intuitively seen from the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a diagram of the microphone placement of the present invention.
[0042] Figure 2 Flowchart of the method of the present invention.
[0043] Figure 3 Graph showing the generalized cross-correlation function of the present invention.
[0044] Figure 4 Schematic diagram of the device of the present invention. DETAILED DESCRIPTION
[0045] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0046] In an empty, uniformly distributed space, sound propagates in a straight line at a constant speed. Based on this principle, a classic time delay estimation algorithm can be used to map the relationship between the distance from the sound source to the microphone and the propagation time.
[0047] like Figure 1 As shown, in the three-dimensional XYZ coordinate system, microphone A is located at the origin, microphones B, C, and D are located at points B, C, and D respectively, and the five-pointed star point O is an arbitrary point in space, serving as the sound source, with the coordinates of the sound source being (x, y, z). The distances from microphones B, C, and D to the center microphone A are all d, the speed of sound is v, and the times at which microphones A, B, C, and D collect the sound signal are respectively 、 、 、 . Then we can list the following system of equations:
[0048]
[0049] In fact, it is difficult to control the generation of sound signals and timing software at the same time, so it is difficult to obtain the time directly. 、 、 、 is very difficult. Instead, the time difference between different microphones is often calculated. Taking the microphone at the origin as the reference, the equations can be changed to:
[0050]
[0051] It represents the time required for the sound to travel from the sound source to microphone A. Among these physical quantities, x, y, z, is an unknown number, d is a constant determined when the device is built, is the time difference between microphone B and microphone A, is the time difference between microphone C and microphone A, is the time difference between microphone D and microphone A. 、 、 All of them are obtained by data analysis, and the above formula can be obtained:
[0052]
[0053] set up is an unknown number, Substitute x, y, and z into the formula In this case, we can obtain the following equation:
[0054]
[0055] Arranging the coefficients, we get the following quadratic equation:
[0056]
[0057] in:
[0058]
[0059] According to the analysis of geometric properties, the relationship between the three sides of the triangle is obtained:
[0060]
[0061] therefore, , is always true, and it also holds true when the sound source is far away from the device (outside the device) Therefore, let the two roots of the quadratic equation be 、 ,satisfy:
[0062] , .
[0063] Since time is a positive number, we can conclude that it must be a positive root. After that, we can calculate x, y, and z:
[0064]
[0065] Expressions based on x, y, and z, 、 、 Precise measurements are required.
[0066] In a real-world environment, sound waves can be reflected by walls, ceilings, and floors. Furthermore, the presence of objects like tables and chairs in a room can easily cause reverberation. Reverberation corresponds to alternative paths, thus calculating different sound path differences. Since a line segment is the shortest distance between two points, a sound signal traveling along a straight line between them will inevitably reach the microphone first, compared to reverberation along a broken line. Therefore, finding the first wave packet on the waveform graph can help accurately locate the sound signal and minimize the effects of reverberation.
[0067] like Figure 2 As shown, the specific process is as follows:
[0068] Sound Signal Acquisition: Using the Virtual Oscilloscope 6000EU, you can simultaneously capture sound signals from four microphones. The oscilloscope's built-in software also supports converting the received analog signals into digital signals and saving them as text files.
[0069] Processing of sound signals and noise: Sound signal processing includes reading sound files, normalization, low-pass filtering, extracting the first envelope of each channel of the sound signal, calculating time difference, calculating the sound source position, and visual display.
[0070] Data reading and normalization: First, convert the data stored in the text document into an array in MATLAB. Then, convert the data from its original format into a computationally scalable array format, restoring the original audio. Next, normalize the data, setting the maximum value to 1 and returning any sound signal less than 5% to zero.
[0071] Filtering (denoising): For sound information, first perform Fourier transform, and then use low-pass filtering to filter out background noise based on the characteristics of the sound signal.
[0072] First Envelope Sound Signal Extraction: Because the generalized cross-correlation algorithm for the entire signal is computationally intensive, to improve computational speed and achieve real-time measurement, it is necessary to select the first small segment of the valid sound signal from the lengthy raw data for calculation. Furthermore, since the sound signal that first reaches the microphone necessarily propagates in a straight line in space, intercepting the sound signal within the small first segment of the valid signal also helps remove the effects of echo and reverberation on the data.
[0073] Using the origin microphone as the reference, use the findpeaks function to find the maximum extreme value of the normalized and filtered data set. Take the first extreme value as the base point, and set the data point corresponding to the base point in the horizontal coordinate of the array in MATLAB as M. The horizontal coordinate M needs to meet the condition: M ≥ 2N 0, N0 is a reference parameter. The following method is used to determine N0: within the measurement distance range of 0.1 to 10 meters, set the sampling rate to Fs, and N0 to d × Fs / v, where Fs is the sampling frequency and v is the speed of sound. If M < 2 N0, zero the sound signal within the range [M, 5N0]. Repeat the above steps to find the extreme value of the zeroed sound signal and determine the value of M. Once M is determined, intercept the sound signal within the range [M-N0, M+N0]. This signal is the extracted first envelope signal. Using the same algorithm, intercept the corresponding length of data for each of the four sound data sets for subsequent calculations.
[0074] Calculate the time difference using the generalized cross-correlation function: Use the generalized cross-correlation function calculation method to calculate the cross-correlation function of microphones B, C, D and microphone A, taking the intercepted sound signal collected by microphone A as a reference. Figure 3 As shown, the cross-correlation functions of microphones B, C, D and microphone A are R BA 、R CA 、R DA There are three curves, each with a peak value, and the horizontal axis corresponding to the peak value is the time difference between the microphones. Figure 3 The three time differences are , , .
[0075] Calculate the sound source position: Using the sound source localization equation and its solution derived above, we can get 、 、 t2, x, y and z can be calculated, and the sound source position (x, y, z) can be obtained, which can be displayed graphically through the software.
[0076] Visualization and integration of results: Using LabVIEW, we can plot a three-dimensional image of the microphone array and the sound source, as well as various required function graphs, such as waveforms and generalized cross-correlation results. Ultimately, by integrating all of these operations into a single LabVIEW program, we create an interface for sound source localization.
[0077] The device of this embodiment is composed of a sound collection module, an amplification module, a power supply module, a virtual oscilloscope, a computer and software, etc. Figure 4 The sound acquisition module consists of a microphone head, and the amplification module uses an amplifier circuit, which is powered by a power supply module. Connect the signals to the CH1 to CH4 ports of the virtual oscilloscope, and connect the virtual oscilloscope to the computer to complete the device.
[0078] To measure the system's accuracy in sound localization at different distances and directions, three exploratory experiments were conducted: measuring the sound location directly in front, at a 45° angle, and at a 30° angle. The multiple measurements were averaged, and the angle θ between the average's coordinate vector and the true value's coordinate vector was calculated. This angle was used as the measurement offset angle. The distance d between the measured and true coordinate points at that location was also calculated.
[0079] Experiment 1 measured the sound source at different distances in front of the device. The three locations were (0.00cm, 100.00cm, 100.00cm), (0.00cm, 80.00cm, 80.00cm), and (0.00cm, 50.00cm, 50.00cm). The actual measurement results of the device at (0.00cm, 100.00cm, 100.00cm) are shown in Table 1 below:
[0080] Table 1
[0081]
[0082] The actual measured results of this device at (0.00cm, 80.00cm, 80.00cm) are shown in Table 2 below:
[0083] Table 2
[0084]
[0085] The actual measured results of this device at (0.00cm, 50.00cm, 50.00cm) are shown in Table 3 below:
[0086] Table 3
[0087]
[0088] Experiment 2 measured the sound source position at different distances in the 45° direction of the device. The three locations were (50.00cm, 50.00cm, 50.00cm), (120.00cm, 120.00cm, 80.00cm), and (140.00cm, 140.00cm, 90.00cm). The actual measured results of this device at (50.00cm, 50.00cm, 50.00cm) are shown in Table 4 below:
[0089] Table 4
[0090]
[0091] The actual measured results of this device at (120.00cm, 120.00cm, 80.00cm) are shown in Table 5 below:
[0092] Table 5
[0093]
[0094] The actual measured results of this device at (140.00cm, 140.00cm, 90.00cm) are shown in Table 6 below:
[0095] Table 6
[0096]
[0097] Experiment 3 measured the sound source position at different distances in the 30° direction of the device. The three locations were (86.00cm, 50.00cm, 50.00cm), (172.00cm, 120.00cm, 90.00cm), and (241.00cm, 140.00cm, 100.00cm). The actual measurement results of this device at (86.00cm, 50.00cm, 50.00cm) are shown in Table 7 below:
[0098] Table 7
[0099]
[0100] The actual measured results of this device at (172.00cm, 120.00cm, 90.00cm) are shown in Table 8 below:
[0101] Table 8
[0102]
[0103] The actual measured results of this device at (241.00cm, 140.00cm, 100.00cm) are shown in Table 9 below:
[0104] Table 9
[0105]
[0106] The relevant calculation results of the previous experiments 1, 2, and 3 are summarized in Table 10 below, and the offset angle, deviation distance, and percentage error of the deviation distance are calculated.
[0107] Table 10
[0108]
[0109] As can be seen from the above table, the offset angle measured by this device at different angles and distances is controlled within 3°, and the percentage error of the deviation distance is controlled within 10%.
[0110] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present invention will not further describe various possible combinations.
Claims
1. A sound source localization method based on four microphones, characterized in that: The specific steps include: Step 1: Use four microphones and place them at right angles in three dimensions to form a triangular pyramid. Specifically, select one microphone and place it at the origin of the three-dimensional coordinate system. This microphone is denoted as A. The microphone placed on the X-axis of the three-dimensional coordinate system is denoted as B. The microphone placed on the Z-axis of the three-dimensional coordinate system is denoted as D. The microphone placed on the Y-axis of the three-dimensional coordinate system is denoted as C. The distances between microphones B, C, and D and microphone A are all d. Step 2: Collect the sounds from the four microphones and save them in a text file; Step 3: Convert the data saved in the text document into an array in MATLAB, and then preprocess the converted data; Step 4: Extract the first envelope sound signal of each channel based on the preprocessed data; Step 5: Using the intercepted sound signal from microphone A as the reference signal, calculate the time difference between the sound signal received by microphone A and the three microphones B, C, and D. Step 6: Locate the sound source based on the time difference calculated in step 5; The extraction of the first envelope sound signal of each channel is specifically as follows: Step 4.1: For the sound signal of any channel, use the findpeaks function to find all the maximum extreme points of the sound signal of the channel; Step 4.2: Take the first maximum extreme point as the base point M and determine whether M is greater than or equal to 2N0. If not, go to step 4.
3. If so, intercept the sound signal in the range [M-N0, M+N0] and use this signal as the first envelope signal to be extracted; N0 is the reference parameter, N0 = d*Fs / v, Fs is the sampling frequency, and v is the speed of sound; Step 4.3: Return the sound signal within the range [M,5N0] to zero, then go to step 4.1 and search for the base point M again.
2. The sound source localization method based on four microphones according to claim 1, characterized in that: In step 2, a virtual oscilloscope 6000EU is used to simultaneously collect the sound signals of the four microphones.
3. The sound source localization method based on four microphones according to claim 1, characterized in that: The preprocessing includes normalization and low-pass filtering.
4. The sound source localization method based on four microphones according to claim 1, characterized in that: In step 5, a generalized cross-correlation function calculation method is used. The sound signal of microphone A is used as a reference signal to obtain a curve of the cross-correlation function of microphones B, C, and D with microphone A. The peak value on the curve is used as the time difference between the sound signal received by the corresponding microphone and microphone A.
5. The sound source localization method based on four microphones according to claim 1, characterized in that: The step 6 is specifically as follows: Step 6.1: Establish the following equation: at O 2 +bt O +c=0; Among them, t O It represents the time required for the sound to travel from the sound source to microphone A. a, b, and c are coefficients. The expressions for a, b, and c are as follows: Where Δt X Indicates the time difference between microphone B and microphone A receiving the sound signal, Δt Y Indicates the time difference between microphone C and microphone A receiving the sound signal, Δt Z It represents the time difference between microphone D and microphone A receiving the sound signal; v represents the speed of sound; Step 6.2: Set the following constraints: Step 6.3: Calculate the two roots of the equation in step 6.
1. Select the root greater than zero and record it as t2. Based on t2, the coordinates (x, y, z) of the sound source are expressed as follows:
6. A system for implementing the four-microphone sound source localization method according to claim 1, characterized in that: include: The sound collection module is used to collect the sounds emitted by the four microphones; Data conversion module, used to convert data into arrays in Matlab; Preprocessing module, used to preprocess the data in the array in matlab; A first envelope sound signal extraction module, configured to extract a first envelope sound signal of each channel; The time difference calculation module is used to calculate the time difference between the sound signal received by microphone A and the three microphones B, C, and D; The sound source localization module is used to calculate the sound source position based on the time difference.
Citation Information
Patent Citations
Three-dimensional space sound source positioning method and positioning system
CN116125389A