Sound source localization method based on microphone array
Through optimization of signal processing algorithms and intelligent optimization technology, combining multiple microphone arrays to collect and process sound field data, and using the least squares method to optimize the positioning results, the problem of the reduction in accuracy and real-time performance of traditional sound source positioning methods in complex environments is solved, and efficient and accurate sound source positioning is achieved.
Patent Information
- Application Number
- CN202411966207.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional sound source positioning methods are susceptible to noise, reverb and multipath effects in complex environments, resulting in reduced positioning accuracy and real-time performance.
Through the optimization of signal processing algorithm and combined with intelligent optimization technology, multiple microphone arrays are used to collect sound field data, perform preprocessing and feature extraction, and optimize positioning results using the least squares method to generate model parameters for sound source positioning.
High-precision sound source positioning is achieved in complex environments, real-time and computing efficiency are improved, and the shortcomings of traditional methods in noise and interference environments are effectively overcome.
Smart Images

Figure CN119936794A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of signal processing and acoustics, and in particular to an efficient sound source localization method based on a microphone array. The method does not rely on deep learning technology, but achieves high-precision and real-time sound source localization through traditional signal processing technology. Background Art
[0002] With the advancement of science and technology, sound source localization technology has been widely used in many fields, including speech recognition, conference systems, smart homes, security monitoring, etc. These application scenarios require the system to accurately and in real time perceive the location of the sound source to improve the quality of interaction and environmental monitoring capabilities. The sound source localization system is usually composed of a microphone array, a signal processing module, and a positioning algorithm. The microphone array is used to capture the sound signals in the environment, the signal processing module pre-processes the sound signals through filtering and feature extraction, and the positioning algorithm uses the processed signal data to calculate the location of the sound source.
[0003] Traditional sound source localization methods mainly rely on technologies such as time difference of arrival (TDOA), beamforming, and high-resolution spectrum estimation. These methods determine the location of the sound source by analyzing the time difference or phase difference between the sound waves arriving at different microphones. The time difference estimation method uses the time difference between the sound waves arriving at the microphone array to infer the location of the sound source. It has high theoretical accuracy, but is susceptible to noise and reverberation in practical applications. Beamforming technology weights and adjusts the phase of the signal received by the microphone array to enhance the sound signal from a specific direction while suppressing interference from other directions. The high-resolution spectrum estimation method achieves high-precision sound source localization by analyzing the spatial spectrum of the signal, but the computational complexity is high.
[0004] These traditional methods can achieve good results in ideal environments, but in complex environments, such as those with multipath reflections, background noise, and dynamic sound sources, the positioning accuracy will drop significantly. In addition, real-time performance and consumption of computing resources are also challenges faced by traditional sound source localization technologies. As application scenarios have increasing requirements for real-time performance and low power consumption, it is particularly important to develop a sound source localization method that can maintain high accuracy and high computational efficiency in complex environments. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides an efficient sound source localization method based on a microphone array. Traditional sound source localization technology is easily affected by factors such as noise, reverberation, and multipath effects in complex environments, resulting in reduced positioning accuracy and real-time performance. The present invention achieves high-precision sound source localization in complex environments by optimizing signal processing algorithms and combining intelligent optimization technology.
[0006] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: a sound source localization method based on a microphone array, comprising the following steps:
[0007] Offline training: The sound field data is collected through a microphone array arranged at a set position, the sound field data is preprocessed, and input into the signal processing model for feature extraction and parameter optimization to generate model parameters for sound source localization;
[0008] Real-time positioning: Collect sound field data in real time, apply the trained signal processing model to obtain the direction and distance of the sound source, and output the positioning results.
[0009] The microphone array is a plurality of microphones arranged at random positions.
[0010] The preprocessing of the sound field data is specifically as follows:
[0011] For the sound field data of multiple microphones, sound field calibration, noise suppression and signal gain adjustment are performed.
[0012] The construction of the signal processing model includes:
[0013] Filtering and denoising module, used to eliminate environmental noise and interference signals;
[0014] A feature extraction module is used to extract the time-frequency characteristics of the sound source signal from the sound field data after filtering and denoising, and identify the key features associated with the sound source;
[0015] The positioning algorithm module calculates the direction and distance of the sound source by analyzing the extracted features and optimizes the positioning results using the least squares method.
[0016] The least square method is used to optimize the positioning result, as follows:
[0017] (xx i ) 2 +(yy i ) 2 =(d i ) 2
[0018] d i =d1+(t1-t i )·V sound
[0019]
[0020] Among them, (x i ,y i ) is the position of the i-th microphone; d i is the distance from the sound source to the i-th microphone; t iis the time when the i-th microphone receives the signal; V sound is the speed of sound, N is the number of microphones, and x and y are the coordinates of the sound source.
[0021] The input signal processing model performs feature extraction and parameter optimization to generate model parameters for sound source localization, specifically:
[0022] The model is trained using a labeled sound source location dataset, and the model parameters, including the coefficients of the least squares method, are obtained through iterative optimization.
[0023] Evaluation indicators used in the model training process include: calculation positioning error, signal-to-noise ratio, and calculation efficiency.
[0024] In the real-time positioning, real-time acquisition of sound field data includes the following steps:
[0025] Collect multi-channel audio signals through microphone array, set sampling rate and recording duration;
[0026] Perform bandpass filtering on the collected audio signal and set the low-pass and high-pass filter frequencies to filter out unnecessary frequency components.
[0027] The method of applying the trained signal processing model to obtain the direction and distance of the sound source includes the following steps:
[0028] The time difference of the sound signal received by each microphone is used to solve the two-dimensional coordinates of the sound source through the least square method;
[0029] Calculate the distance and azimuth between the sound source and the center of the microphone array and output them.
[0030] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for sound source localization based on a microphone array is implemented.
[0031] The present invention has the following beneficial effects and advantages:
[0032] The present invention can effectively overcome the shortcomings of the current traditional sound source localization method that is easily affected by environmental noise and interference, and at the same time improves the accuracy and real-time performance of sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of device coordinates
[0034] Figure 2 It is a detailed flow chart of the method of the present invention. DETAILED DESCRIPTION
[0035] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0036] A method for sound source localization based on a microphone array comprises the following steps:
[0037] 1. First receive the sound source signal:
[0038] Arrange microphone arrays to collect target sound source signals in the environment
[0039] The microphones are arranged in a rectangular array to collect the target sound source signal in the environment; the signal of the microphone channel is collected as the initial signal.
[0040] 2. Secondly, pre-process the sound source signal:
[0041] The obtained sound source signal is detected to determine the frequency range and threshold of the sound source signal. The frequency range and threshold of the sound source signal calculated by the detection are set to the prefabricated bandpass filter module, the parameters of the complete bandpass filter are configured, and the threshold of the sound source signal is set. Finally, the bandpass filter module and the denoising module in the sound source signal preprocessing module are combined to remove unnecessary frequency components and reduce the impact of environmental noise on the test.
[0042] 3. Rough determination of sound source coordinates and distance:
[0043] According to the time when the four microphones receive the sound source signal in the target signal, the time difference between different microphones receiving the sound source signal is calculated. The time difference is used to preliminarily estimate the coordinates and distance of the sound source.
[0044] 4. Determine the more accurate coordinates and distance of the sound source:
[0045] According to the preliminary estimation of the sound source coordinates and distance values, the parameters are updated iteratively for multiple times to gradually reduce the objective function, and finally the sound source coordinate position and distance with the minimum residual square sum are obtained as more accurate coordinates and distance.
[0046] The required formulas include:
[0047] The Fast Fourier Transform is used to convert a signal from the time domain to the frequency domain:
[0048]
[0049] Among them, X(f) is the frequency form of the sound source signal; x(n) is the time domain form of the sound source signal; N is the number of sampling points of the signal, f represents the frequency index, and n is the sampling point index of the time domain signal x(n).
[0050] Calculate the average power of the signal:
[0051]
[0052] The frequency domain signal after fast Fourier transformation is screened by a set average power threshold, and the frequency domain signal greater than the average power threshold is processed in the next step.
[0053] The coordinates and distance of the sound source are roughly determined by the coordinates of the sound source and the distances of each microphone;
[0054] By optimizing the iterative residual function and minimizing the objective function, the sound source position is obtained.
[0055] The detailed process of the present invention is as follows Figure 2 As shown:
[0056] Firstly, a microphone array consisting of four microphones is used to collect sound source signals in the environment to obtain an initial sound source signal with target signal information. Signal measurement and analysis are performed on the collected initial sound source signal to obtain the frequency range and sound threshold of the target signal. The obtained characteristic information, such as the upper and lower frequency limits and threshold of the target signal, is used as parameters of the filter module to preprocess the initial sound source signal with target signal information. The filter module and the denoising module are used to filter out interference signals outside the target signal frequency range and that do not meet the sound threshold.
[0057] like Figure 1 As shown in the figure, after preprocessing the initial signal, the main steps of sound source localization begin; the sound source position and microphone position, the sound source signal is analyzed to obtain the initial time when the four microphones receive the sound source signal, recorded as ΔT1, ΔT2, ΔT3, ΔT4, and the time difference between the two microphones receiving the sound source signal is calculated and recorded as ΔT 12 ,ΔT 13 ,ΔT 14 Roughly calculate the coordinates of the sound source, the formula is as follows
[0058] The distance equation between the sound source coordinates and the first microphone is:
[0059]
[0060] The distance equation between the sound source coordinates and the second microphone is:
[0061] (x p -X2) 2 +(y p -Y2) 2 =(d1-ΔT 12 ·V sound ) 2
[0062] The distance equation between the sound source coordinates and the third microphone is:
[0063] (x p -X3)2 +(y p -Y3) 2 =(d1-ΔT 13 ·V sound ) 2
[0064] The distance equation between the sound source coordinates and the fourth microphone is:
[0065] (x p -X4) 2 +(y p -Y4) 2 =(d1-ΔT 14 ·V sound ) 2
[0066] Among them, (x p ,y p ) is the coordinate of the sound source signal; (X1, Y1) (X2, Y2) (X3, Y3) (X4, Y4) are the coordinates of the four microphones; t1, t2, t3, t4 are the times when the four microphones receive the sound source signal; V sound is the speed of sound in the environment.
[0067] After obtaining the rough sound source coordinates and distance information, they are used as the initial value for accurate positioning of the sound source coordinates. Through multiple iterations, the iteration parameters are updated to finally obtain accurate sound source coordinate information.
[0068] The calculation formula of the optimization iterative algorithm is as follows:
[0069] The residual function of the distance equation between the sound source coordinates and the first microphone is:
[0070]
[0071] The residual function of the distance equation between the sound source coordinates and the second microphone is:
[0072] r2=(x p -X2) 2 +(y p -Y2) 2 -(d1-ΔT 12 ·V sound ) 2
[0073] The residual function of the distance equation between the sound source coordinates and the third microphone is:
[0074] r3=(x p -X3) 2 +(y p -Y3) 2 -(d1-ΔT13 ·V sound ) 2
[0075] The residual function of the distance equation between the sound source coordinates and the fourth microphone is:
[0076] r4=(x p -X4) 2 +(y p -Y4) 2 -(d1-ΔT 14 ·V sound ) 2
[0077] The minimization objective function is as follows:
[0078]
[0079] The optimal solution of the two-dimensional coordinates of the sound source is obtained by minimizing the objective function, and the distance between the sound source and the center of the microphone array (such as Figure 1 The distance and azimuth of the (0,0) point in the image are output.
Claims
1. A sound source localization method based on a microphone array, characterized in that: The steps include: Offline training: The sound field data is collected through a microphone array arranged at a set position, the sound field data is preprocessed, and input into the signal processing model for feature extraction and parameter optimization to generate model parameters for sound source localization; Real-time positioning: Collect sound field data in real time, apply the trained signal processing model to obtain the direction and distance of the sound source, and output the positioning results.
2. A method for sound source localization based on microphone array according to claim 1, characterized in that: The microphone array is a plurality of microphones arranged at random positions.
3. The method for sound source localization based on microphone array according to claim 1, characterized in that: The preprocessing of the sound field data is specifically as follows: For the sound field data of multiple microphones, sound field calibration, noise suppression and signal gain adjustment are performed.
4. The method for sound source localization based on microphone array according to claim 1, characterized in that: The construction of the signal processing model includes: Filtering and denoising module, used to eliminate environmental noise and interference signals; A feature extraction module is used to extract the time-frequency characteristics of the sound source signal from the sound field data after filtering and denoising, and identify the key features associated with the sound source; The positioning algorithm module calculates the direction and distance of the sound source by analyzing the extracted features and optimizes the positioning results using the least squares method.
5. The method for sound source localization based on microphone array according to claim 1, characterized in that: The least square method is used to optimize the positioning result, as follows: (x-x i ) 2 +(y-y i ) 2 =(d i ) 2 d i =d1+(t1-t i )·V sound Among them, (x i ,y i ) is the position of the i-th microphone; d i is the distance from the sound source to the i-th microphone; t i is the time when the i-th microphone receives the signal; V sound is the speed of sound, N is the number of microphones, and x and y are the coordinates of the sound source.
6. The method for sound source localization based on microphone array according to claim 1, characterized in that: The input signal processing model performs feature extraction and parameter optimization to generate model parameters for sound source localization, specifically: The model is trained using a labeled sound source location dataset, and the model parameters, including the coefficients of the least squares method, are obtained through iterative optimization.
7. The method for sound source localization based on microphone array according to claim 1, characterized in that: Evaluation indicators used in the model training process include: calculation positioning error, signal-to-noise ratio, and calculation efficiency.
8. The method for sound source localization based on microphone array according to claim 1, characterized in that: In the real-time positioning, real-time acquisition of sound field data includes the following steps: Collect multi-channel audio signals through microphone array, set sampling rate and recording duration; Perform bandpass filtering on the collected audio signal and set the low-pass and high-pass filter frequencies to filter out unnecessary frequency components.
9. The method for sound source localization based on microphone array according to claim 1, characterized in that: The method of applying the trained signal processing model to obtain the direction and distance of the sound source includes the following steps: The time difference of the sound signal received by each microphone is used to solve the two-dimensional coordinates of the sound source through the least square method; Calculate the distance and azimuth between the sound source and the center of the microphone array and output them.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the sound source localization method based on a microphone array as described in any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Vehicle automatic driving environment sensing sound source positioning method and device
CN120143054A