A kind of dynamic multi-sound source positioning method of vehicle-mounted microphone array fusing doppler effect compensation
By employing particle swarm optimization algorithm and density clustering technology, the Doppler effect and multi-source aliasing problems of vehicle microphone arrays in dynamic environments were solved, achieving high-precision multi-source localization, which is suitable for environmental perception in intelligent vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV OF SCI & TECH
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
In dynamic vehicle environments, the Doppler effect causes signal distortion and multi-source aliasing, making it difficult for existing technologies to achieve high-precision multi-source localization.
By employing a closed-loop iterative optimization mechanism based on particle swarm optimization and a ring topology niche mechanism, combined with a density clustering algorithm, Doppler effect compensation and multi-source decoupling are achieved. Signal stability is restored through nonlinear time mapping and amplitude correction, and the locations of multiple sound sources are detected in parallel.
It achieves high-precision multi-sound source positioning in dynamic vehicle environments with a positioning error of less than 0.45 meters, exhibits good noise resistance and real-time performance, and is adaptable to complex road scenarios.
Smart Images

Figure CN122362287A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle acoustic perception and intelligent driving technology, specifically relating to a dynamic multi-sound source localization method and system for vehicle microphone arrays for intelligent vehicle environmental perception, which is suitable for solving the problems of signal distortion and multi-sound source aliasing caused by the Doppler effect when the vehicle is in motion. Technical Background
[0002] With the rapid development of intelligent connected vehicle technology, environmental perception systems, as a core component of autonomous driving, directly determine the safety and reliability of vehicles. Currently, mainstream perception solutions primarily rely on line-of-sight sensors such as cameras, LiDAR, and millimeter-wave radar. However, these sensors have physical limitations in non-line-of-sight scenarios. In contrast, sound source localization technology, as a passive perception method, has advantages such as being unaffected by lighting conditions, having a wide coverage area, and possessing non-line-of-sight detection capabilities, making it an important supplement to visual perception. However, applying sound source localization technology to in-vehicle mobile environments faces two major technical challenges:
[0003] (1) Doppler effect interference. The microphone array moves with the vehicle, causing frequency shift and amplitude modulation of the received signal, which disrupts the short-time stationarity of the signal and causes a sharp decline in the performance of traditional time delay estimation algorithms based on generalized cross-correlation.
[0004] (2) Multi-source aliasing problem. In real-world road scenarios, there are often multiple independent sound sources (such as multiple vehicles honking their horns). The signals are mixed in the time and frequency domains, resulting in the objective function exhibiting complex multi-peak characteristics. Traditional localization methods are difficult to achieve multi-target separation and localization.
[0005] Existing research on Doppler effect compensation methods mainly includes frequency domain filtering and time domain interpolation. However, these methods often require prior assumptions about the sound source location, resulting in a nonlinear coupling problem where compensation and localization are interdependent. For multi-source localization, traditional beamforming or subspace methods are computationally complex and difficult to adapt to the real-time requirements of dynamic environments. Therefore, there is an urgent need for a localization method that can simultaneously solve the problems of Doppler effect compensation and multi-source decoupling. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the problems of insufficient Doppler effect processing capability, low multi-source localization accuracy and easy getting trapped in local optima in the dynamic environment of vehicle, and to provide a multi-source localization method and system based on Doppler effect compensation of vehicle microphone array, so as to realize high-fidelity reconstruction of received signals and parallel localization of multiple target sound sources in motion.
[0007] Technical solution
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0009] 1. Modeling of received signals from a vehicle-mounted mobile microphone array
[0010] First, establish such Figure 1 The diagram shows a motion receiving geometry model. It assumes the stationary monopole point source is located at a fixed position, and the vehicle-mounted microphone moves along... With a constant speed in the positive direction of the axis It moves at a constant velocity in a straight line. In an ideal, uniform, and stationary free fluid medium, the speed of sound is... The density of the medium is According to sound field theory, the sound pressure field of a stationary monopole point source satisfies the non-homogeneous wave equation with source terms. Solving using the Green's function, the sound field radiation formula for a stationary point source is:
[0011] (1-1)
[0012] in, The distance between the sound source and the observation point. This represents the total mass flow rate out of the sound source, which is the inherent intensity function of the sound source.
[0013] The instantaneous distance between the sound source and the microphone when the microphone moves with the vehicle. The signal is a time-varying function. Based on the superposition principle and radiation characteristics of the sound field, the movement of the microphone array does not change the radiation mechanism of the sound source itself or the sound pressure distribution in the free field. The moving array only acts as a moving observation point to perform time-varying sampling of the given sound field. Therefore, substituting the time-varying distance into equation (1-1), we obtain the received signal model of the moving microphone:
[0014] (1-2)
[0015] in, This refers to the sound pressure signal received by the motion microphone; This is the distance between the sound source and the microphone. Because... Due to its time-varying characteristics, the received signal undergoes both amplitude modulation and frequency shift, i.e., the Doppler effect.
[0016] 2. Doppler effect compensation algorithm
[0017] For the same sound source, regardless of whether the microphone array moves, its sound source intensity function It is unique. Therefore, by comparing the static receiver model (1-1) and the motion receiver model (1-2), the amplitude correction relationship can be obtained:
[0018] (1-3)
[0019] (1-4)
[0020] To reconstruct the equivalent stationary signal The independent variables of the sound source intensity function in both equations must correspond to the same emission time. If the virtual stationary array is at time... The received signal and the actual moving array at time If the received signal was emitted by the sound source at the same historical moment, then:
[0021] (1-5)
[0022] Equation (1-5) is the amplitude correction model. This equation can be used to dynamically compensate for signal amplitude fluctuations caused by array motion, so that the amplitude characteristics of the reconstructed signal are equivalent to the state when observed at rest.
[0023] To realize the correspondence in equation (1-5), a virtual receiving time must be established. With actual reception time A precise mapping between them. Based on the principle of the uniqueness of emission time, for any sound source at time... The emitted pulse, regardless of whether the microphone array moves, has an objectively unique emission time, which is:
[0024] (1-6)
[0025] in, The moment when the virtual stationary array receives the signal. This refers to the moment when the actual moving array receives the signal.
[0026] Rearranging equations (1-6) to reflect the actual reception time The equation:
[0027] (1-7)
[0028] For any virtual static reception time Solving equation (1-7) yields the corresponding actual motion reception time. Therefore, a nonlinear mapping relationship from the reconstructed time axis to the actual measurement time axis was established (the time-frequency domain comparison before and after compensation is shown in Figure 2).
[0029] 3. "Hypothesis-Compensation-Verification" Closed-Loop Positioning Model
[0030] To decouple the interdependence between Doppler effect compensation and position calculation, this invention proposes a closed-loop iterative optimization mechanism based on the Particle Swarm Optimization (PSO) algorithm.
[0031] Define position deviation function As the optimization objective, for any assumed sound source location... Perform the following operations:
[0032] Will Substitute the data into the motion model and calculate the instantaneous distance between each array element and the assumed sound source;
[0033] The Doppler effect is compensated for using the compensation algorithm in step 2;
[0034] The GCC-PHAT algorithm is used to estimate the time difference of arrival of the compensated signal.
[0035] The location of the sound source was calculated using the Chan algorithm. .
[0036] The position deviation function is defined as follows:
[0037] (1-8)
[0038] When assuming position When approximating the actual sound source location, the Doppler effect is precisely eliminated, and the location is solved. and Towards consistency, positional deviation function The value approaches zero. Therefore, the sound source localization problem is transformed into minimizing a function in continuous space. The optimization problem.
[0039] The PSO algorithm is used to solve the above optimization problem. Assume the population contains... The particle, the first The position vector of each particle Represents the coordinates and velocity vector of a hypothetical sound source. This indicates its update step size. The particle updates by tracking its historical best position. and the global optimal position Perform iterative updates:
[0040] (1-9)
[0041] (1-10)
[0042] in, This represents the number of iteration steps. Inertial weight; , For learning factors; , These are random numbers distributed between [0,1].
[0043] 4. Multi-peak detection based on Niche Particle Swarm Optimization (NPSO)
[0044] To address the multi-peak characteristics of the objective function in multi-source scenarios, traditional PSO is prone to getting trapped in local optima due to global extrema. This invention introduces a ring-shaped topological niche mechanism.
[0045] For each particle Define a restricted ring neighborhood It contains only the particle itself and its neighboring particles on the logic ring:
[0046] (1-11)
[0047] in, To enable the modulo operation, the closed connectivity of the logical loop at the beginning and end of the index is guaranteed.
[0048] Particle velocity updates no longer depend on the global optimum. Instead, it adopts neighborhood optimality. :
[0049] (1-12)
[0050] in, For particles neighborhood The position of the particle with the best fitness among all the individual best positions in history:
[0051] (1-13)
[0052] This mechanism induces the particle swarm to spontaneously differentiate into multiple independent subpopulations, searching different extreme value regions in parallel. The algorithm flow is as follows: Figure 3 As shown.
[0053] 5. Multi-source coordinate extraction based on density-based spatial clustering of applications with noise (DBSCAN)
[0054] Once the NPSO algorithm reaches its maximum number of iterations, it collects the individual historical best positions of all particles to form a candidate solution set. This set contains dense clusters corresponding to multiple sound sources and a small number of outliers caused by residual noise. To automatically extract the sound source coordinates, the DBSCAN density clustering algorithm is used for post-processing.
[0055] Input candidate solution set Set neighborhood radius and minimum points ;
[0056] Calculate each point Neighborhood ;
[0057] like Then mark As the core point;
[0058] Based on density reachability, core points and points connected by density are grouped into the same cluster. ;
[0059] Points that do not belong to any cluster are marked as noise and removed;
[0060] For each valid cluster Calculate its geometric center as the first The final estimated coordinates of each sound source:
[0061] (1-14)
[0062] The beneficial effects of this invention are:
[0063] 1. High-precision Doppler effect compensation: Based on the principle of source intensity function consistency, signal stationarity is effectively restored through joint compensation of nonlinear time mapping and amplitude correction. Simulation results show that the correlation coefficient between the compensated signal and the ideal signal reaches 0.9959, and the compensation effect remains stable within the vehicle speed range of 20 km / h to 80 km / h. It also exhibits good noise resistance robustness within a wide signal-to-noise ratio range of 0 dB to 30 dB.
[0064] 2. Multi-source parallel localization capability: The ring topology NPSO algorithm is introduced to overcome the limitation of traditional PSO being prone to getting trapped in local optima, and to achieve multi-peak parallel detection. Simulation experiments show that under the condition of random distribution of dual sound sources, the maximum localization error does not exceed 0.2 m, and the localization accuracy is not affected by changes in vehicle speed.
[0065] 3. Strong robustness: Combined with DBSCAN clustering to automatically remove noise interference, the dual sound source localization error was controlled within 0.45 m in the real vehicle dynamic test, which verified the effectiveness of the algorithm in the real environment. Attached Figure Description
[0066] Figure 1 Geometric model diagram of motion receiving of vehicle-mounted microphone array
[0067] Figure 2: Time-frequency domain comparison of signals before and after Doppler effect compensation: (a) Time-domain waveform before compensation; (b) Time-domain waveform after compensation; (c) Spectrum before compensation; (d) Spectrum after compensation
[0068] Figure 3 Flowchart of a multi-source localization algorithm based on NPSO and DBSCAN
[0069] Figure 4 Schematic diagram of the geometric position and status of a multi-source localization simulation scene.
[0070] Figure 5 Schematic diagram of cluster distribution after particle swarm optimization iteration
[0071] Figure 6 Actual vehicle test platform image
[0072] Figure 7 Physical image of a cross-shaped microphone array
[0073] Figure 8 Error analysis chart of dual-source measurement and positioning results
[0074] Figure 9 Flowchart of the steps in an embodiment of the present invention Detailed Implementation
[0075] Example 1: Simulation Experiment Verification
[0076] To verify the effectiveness of the method of this invention, a two-dimensional free field simulation environment (such as...) was constructed on the MATLAB platform. Figure 4 (As shown). A 5-element cross-shaped microphone array is used, with an element spacing d = 0.25 m. The array center starts at (-5, 0) m and moves at a speed of 20 km / h along... The axis moves at a constant speed to (5,0) m, and the effective observation distance is 10 m. Two stationary sound sources are set up. (0,3) m transmits an LFM signal of 1000-2000 Hz. The (-2,3) m transmitter emits an LFM signal at 2500-3500 Hz. The sampling rate is 40960 Hz, and the signal-to-noise ratio is 20 dB.
[0077] The location calculation using the method of this invention involves the following steps:
[0078] 1. Initialize NPSO parameters;
[0079] 2. Construct a ring-shaped topological neighborhood and randomly initialize the particle position and velocity.
[0080] 3. For each particle position Using the assumed sound source coordinates as a basis, Doppler effect compensation is performed to obtain the compensated signal.
[0081] 4. Perform GCC-PHAT time delay estimation on the compensated signal and substitute it into the Chan algorithm to calculate the position. .
[0082] 5. Calculate the position deviation according to formula (1-8). Update the individual's historical best and neighborhood optimal .
[0083] 6. Update the particle velocity and position according to equations (1-12) and (1-10).
[0084] 7. After reaching the maximum number of iterations, collect the individual best positions of all particles to form a candidate solution set. .
[0085] 8. The DBSCAN algorithm is used to... Clustering is performed to remove noise points, resulting in two effective clusters (e.g., ...). Figure 5 (As shown).
[0086] 9. Calculate the geometric center of each cluster as the final positioning result.
[0087] The simulation results are as follows:
[0088] sound source Estimated coordinates: (0.0627, 3.0467) m, absolute error 0.078 m;
[0089] sound source Estimated coordinates: (-2.0928, 2.8774) m, absolute error 0.154 m.
[0090] To verify the spatial robustness of the algorithm, five sets of dual sound sources at different locations were randomly selected for testing. The results showed that the maximum positioning error was less than 0.184 m, and the average error was approximately 0.125 m. Comparative experiments at different vehicle speeds (20 km / h to 80 km / h) showed that the positioning accuracy was not affected by changes in vehicle speed. Tests at different signal-to-noise ratios (0 dB to 30 dB) showed that under normal operating conditions above 10 dB, the positioning error remained stable within 0.25 m, demonstrating strong noise resistance.
[0091] Example 2: Real Vehicle Dynamic Test
[0092] To verify the effectiveness of the method of the present invention in a real environment, a system was built as follows: Figure 6 The actual vehicle dynamic testing platform shown. It uses a cross-shaped array of GRAS46AE free-field microphones (e.g., Figure 7(As shown), the array element spacing was 0.25 m. Data acquisition was performed using the LMS Test.Lab system with a sampling rate of 40960 Hz. The experiment was conducted on an open asphalt road surface, with the test vehicle traveling at a constant speed of 20 km / h along a straight line, and an effective observation distance of 10 m. Two Bluetooth speakers were set up as target sound sources, emitting linear frequency modulated signals of 1000-2000 Hz and 2500-3500 Hz respectively, and were fixed at predetermined coordinate positions on both sides of the driving trajectory.
[0093] Five sets of dual-source localization experiments with different spatial distributions were conducted. After acquiring multi-channel signals for each experiment, the signals were imported into the MATLAB platform and processed offline using the method of this invention. Experimental results show that the localization error of all test samples is within 0.45m (error distribution is shown below). Figure 8 (As shown). The sound source... The average error is approximately 0.276 m, and the sound source... The average error is approximately 0.354 m. The algorithm effectively overcomes multi-source aliasing and Doppler effect interference even in real-world complex environments, without target loss or positioning divergence, verifying the practical engineering applicability and robustness of the method.
Claims
1. A method for multi-source localization using a vehicle-mounted microphone array based on Doppler effect compensation, characterized in that, Includes the following steps: Step 1: Establish a received signal model of the vehicle-mounted mobile microphone array and analyze the influence characteristics of the Doppler effect on signal amplitude and frequency; Step 2: Based on the inherent physical properties of the sound source intensity function, a Doppler effect compensation algorithm combining nonlinear time mapping and amplitude correction is proposed to perform time-domain inverse resampling of the motion received signal, eliminate Doppler distortion, and restore the short-time stationarity of the signal. Step 3: Construct a closed-loop localization model based on the "hypothesis-compensation-verification" mechanism, transforming the sound source location calculation into an optimization problem with the location deviation function as the objective function; Step 4: Introduce the Niche Particle Swarm Optimization (NPSO) algorithm. By constructing a one-dimensional ring-shaped neighborhood topology, the social learning range of particles is limited, and spontaneous differentiation of the population is induced, so as to achieve parallel detection of multi-modal objective functions. Step 5: Combine the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm to perform cluster analysis on the candidate solution set after particle swarm iteration, automatically identify the clusters corresponding to multiple sound sources, remove isolated noise points, and calculate the geometric center of each cluster as the spatial coordinate estimate of multiple sound sources. Step 6: Output the spatial location estimation results of the multiple sound sources.
2. The method according to claim 1, characterized in that, The signal receiving model of the vehicle-mounted mobile microphone array established in step 1 is specifically as follows: (1-1) in, This refers to the sound pressure signal received by the motion microphone; Let be the sound source intensity function, representing the total mass flow rate flowing out of the sound source; The time-varying distance between the sound source and the microphone; The speed of sound.
3. The method according to claim 1, characterized in that, The Doppler effect compensation algorithm in step 2 specifically includes: Amplitude correction: Utilizing the consistency of the sound source intensity function, an amplitude correction formula is established: (1-2) in, The signal received by the virtual stationary array; The distance between the virtual stationary array and the sound source; Nonlinear time mapping: Based on the principle of uniqueness of launch time, time-domain constraint equations are established: (1-3) in, This refers to the moment when the virtual stationary array receives the signal; This represents the actual moment when the moving array receives the signal. The actual reception moment is determined by solving this equation. With the moment of reconstruction The mapping relationship; Signal reconstruction: Based on the mapping relationship determined in the above steps, the original distorted signal is non-uniformly resampled, and combined with the amplitude correction formula, a compensated stable signal is obtained.
4. The method according to claim 1, characterized in that, The position deviation function constructed in step 3 Defined as: (1-4) in, Assuming the location of the sound source; For based on After Doppler effect compensation, the position is obtained through time delay estimation using the Generalized Cross-Correlation-Phase Transform (GCC-PHAT) and the Chan algorithm, assuming the position approximates the true sound source. Approaching zero.
5. The method according to claim 1, characterized in that, The particle velocity update formula for the annular topology niche particle swarm optimization algorithm introduced in step 4 is as follows: (1-5) in, This represents the current iteration step number; Inertial weights; , For learning factors; , These are random numbers distributed between [0,1]. For particles The local optimum within the ring neighborhood replaces the global optimum in the traditional Particle Swarm Optimization (PSO) algorithm. .
6. The method according to claim 5, characterized in that, The ring neighborhood topology is defined as follows: For population size The particle swarm, the first The ring neighborhood of each particle for: (1-6) in, To ensure the closed connectivity of the logical loop at the beginning and end of the index, the modulo operation is performed. The social learning range of each particle is strictly limited to its neighborhood, specifically its optimal neighborhood position. Determined by the particle with the best fitness among all its historical best positions in the neighborhood: (1-7)。 7. The method according to claim 1, characterized in that, The DBSCAN density clustering algorithm in step 5 specifically includes the following steps: Neighborhood parameter settings: Set the neighborhood search radius With minimum number of points threshold ; Key point determination: For the candidate solution set Particles in If its - The number of samples contained in the neighborhood satisfies If so, it is determined to be the core point; Cluster generation: Based on density reachability, core points and their neighboring points are grouped into independent clusters. ; Noise removal: Isolated points that do not belong to any cluster are marked as noise and removed; Coordinate calculation: for the first 1 effective cluster Calculate its geometric center as the estimated value of the sound source coordinates; (1-8) in, This represents the total number of effective particles contained within the cluster.