Vehicle automatic driving environment sensing sound source positioning method and device

By using microphone arrays and deep learning models in the autonomous driving system to calculate the three-dimensional coordinates of the sound source, the problem that the autonomous driving system is difficult to capture emergency sound sources in complex acoustic scenarios is solved, and high-precision and real-time sound source positioning is achieved, which improves driving safety.

CN120143054APending Publication Date: 2025-06-13FUJIAN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510615660.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing autonomous driving systems are difficult to effectively capture emergency sound sources in complex acoustic scenarios, such as ambulance alarm sounds, pedestrian shouts, etc., and the sound source positioning technology is difficult to adapt to vehicle dynamic scenarios, has a low signal-to-noise ratio, high calculation complexity, and is difficult to meet real-time requirements.

Method used

Acoustic signals are collected using a microphone array, the propagation time difference and distance are calculated through adaptive generalized cross-correlation method, the azimuth angle of the sound source is calculated, the acoustic signals are enhanced using the minimum variance and no distortion response algorithm, and the acoustic signals are classified and positioned through a deep learning model, and the three-dimensional coordinates of the sound source are calculated.

Benefits of technology

It realizes high-precision and real-time sound source positioning, can capture sound information that cannot be perceived by vision and radar sensors, enriches the vehicle's perception of the surrounding environment, improves driving safety, and can accurately separate and locate multiple sound sources in complex scenarios to suppress noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143054A_ABST
    Figure CN120143054A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle automatic driving environment sensing sound source positioning method and device. The method comprises the steps that sound signals around a vehicle are collected through a microphone array; calculating a propagation time difference between target microphone pairs in the microphone array through an adaptive generalized cross-correlation method, and obtaining a distance between the target microphone pairs; calculating a sound source azimuth angle of the sound signal according to the propagation time difference and the distance; enhancing a sound signal from the azimuth angle of the sound source through a minimum variance undistorted response algorithm; classifying the enhanced sound signals through a deep learning model to obtain the types of the sound signals; judging whether the type of the sound signal is a target type or not, and if yes, calculating a sound source distance through the enhanced sound signal; and calculating the three-dimensional coordinates of the sound signal according to the sound source azimuth angle and the sound source distance. And high-precision and high-real-time sound source positioning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method and device for sound source localization for environmental perception in vehicle autonomous driving. Background Art

[0002] The visual and radar sensors adopted by existing autonomous driving systems have many limitations in practical applications. In complex acoustic scenarios, such as tunnels, urban dense areas and other environments, the autonomous driving system will not be able to effectively capture emergency sound sources, such as important information like ambulance sirens and pedestrian shouts.

[0003] At the same time, there are special problems such as the Doppler effect and high-speed moving noise interference in the propagation of sound sources in the vehicle dynamic scenario. Most of the existing sound source localization technologies are developed for static or indoor environments, so it is difficult to directly apply them to the vehicle dynamic scenario. Moreover, the engine noise, wind noise generated by the vehicle itself and environmental traffic noise, etc., will lead to a decrease in the signal-to-noise ratio and increase the difficulty of sound source localization.

[0004] In addition, sound source localization faces a series of complex problems such as environmental noise interference, sound wave reflection and multi-source sound signal separation. In the actual road environment, there are often multiple sound sources around the vehicle, such as the honking of adjacent vehicles and emergency alarms. The interference between sound sources makes it extremely difficult to accurately separate and localize. The autonomous driving system needs to have a response speed of milliseconds to cope with various emergencies. However, the existing sound source localization algorithms have a high computational complexity and it is difficult to meet the real-time requirements. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to provide a method and device for sound source localization for environmental perception in vehicle autonomous driving, so as to achieve high-precision and strong real-time sound source localization.

[0006] To solve the above technical problem, the technical solution adopted by the present invention is: A method for sound source localization for environmental perception in vehicle autonomous driving, comprising: Collecting sound signals around the vehicle through a microphone array; Calculating the propagation time difference between target microphone pairs in the microphone array through the adaptive generalized cross-correlation method, and obtaining the distance between the target microphone pairs; Calculating the sound source azimuth angle of the sound signal according to the propagation time difference and the distance; Enhancing the sound signal from the sound source azimuth angle through the minimum variance distortionless response algorithm; Classifying the enhanced sound signal through a deep learning model to obtain the sound signal type; Determine whether the type of the sound signal is the target type. If so, calculate the sound source distance based on the enhanced sound signal. Calculate the three-dimensional coordinates of the sound signal based on the sound source azimuth angle and the sound source distance.

[0007] To solve the above technical problems, another technical solution adopted by the present invention is: A sound source localization device for vehicle autonomous driving environment perception includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above-mentioned method for sound source localization in vehicle autonomous driving environment perception.

[0008] The beneficial effects of the present invention are as follows: The microphone array can capture sound information that cannot be perceived by visual and radar sensors, enriching the vehicle's perception ability of the surrounding environment; after calculating the propagation time difference between target microphone pairs by the adaptive generalized cross-correlation method, the sound source azimuth angle is calculated in combination with the distance between the target microphone pairs, and then the sound signal is enhanced by the minimum variance distortionless response algorithm, so as to enhance the target sound signal in the sound source while suppressing noise and interference, and classify the sound signal through the deep learning model to realize the separation and classification of multi-source signals in complex scenarios. After obtaining the target sound signal, calculate the sound source distance based on the sound signal and calculate the three-dimensional coordinates of the sound signal in combination with the sound source azimuth angle to realize the localization of the target sound signal. Description of the Drawings

[0009] Figure 1 It is a step flow chart of a method for sound source localization in vehicle autonomous driving environment perception in an embodiment of the present invention; Figure 2 It is another step flow chart of a method for sound source localization in vehicle autonomous driving environment perception in an embodiment of the present invention; Figure 3 It is a scene schematic diagram of a method for sound source localization in vehicle autonomous driving environment perception in an embodiment of the present invention; Figure 4 It is a structural schematic diagram of a sound source localization device for vehicle autonomous driving environment perception in an embodiment of the present invention. Detailed Embodiments

[0010] To describe in detail the technical content, achieved objectives and effects of the present invention, the following is described in conjunction with the embodiments and accompanied by the drawings.

[0011] A method for sound source localization in vehicle autonomous driving environment perception includes: Collect the sound signals around the vehicle through a microphone array; Calculate the propagation time difference between target microphone pairs in the microphone array through the adaptive generalized cross-correlation method, and obtain the distance between the target microphone pairs; Calculate the sound source azimuth angle of the sound signal based on the propagation time difference and the distance; Enhance the sound signal from the sound source azimuth angle through the minimum variance distortionless response algorithm; Classify the enhanced sound signal through a deep learning model to obtain the sound signal type; Judge whether the sound signal type is the target type. If so, calculate the sound source distance through the enhanced sound signal; Calculate the three-dimensional coordinates of the sound signal based on the sound source azimuth angle and the sound source distance.

[0012] As can be seen from the above description, the beneficial effects of the present invention are as follows: The microphone array can capture sound information that cannot be perceived by visual and radar sensors, enriching the vehicle's perception ability of the surrounding environment; after calculating the propagation time difference between target microphone pairs through the adaptive generalized cross-correlation method, the sound source azimuth angle is calculated by combining the distance between the target microphone pairs, and then the sound signal is enhanced through the minimum variance distortionless response algorithm, so as to enhance the target sound signal in the sound source while suppressing noise and interference, and classify the sound signal through a deep learning model to realize the separation and classification of multi-source signals in complex scenes. After obtaining the target sound signal, the sound source distance is calculated through the sound signal and the three-dimensional coordinates of the sound signal are calculated by combining the sound source azimuth angle, realizing the positioning of the target sound signal.

[0013] Furthermore, each microphone pair in the microphone array satisfies the constraint condition:

[0014] wherein, is the distance between microphone i and microphone j in the target microphone pair in the microphone array; c is the propagation speed of sound waves in the medium; is the highest operating frequency of the microphone array.

[0015] As can be seen from the above description, the distance between any two microphones in the microphone array is constrained based on the constraint condition, so that each microphone can form an effective microphone pair with other microphones. Such a layout can ensure that multi-directional sound signals can be effectively captured within the 360° sound field range, providing a basis for accurately calculating the time delay difference.

[0016] Furthermore, the calculating the propagation time difference between target microphone pairs in the microphone array through the adaptive generalized cross-correlation method includes:

[0017] where τ represents the search volume, represents the propagation time difference between the target microphone pairs; is the signal of microphone i in the microphone pair in the frequency domain state; is the signal of microphone j in the microphone pair in the frequency domain state; is the noise suppression coefficient.

[0018] As can be seen from the above description, by introducing an adaptive weighting function in the Generalized Cross-Correlation method (GCC-PHAT) to estimate the time delay difference between microphone pairs, the problem that the time delay difference between microphone pairs cannot be accurately obtained due to factors such as the Doppler effect and high-speed moving noise interference affecting sound propagation is solved.

[0019] Furthermore, it also includes: dynamically adjusting the noise suppression coefficient through a Kalman filter.

[0020] As can be seen from the above description, by dynamically adjusting the real-time noise suppression coefficient through a Kalman filter, the weighting can be optimized in different noise environments according to the current noise observation value, effectively suppressing the interference of noise on time delay estimation and improving the accuracy of time delay estimation.

[0021] Furthermore, the calculation of the sound source azimuth angle of the sound signal based on the propagation time difference and the distance includes:

[0022] where, is the sound source azimuth angle; c is the propagation speed of sound waves in the medium; represents the propagation time difference between the target microphone pairs; is the distance between microphone i and microphone j in the target microphone pair in the microphone array.

[0023] As can be seen from the above description, based on the propagation time difference and the distance between the target microphone pairs, the sound source azimuth angle of the sound signal can be accurately calculated.

[0024] Furthermore, the enhancement of the sound signal from the sound source azimuth angle through the Minimum Variance Distortionless Response algorithm includes: ; ; where, is the weight vector of the beamformer; R is the covariance matrix of the sound signal; is the array response vector at the sound source azimuth angle ; is the conjugate transpose of; is the variance of the enhanced acoustic signal, is the conjugate transpose of the beamformer weights.

[0025] As can be seen from the above description, by enhancing the signal of the target sound source with minimum variance distortionless response, the response of the microphone array to the target sound source can be maximized by adjusting the weights of the array, while suppressing the noise and interference in other directions.

[0026] Furthermore, calculating the sound source distance based on the enhanced acoustic signal includes:

[0027] where D is the sound source distance; is the intensity of the acoustic signal received by the microphone array, is the initial intensity of the acoustic signal of the sound source, is the air attenuation coefficient, is the propagation path.

[0028] As can be seen from the above description, calculating the sound source distance based on the intensity of the acoustic signal received by the microphone array, the initial intensity of the sound source, the air attenuation coefficient, and the propagation path can obtain an accurate sound source distance.

[0029] Furthermore, calculating the three-dimensional coordinates of the acoustic signal based on the sound source azimuth angle and the sound source distance includes: ; where D is the sound source distance; is the sound source azimuth angle; is the angle between the sound source and the front of the vehicle.

[0030] As can be seen from the above description, calculating the three-dimensional coordinates through the sound source distance, the sound source azimuth angle, and the angle between the sound source and the front of the vehicle can obtain the accurate three-dimensional coordinates of the acoustic signal.

[0031] Furthermore, determining whether the type of the acoustic signal is the target type further includes: If the target type is noise, the recursive least squares method is used to process the enhanced acoustic signal to obtain a noise spectrum model.

[0032] As can be seen from the above description, after obtaining the noise signal, the recursive least squares method is used to process the noise to obtain a noise spectrum model, so that the noise signal in the acoustic signal can be effectively identified based on the noise spectrum model.

[0033] Another embodiment of the present invention provides a sound source localization device for vehicle autonomous driving environment perception, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, each step in a sound source localization method for vehicle autonomous driving environment perception as described above is implemented.

[0034] The sound source localization method and device for vehicle autonomous driving environment perception provided by the present invention can be applied to the scenario of autonomous driving, which will be described below through specific embodiments: Embodiment 1 Please refer to Figure 1 and Figure 2 , a sound source localization method for vehicle autonomous driving environment perception, including: S1. Collect the sound signals around the vehicle through a microphone array; in order to cover a 360° sound field and comprehensively capture multi-directional sound signals, in this embodiment, a hybrid layout of a circular array and a linear array is adopted to set the microphone array; this hybrid layout combines the advantages of the circular array in omnidirectional perception and the high-resolution characteristics of the linear array in a specific direction; taking the microphone array evenly distributed around the vehicle as an example, each microphone can form a microphone pair with other microphones to ensure that multi-directional sound signals can be effectively captured within the 360° sound field range, providing a basis for accurately calculating the time delay difference; among them, each microphone pair in the microphone array satisfies the constraint condition:

[0035] Among them, is the distance between microphone i and microphone j in the target microphone pair in the microphone array (the distance between any two microphone sensors); c is the propagation speed of sound waves in the medium; is the highest operating frequency of the system, that is, the maximum frequency that the microphone array can receive and process. The above constraint condition describes the distance constraint condition between microphone sensors in the array, which is mainly related to the propagation speed of sound waves, the design of the array, and the frequency characteristics of the signal.

[0036] For the selection of microphones, in this embodiment, in-vehicle anti-noise microphones are selected, whose dynamic range ≥ 120dB, and can adapt to a wide range of changes from weak sounds to high-intensity sounds; the frequency response covers 50Hz - 20kHz, and various sound signals within the human auditory range can be captured; each microphone records the sound signal , where i is the microphone number, is the signal in the time domain.

[0037] S2. Calculate the propagation time difference between target microphone pairs in the microphone array through the adaptive generalized cross-correlation method, and obtain the distance between the target microphone pairs. Since sound propagation in complex vehicle dynamic scenarios is affected by factors such as the Doppler effect and high-speed moving noise interference, it is difficult for traditional methods to accurately obtain the time delay difference between microphone pairs. In this embodiment, an adaptive weighting function is introduced into the generalized cross-correlation method (GCC-PHAT) to estimate the time delay difference between microphone pairs, and the noise suppression coefficient in the adaptive weighting function can be dynamically adjusted in real time through a Kalman filter, and the dynamically adjusted noise suppression coefficient can optimize the weighting according to the current noise observation value in different noise environments, effectively suppressing the interference of noise on time delay estimation, thereby improving the accuracy of time delay estimation. The algorithm is specifically expressed as follows:

[0038] where τ is the search quantity, represents the propagation time difference between target microphone pairs; is the signal of microphone in the microphone pair in the frequency domain state; is the signal of microphone j in the microphone pair in the frequency domain state. Since there are multiple τ values in argmax, the maximum τ value obtained is denoted as τij; is the noise suppression coefficient; ; ; where, is the noise suppression coefficient at time k, is the noise suppression coefficient at the next moment; K is the Kalman gain, is the current noise observation value, is the observation matrix; is the current noise observation value, which is the key data reflecting the noise condition at the current moment; where is the true state of the microphone array at time k, is the measurement noise at time k. At the same time, further determine whether the calculated time delay difference (propagation time difference) is valid. If it is invalid, check the acquisition device of the error point information, re-upload the error point and keep the relevant devices in a real-time update state.

[0039] S3. Calculate the sound source azimuth angle of the sound signal according to the propagation time difference and the distance, and the formula is expressed as follows:

[0040] where, is the azimuth angle of the sound source; c is the propagation speed of sound waves in the medium; represents the propagation time difference between the target microphone pairs; is the distance between microphone i and microphone j in the target microphone pair in the microphone array.

[0041] S4. Enhance the sound signal from the azimuth angle of the sound source through the minimum variance distortionless response algorithm; that is, in this embodiment, the minimum variance distortionless response (MVDR) algorithm in beamforming technology is combined to enhance the signal of the target sound source while suppressing noise and interference; beamforming technology can adjust the weights of the array to maximize the response of the array to the target sound source and suppress noise and interference in other directions; the goal of the MVDR algorithm is to minimize the power of the output signal on the premise of ensuring that the desired signal passes through without distortion, so as to suppress interference and noise.

[0042] Given a microphone sensor array containing N microphones, assuming the covariance matrix of the signal is R and the direction of the target signal is θ, the goal of beamforming is to minimize the variance of the output signal while ensuring a distortionless response to the target signal. Then the MVDR beamforming weight can be calculated by the following formula: ; where, is the weight vector of the beamformer; is the covariance matrix of the sound signal; is at the azimuth angle of the sound source the array response vector under; the direction represents the azimuth angle of the sound source relative to the vehicle. There are multiple microphones in the microphone array. Therefore, different directions correspond to different combinations of signals received by the microphones; in the actual scenario, when sound sources in different directions emit sounds and reach the microphone array, the time and intensity of the sounds received by each microphone will be different; due to the different positions of the microphones in the array, taking a sound source in a specific direction as an example, the microphones on the side closer to the sound source will receive the sound first and the signal intensity will be relatively large; while the microphones on the other side will receive the sound with a time lag and the signal intensity will also decay due to the increased propagation distance; is the conjugate transpose of.

[0043] The variance of the output signal calculated by the MVDR beamformer is the smallest and responds to the target signal without distortion. The variance of the final output signal is expressed as: ; is the variance of the enhanced sound signal, is the conjugate transpose of the beamformer weights.

[0044] S5. Classify the enhanced acoustic signal through a deep learning model to obtain the type of the acoustic signal. For example, a deep learning model of a convolutional recurrent neural network (CRNN) is used to separate and classify multi-source signals in complex scenarios. To further improve the multi-source separation accuracy, spatio-temporal attention weights are introduced in the feature fusion layer , and the key source signal is dynamically focused through the attention mechanism to improve the multi-source separation accuracy. The specific formula is as follows: ; where is the time-domain feature vector and is the frequency-domain feature vector, and is the trainable weight matrix.

[0045] Meanwhile, to improve the generalization ability of the model in a dynamic environment, a vehicle noise dataset such as engine noise, wind noise, etc. can be introduced for adversarial training, so that the model can better adapt to various noise interferences in the actual road environment. Among them, when it is determined that the acoustic signal classification fails or the localization recognition fails, the beamforming parameters are adjusted, and the model is retrained or the model parameters are optimized.

[0046] S6. Determine whether the type of the acoustic signal is the target type. If so, calculate the source distance through the enhanced acoustic signal. For example, when the target type is the sound of a car horn, the following calculation is performed:

[0047] where D is the source distance; P 0 is the initial acoustic signal intensity of the source, α is the air attenuation coefficient, and d is the propagation path; P r is the acoustic signal intensity received by the microphone array. The specific calculation method is as follows: ; where E is the energy, which is obtained by integrating the time-domain signal S i (t) over the time period from t 1 to t 2 , or by integrating the denoised time-domain signal over the time period from t 1 to t 2 ; T = t 2 - t 1 ; A is the area where the acoustic energy is evenly distributed when the sound propagates to the position of the microphone. Since the sound source is a point source, when the sound wave propagates in all directions, the acoustic energy is evenly distributed in an expanding spherical region. Therefore, the distribution of the acoustic energy in space is spherically symmetric, that is, the area is expressed as A = 4 .

[0048] S7. Calculate the three-dimensional coordinates of the acoustic signal based on the azimuth angle and distance of the sound source. The specific calculation method is as follows: ; where D is the distance of the sound source; is the azimuth angle of the sound source; is the angle between the sound source and the front of the vehicle, which can be calculated through the azimuth angle of the sound source For example, if the azimuth angle of the sound source is 90°, then is also 90°; for example, if the azimuth angle of the sound source is 270°, then is still 90°. As shown in Figure 3 , the three-dimensional coordinates (x1, y1, z1) of the acoustic signal emitted by the first vehicle in the front and the three-dimensional coordinates (x2, y2, z3) of the acoustic signal emitted by the second vehicle can be respectively identified.

[0049] Meanwhile, update the background noise model in real time according to the scene, and introduce data augmentation technology to expand the training data. Then, use the recursive least squares method to process the enhanced acoustic signal to obtain the noise spectrum model, which is specifically expressed as follows: ; where is the noise covariance matrix, is the forgetting factor (usually taken as 0.95 - 0.99), is the frequency domain vector of the current frame noise signal; the recursive update strategy can quickly adapt to the sudden noise environment. In practical applications, the background noise model can be updated in real time; for example, update the background noise spectrum every 100 ms, and use the subspace tracking algorithm to separate the transient noise and the steady-state noise to enhance the robustness of the system in a noisy environment. At the same time, introduce data augmentation technology to expand the training data, use a vehicle-mounted simulation platform (such as CARLA) to generate multi-scene acoustic data such as rainy days and tunnel echoes to expand the training set, and improve the adaptability of the model to various sound sources. And upload the data to the cloud platform for recording, so that the driverless vehicle makes corresponding adjustments and feeds back to the driverless system in real time to update the state.

[0050] Embodiment 2 Please refer to Figure 4 , a sound source localization device for vehicle autonomous driving environment perception, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes each step in a method for sound source localization in a vehicle autonomous driving environment perception as described in Embodiment 1.

[0051] In summary, the method and device for sound source localization in the vehicle's autonomous driving environment provided by the present invention innovatively introduce auditory information into the autonomous driving system, bringing a new dimension to the vehicle's perception ability. Compared with traditional autonomous driving systems that only rely on visual and radar sensors, introducing auditory information into the autonomous driving system can capture sound information that cannot be perceived by visual and radar sensors, such as emergency vehicle sirens, pedestrian shouts, etc., thus enriching the vehicle's perception ability of the surrounding environment and improving driving safety. And through the optimized microphone array layout, accurate time delay estimation method, and multi-source separation and localization technology, the position of the sound source can be accurately determined, providing a reliable decision-making basis for the autonomous driving system. Combined with the deep learning separation and filtering technology, it effectively addresses the problems of multi-source interference and noise, enabling accurate separation and localization of multiple sound sources in complex road environments such as urban dense areas and noisy traffic intersections, while suppressing various noise interferences to ensure the robustness of the system in complex scenarios.

[0052] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for locating sound sources in an environment-aware vehicle autonomous driving system, characterized in that: include: Collecting acoustic signals around the vehicle through a microphone array; Calculating the propagation time difference between the target microphone pairs in the microphone array by an adaptive generalized cross-correlation method, and obtaining the distance between the target microphone pairs; Calculating the sound source azimuth of the sound signal according to the propagation time difference and the distance; enhancing the acoustic signal from the azimuth of the sound source by a minimum variance distortion-free response algorithm; Classifying the enhanced acoustic signal through a deep learning model to obtain the acoustic signal type; Determine whether the acoustic signal type is a target type, and if so, calculate the sound source distance using the enhanced acoustic signal; The three-dimensional coordinates of the sound signal are calculated according to the sound source azimuth and the sound source distance.

2. The method for locating a sound source in an environment-aware vehicle autonomous driving system according to claim 1, characterized in that: Each microphone pair in the microphone array satisfies the constraint: in, is the distance between microphone i and microphone j in the target microphone pair in the microphone array; c is the propagation speed of sound waves in the medium; is the highest operating frequency of the microphone array.

3. The method for locating a sound source for vehicle autonomous driving environment perception according to claim 1, characterized in that: The calculating the propagation time difference between the target microphone pairs in the microphone array by the adaptive generalized cross-correlation method comprises: Among them, τ represents the search volume, represents the propagation time difference between the target microphone pair; is the signal of microphone i in the microphone pair in the frequency domain; is the signal of microphone j in the microphone pair in the frequency domain; where f is the actual frequency; is the noise suppression coefficient.

4. The method for locating a sound source for vehicle automatic driving environment perception according to claim 3, characterized in that: Also includes: The noise suppression coefficient is dynamically adjusted through a Kalman filter.

5. The method for locating a sound source for vehicle autonomous driving environment perception according to claim 1, characterized in that: The step of calculating the sound source azimuth of the sound signal according to the propagation time difference and the distance comprises: in, is the azimuth of the sound source; c is the propagation speed of the sound wave in the medium; represents the propagation time difference between the target microphone pair; is the distance between microphone i and microphone j in the target microphone pair in the microphone array.

6. The method for locating a sound source for vehicle autonomous driving environment perception according to claim 1, characterized in that: The method of enhancing the acoustic signal from the azimuth of the sound source by using the minimum variance distortion-free response algorithm comprises: ; ; in, is the beamformer weight vector; is the covariance matrix of the acoustic signal; is the azimuth of the sound source The array response vector under; yes The conjugate transpose of ; is the variance of the enhanced acoustic signal, is the conjugate transpose of the beamformer weights.

7. The method for locating a sound source for vehicle automatic driving environment perception according to claim 1, characterized in that: Calculating the sound source distance by using the enhanced sound signal comprises: Where D is the distance from the sound source; is the acoustic signal strength received by the microphone array, is the initial sound signal strength of the sound source, is the air attenuation coefficient, For the propagation path.

8. The method for locating a sound source in an environment-aware vehicle autonomous driving system according to claim 1, characterized in that: The step of calculating the three-dimensional coordinates of the sound signal according to the sound source azimuth and the sound source distance comprises: ; Where D is the distance from the sound source; is the azimuth of the sound source; is the angle between the sound source and the front of the vehicle.

9. The method for locating a sound source in an environment-aware vehicle autonomous driving system according to claim 1, characterized in that: The determining whether the acoustic signal type is a target type further includes: If the target type is noise, the enhanced acoustic signal is processed using a recursive least squares method to obtain a noise spectrum model.

10. A vehicle automatic driving environment perception sound source localization device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, each step of the method for localizing sound sources in a vehicle autonomous driving environment perception as described in any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Method for locating a sound source, and humanoid robot using such a method

    CN106030331A

  • Sound source localization method based on microphone array

    CN119936794A

  • Sound source detection device, especially for use in a motor vehicle driver assistance system, detects phase displacements or signal time of flight differences from sound sources that are detected using multiple microphones

    DE102004045690A1

  • Using classified sounds and localized sound sources to operate an autonomous vehicle

    US20200241552A1