A method and device for fusing visual and auditory information in vehicle autonomous driving
By projecting auditory information into the coordinate system of visual information and fusion, combining deep learning models and Bayesian decision theory, the perception problem of autonomous driving systems in occlusion and noise environments is solved, achieving higher perception accuracy and robustness.
Patent Information
- Application Number
- CN202510612651.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing autonomous driving perception system mainly relies on vision and lidar, ignoring the value of auditory signals, resulting in failed target detection under occlusion or inclement weather conditions, and it is difficult for the auditory system to accurately identify the sound source in a noisy environment, affecting the robustness and perception ability of the system.
By obtaining the visual and auditory information of the vehicle, the auditory information is projected into the coordinate system of the visual information, and the attention mechanism is used to fuse it. The camera internal reference matrix and sound source distance are used to convert the sound source direction into the coordinates under the visual coordinate system, and the deep learning model is combined for object detection and positioning.
The perception accuracy and robustness of the autonomous driving system in complex scenarios are improved, and target detection and positioning are carried out by integrating features, and vehicle driving behavior is optimized to improve safety and robustness.
Smart Images

Figure CN120123996B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and device for fusing visual and auditory information for autonomous driving of a vehicle. Background Art
[0002] Current autonomous driving perception systems mainly rely on visual (such as cameras) and laser radar (LiDAR) information, ignoring the value of auditory signals in the perception process. Although some systems apply auditory signals, auditory signals usually work independently from visual signals. The application of a single visual system or auditory system has certain limitations. For example, in the case of occlusion, visual information may not be able to obtain the target object; in severe weather conditions such as rain, snow, and fog, the image quality of the camera decreases, resulting in target detection failure. In noisy environments such as urban traffic noise, it is difficult for the auditory system to accurately identify the sound source, resulting in information loss. Therefore, how to effectively integrate visual and auditory information to improve the robustness and perception capabilities of the autonomous driving system has become a major technical difficulty. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a method and device for fusing visual and auditory information of vehicle automatic driving, so as to improve the accuracy of target detection and positioning of the automatic driving system.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0005] Acquiring visual information and auditory information collected by the vehicle; the auditory information includes the direction of the sound source;
[0006] Projecting the auditory information into the coordinate system of the visual information;
[0007] The visual information and the auditory information in the same coordinate system are fused through an attention mechanism to obtain a fused feature;
[0008] The projecting the auditory information into the visual coordinate system of the visual information comprises:
[0009] Get the intrinsic parameter matrix of the vehicle camera;
[0010] Obtaining a distance value from the vehicle camera to the sound source according to the direction of the sound source;
[0011] The sound source direction is converted into coordinates in the visual coordinate system according to the intrinsic parameter matrix and the distance value.
[0012] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0013] A vehicle autonomous driving visual and auditory information fusion device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above-mentioned vehicle autonomous driving visual and auditory information fusion method.
[0014] The beneficial effects of the present invention are as follows: After obtaining the visual information and auditory information of the vehicle, the auditory information is projected into the coordinate system of the visual information, so that the auditory information and visual information can be effectively fused in the same coordinate system, thereby obtaining accurate fusion features. Finally, target detection and positioning are performed through the fusion features, improving the perception accuracy and robustness of the vehicle in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flowchart of the steps of a vehicle autonomous driving visual and auditory information fusion method in an embodiment of the present invention;
[0016] Figure 2 It is another flowchart of the steps of a vehicle autonomous driving visual and auditory information fusion method in an embodiment of the present invention;
[0017] Figure 3 It is a schematic diagram of the application scenario of a vehicle autonomous driving visual and auditory information fusion method in an embodiment of the present invention;
[0018] Figure 4 It is the structure of a vehicle autonomous driving visual and auditory information fusion device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To describe in detail the technical content, achieved objectives, and effects of the present invention, the following is described in conjunction with the embodiments and accompanied by the drawings.
[0020] A vehicle autonomous driving visual and auditory information fusion method includes:
[0021] Obtain the visual information and auditory information collected by the vehicle; the auditory information includes the sound source direction;
[0022] Project the auditory information into the coordinate system of the visual information;
[0023] Fuse the visual information and the auditory information in the same coordinate system through an attention mechanism to obtain a fusion feature;
[0024] The projecting the auditory information into the visual coordinate system of the visual information includes:
[0025] Obtain the internal parameter matrix of the vehicle camera;
[0026] Obtain the distance value from the vehicle camera to the sound source according to the sound source direction;
[0027] Convert the sound source direction into coordinates in the visual coordinate system according to the internal reference matrix and the distance value.
[0028] As can be seen from the above description, the beneficial effect of the present invention is that after obtaining the visual information and auditory information of the vehicle, based on the internal reference matrix of the camera and the distance value from the vehicle camera to the sound source, the sound source direction is converted into coordinates in the visual coordinate system, that is, by projecting the sound source direction onto the image plane coordinate system, the auditory information and visual information can be effectively fused in the same coordinate system, so as to obtain accurate fusion features, and finally the target detection and positioning are carried out through the fusion features, improving the perception accuracy and robustness of the vehicle to complex scenes.
[0029] Further, the auditory information obtained by the vehicle includes:
[0030] Collect sound signals through a microphone array;
[0031] Extract the sound source direction and auditory features in the sound signal;
[0032] Use the sound source direction and auditory features as the auditory information.
[0033] As can be seen from the above description, the microphone array can effectively collect the audio signals around the vehicle, and extract the sound source direction and auditory features in the sound signal as the auditory information, so as to accurately describe the sound environment around the vehicle.
[0034] Further, the conversion of the sound source direction into coordinates in the visual coordinate system according to the internal reference matrix and the distance value includes:
[0035] ;
[0036] ;
[0037] ;
[0038] where K is the internal reference matrix, f x and f y are the focal lengths, (c x , c y ) are the principal point coordinates; r, Ф are the polar coordinate parameters of the sound source direction; Z c represents the distance value from the vehicle camera to the sound source; x, y are the converted coordinates.
[0039] As can be seen from the above description, after converting the sound source direction into polar coordinates and combining the parameters in the camera internal reference matrix to project and convert the sound source direction, the sound source direction can be effectively converted into coordinates in the visual coordinate system.
[0040] Further, the fusion of the visual information and the auditory information in the same coordinate system through the attention mechanism includes:
[0041] Obtain the visual weight corresponding to the visual information and the auditory weight corresponding to the auditory information;
[0042] Perform weight calculation on the visual information and the auditory information according to the visual weight and the auditory weight to obtain the fusion feature.
[0043] As can be seen from the above description, by separately obtaining the weights corresponding to the visual information and the auditory information and then performing weight calculation according to the weight values, the weights corresponding to the visual information and the auditory information can be dynamically adjusted according to various complex scenarios, improving the adaptability of the fusion feature in complex scenarios.
[0044] Further, the obtaining of the visual weight corresponding to the visual information and the auditory weight corresponding to the auditory information includes:
[0045] ;
[0046] ;
[0047] where, w v represents the visual weight, w a represents the auditory weight; f w is the weight calculation function, W is the weight matrix, and b is the bias term; F v represents the visual feature, and F a represents the auditory feature.
[0048] As can be seen from the above description, by using the weight calculation function to calculate the contribution degrees of the visual information and the auditory information in various complex scenarios respectively, and then determining the weight values of the visual information and the auditory information based on their respective contribution degrees, the adaptability of the fusion feature in complex scenarios can be improved.
[0049] Further, it further includes: performing object detection and localization on the fusion feature through a deep learning model to obtain a detection result.
[0050] As can be seen from the above description, by processing the fusion feature based on the deep learning model, accurate object detection and localization results can be obtained.
[0051] Further, after obtaining the detection result, it further includes:
[0052] Obtain the detection result of the object detection and localization;
[0053] Optimize the vehicle behavior based on the detection results through Bayesian decision theory to obtain a vehicle decision.
[0054] As can be seen from the above description, after obtaining the detection results of target detection and positioning, the driving behavior of the vehicle is optimized through Bayesian decision theory combined with the detection results, so as to obtain the best driving strategy.
[0055] Further, the optimizing the vehicle behavior based on the detection results through Bayesian decision theory includes:
[0056] ;
[0057] ;
[0058] wherein, represents the action that can maximize the expected utility among all possible actions; U is the utility function, A is the action set; R(a, O(t)) represents the reward for choosing action a in state O(t); C(a) represents the cost of choosing action a; w1 and w2 are weight coefficients.
[0059] As can be seen from the above description, based on the above formula, the rewards corresponding to different vehicle actions can be effectively described, so as to obtain the best driving strategy.
[0060] Further, before projecting the auditory information into the coordinate system of the visual information, it further includes:
[0061] Align the visual information and the auditory information in time.
[0062] As can be seen from the above description, by aligning the visual information and the auditory information in time, the visual information and the auditory information can represent data at the same moment.
[0063] Another embodiment of the present invention provides a vehicle autonomous driving visual and auditory information fusion device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above-mentioned vehicle autonomous driving visual and auditory information fusion method.
[0064] The vehicle autonomous driving visual and auditory information fusion method and device provided by the present invention can be applied to the vehicle autonomous driving scenario, which will be described below through specific embodiments:
[0065] Embodiment 1
[0066] Please refer to Figure 1 , a vehicle autonomous driving visual and auditory information fusion method, including:
[0067] S1. Obtain the visual information and auditory information collected by the vehicle. Among them, for the collection of visual information, a high-resolution camera is used to collect the RGB image sequence I(t), and the target detection box and feature F v (t) are extracted through a deep learning model (such as YOLO). To perform target detection using the YOLO model, the model structure is as follows:
[0068] Input layer: Receive the image I(t).
[0069] Feature extraction layer: Use a convolutional neural network (CNN) to extract features.
[0070] Detection layer: Output the target detection box and class probability.
[0071] Feature extraction formula: F v (t)=f v (I(t))=CNN(I(t)); where f v is the visual feature extraction function, and CNN is the convolutional neural network.
[0072] For the collection of auditory information, a microphone array is used to collect the acoustic signal S(t), and the sound source direction θ(t) and feature F a (t) are estimated through the generalized cross-correlation method (GCC-PHAT). Specifically:
[0073] Sound source direction estimation: Use the GCC-PHAT algorithm for sound source localization, and the formula is as follows:
[0074] ;
[0075] Among them, S1 and S2 are the two microphone signals of the microphone array, and τ is the delay.
[0076] Feature extraction formula:
[0077] F a (t)=f a (S(t))=GCC-PHAT(S(t));
[0078] Among them, f a is the auditory feature extraction function; GCC-PHAT() is the generalized cross-correlation method formula.
[0079] S2. Align the visual information and the auditory information in time. Specifically:
[0080] Obtain the time stamp T v corresponding to the visual information, and the time stamp T a corresponding to the auditory information. According to the time stamps T v and T a , align the visual and auditory features in time:
[0081] ;
[0082] Among them, is the allowable time error range, usually set to 10 milliseconds.
[0083] S3. Project the aligned auditory information into the coordinate system of the visual information; in order to enable the auditory information and the visual information to be fused and processed in the same coordinate system, it is necessary to project the sound source direction θ(t) into the image plane coordinate system, expressed as:
[0084] ;
[0085] Among them, F x , F y is the spatial transformation function; the specific form of the spatial transformation function is as follows:
[0086] Let the internal parameter matrix of the camera be K, and its form is:
[0087] ;
[0088] Among them, f x and f y are the focal lengths, (c x , c y ) is the principal point coordinate;
[0089] The sound source direction θ(t) is represented in polar coordinates as (r, Ф), where r is the distance and Ф is the angle relative to a certain reference direction, then the projection formula is expressed as:
[0090] ;
[0091] ;
[0092] Among them, Z c represents the distance value from the vehicle camera to the sound source; the distance value Z c is obtained from the distance r; for example, it is calculated according to the distance r, the distance difference and the angle difference between the camera and the microphone; or taking the vehicle as a whole as the reference point, the distance r can be directly used as the distance value Z c .
[0093] x, y are the transformed coordinates.
[0094] S4. Fuse the visual information and the auditory information in the same coordinate system through the attention mechanism to obtain the fused feature; that is, use the attention mechanism to fuse the visual feature F v (t) and the auditory feature F a(t) to enhance the perception ability of multimodal information. Specifically:
[0095] S41. Obtain the visual weight corresponding to the visual information and the auditory weight corresponding to the auditory information. In this embodiment, the attention weights w v and w a are used to dynamically adjust the contributions of audiovisual features, and the specific formula is expressed as:
[0096] ;
[0097] where f w is a weight calculation function, and in the form based on the attention mechanism is as follows:
[0098] ;
[0099] where w v represents the visual weight, w a represents the auditory weight; f w is a weight calculation function, W is a weight matrix, and b is a bias term.
[0100] S42. Calculate the weights of the visual information and the auditory information according to the visual weight and the auditory weight to obtain the fusion feature, and the specific form is as follows:
[0101] .
[0102] S5. Perform object detection and localization through the fusion feature; that is, based on the fused feature F f (t), perform joint detection and localization on the object. In this embodiment, an object detection and localization are performed on the fusion feature through a deep learning model to obtain a detection result:
[0103] O(t)=f d (F f (t));
[0104] where f d is a detection model, and O(t) is the category and position of the object.
[0105] In this embodiment, the driving decision of the autonomous vehicle is also optimized based on the detection result; that is, after performing object detection and localization through the fusion feature, the following steps are further included:
[0106] Please refer to Figure 2 and Figure 3As shown in the figure, when the vehicle reaches an intersection, step S1 is executed. At this time, the autonomous driving audio-visual fusion system collects visual information in the environment through a camera, such as vehicles, pedestrians, road signs, etc.; at the same time, a microphone array is used to collect traffic noise and sound source signals, such as car honks and emergency vehicle sirens. These data are respectively processed into visual features and auditory features to support subsequent information fusion. Subsequently, step S2 is executed. The visual features and auditory features are synchronously processed through timestamps to ensure the validity of the two types of data in the same time dimension. Then step S3 is executed. Using the spatial position relationship between the camera and the microphone, the auditory information is mapped to the visual coordinate system so that the two types of perception information can be uniformly expressed in the same coordinate system. Step S4 is executed. The weights of the visual features and auditory features are calculated through an attention mechanism, and the influence degrees of the two modalities are dynamically allocated. Then the weighted visual features and auditory features are fused to generate a unified feature representation. Finally, step S5 is executed. The fused features are used for target detection and positioning.
[0107] S6. Obtain the detection results of target detection and positioning;
[0108] S7. Optimize the vehicle behavior based on the detection results through the Bayesian decision theory to obtain a vehicle decision; that is, the system analyzes the fused perception information through the Bayesian theory to predict the potential benefits or risks of different driving behaviors; and selects the optimal driving action according to the target detection results and environmental data to improve the decision-making accuracy and safety of the autonomous driving vehicle in complex scenarios. Specifically:
[0109] Based on the detection result O(t), optimize the vehicle behavior through the Bayesian decision theory:
[0110] ;
[0111] Among them, represents the action that can maximize the expected utility among all possible actions; U is the utility function, and A is the set of actions. Vehicle behavior usually includes multiple parameters, such as speed, direction, acceleration, etc.; by selecting the optimal action , it can be specifically reflected in the following aspects:
[0112] (1) Speed optimization: Assume that the set of actions A contains different speed adjustment options such as accelerating, decelerating, and maintaining the current speed. By calculating the utility value of each action, select the speed adjustment action that can maximize safety and efficiency. For example:
[0113] Acceleration: In the absence of obstacles, choose to accelerate to the maximum safe speed; or choose to accelerate to a specific speed.
[0114] Deceleration: When approaching a traffic light or an obstacle, choose to decelerate to ensure safety; or choose to decelerate to a specific speed.
[0115] Maintain the current speed: Do not change the current speed and continue to drive at the same rate.
[0116] (2) Direction optimization: The action set A can also include direction adjustments such as turning left, turning right, going straight, etc.; by calculating the utility value of each direction adjustment action, select the direction that can maximize driving efficiency and safety, for example:
[0117] Turn left: Choose to turn left when having the right of way at an intersection; or turn left according to the navigation route.
[0118] Turn right: Choose to turn right when the traffic flow is less; or turn right according to the navigation route.
[0119] Go straight: Choose to continue going straight.
[0120] (3) Other behaviors:
[0121] Stop: Choose to come to a complete stop.
[0122] Change lanes: Choose to change from the current lane to the left or right lane.
[0123] Overtake: Choose to overtake the vehicle in front when it is safe.
[0124] Among them, the utility function U(a|O(t)) quantifies the "goodness or badness" of choosing a certain action in a specific state; its specific form can vary according to different application scenarios. The following is an optional example:
[0125] ;
[0126] In the formula, R(a, O(t)) represents the reward for choosing action a in state O(t), which is usually related to factors such as safety and efficiency; C(a) represents the cost of choosing action a, which is usually related to factors such as time and energy consumption; w1 and w2 are weight coefficients used to balance the reward and cost.
[0127] Embodiment 2
[0128] Please refer to Figure 4 , a vehicle autonomous driving visual and auditory information fusion device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes each step in a vehicle autonomous driving visual and auditory information fusion method as described in Embodiment 1.
[0129] In summary, the method and device for fusing visual and auditory information in vehicle autonomous driving provided by the present invention obtain the visual information and auditory information of the vehicle, align the visual information and auditory information in the time dimension, and then project the auditory information into the coordinate system of the visual information, so that the auditory information and visual information can be effectively fused in the same coordinate system. During the fusion process, the attention mechanism is used to dynamically adjust the modal contribution degree according to various complex scenarios to achieve the adaptability to complex scenarios, thereby obtaining accurate fusion features. Finally, object detection and positioning are performed through the fusion features, improving the perception accuracy and robustness of the vehicle for complex scenarios. On this basis, combined with the Bayesian decision theory, the real-time decision-making ability of autonomous driving vehicles is optimized, thus significantly improving the safety and robustness of the autonomous driving system in occlusion and multi-noise scenarios, and further promoting the development of intelligent driving technology.
[0130] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the relevant technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A method for fusing visual and auditory information in vehicle autonomous driving, characterized in that, Including: Obtaining visual information and auditory information collected by the vehicle; the auditory information includes the sound source direction; Projecting the auditory information into the coordinate system of the visual information; Fusing the visual information and the auditory information in the same coordinate system through an attention mechanism to obtain a fused feature; The projecting the auditory information into the visual coordinate system of the visual information includes: Obtaining the internal parameter matrix of the vehicle camera; Obtaining the distance value from the vehicle camera to the sound source according to the sound source direction; Converting the sound source direction into the coordinates in the visual coordinate system according to the internal parameter matrix and the distance value; The converting the sound source direction into the coordinates in the visual coordinate system according to the internal parameter matrix and the distance value includes: where K is the intrinsic matrix, f x and f y are the focal lengths, (c x , c y ) are the principal point coordinates; r, Ф are the polar coordinate parameters of the sound source direction; Z c represents the distance value from the vehicle camera to the sound source; x, y are the transformed coordinates; The fusing the visual information and the auditory information in the same coordinate system through an attention mechanism includes: Obtaining the visual weight corresponding to the visual information and the auditory weight corresponding to the auditory information; Performing weight calculation on the visual information and the auditory information according to the visual weight and the auditory weight to obtain the fused feature; The obtaining the visual weight corresponding to the visual information and the auditory weight corresponding to the auditory information includes: f w (F) = softmax(W·F x + b), x = v or a; Among them, w v represents the visual weight, and w a represents the auditory weight; f w is the weight calculation function, W is the weight matrix, and b is the bias term; F v represents the visual feature, and F a represents the auditory feature.
2. The vehicle automatic driving visual and auditory information fusion method according to claim 1, characterized in that The obtained auditory information of the vehicle includes: Collecting sound signals through a microphone array; Extracting the sound source direction and the auditory feature in the sound signals; Taking the sound source direction and the auditory feature as the auditory information.
3. A method for fusing visual and auditory information in vehicle autonomous driving according to claim 1, characterized in that, It further includes: Performing target detection and positioning on the fused feature through a deep learning model to obtain a detection result.
4. A method for fusing visual and auditory information for vehicle autonomous driving according to claim 3, characterized in that After obtaining the detection result, it further includes: Obtaining the detection result of the target detection and positioning; Optimizing the vehicle behavior by the detection result through the Bayesian decision theory to obtain a vehicle decision.
5. A method for fusing visual and auditory information for vehicle autonomous driving according to claim 4, characterized in that, The optimizing the vehicle behavior by the detection result through the Bayesian decision theory includes: U(a|O(t)) = w1·R(a,O(t)) - w2·C(a); where a * represents the action that can maximize the expected utility among all actions; U is the utility function, A is the set of actions; R(a, O(t)) represents the reward for choosing action a in state O(t); C(a) represents the cost of choosing action a; w1 and w2 are weight coefficients.
6. The vehicle automatic driving visual and auditory information fusion method according to claim 1, wherein Before projecting the auditory information into the coordinate system of the visual information, it further includes: Performing time alignment on the visual information and the auditory information.
7. An audiovisual information fusion device for vehicle autonomous driving, comprising a memory, a processor, and a computer program stored on the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements each step in a method for fusing visual and auditory information for vehicle autonomous driving according to any one of claims 1-6.
Citation Information
Patent Citations
Automatic driving vector map online construction method based on crowd source visual image
CN116753936A
Visual and auditory concentration judgment method suitable for driver of vehicle, host and driver monitoring system
CN119489840A