Automatic driving vehicle control method and system based on multi-modal data fusion
Through the multimodal data fusion network MDF-Net, combining three-view cameras, long-range cameras and lidar data, extracting and fusion features, the shortcomings of the autonomous driving system in perceiving complex driving scenarios and predicting vehicle trajectory are solved, and more efficient and accurate autonomous driving control is achieved.
Patent Information
- Application Number
- CN202510184852.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
AI Technical Summary
When existing autonomous driving systems perceive complex driving scenarios, especially the perception of traffic light status, they are prone to misjudgment or misjudgment, and it is difficult to effectively extract and integrate dynamic interaction characteristics of traffic participants, resulting in low accuracy of trajectory prediction.
The multimodal data fusion network MDF-Net is used to input the three-view camera image, telecamera image and lidar image into the network. Through the signal light perception module, scene perception module and trajectory prediction module, the features are extracted and fused, and the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control command is generated.
It improves the accuracy of perception and scenario understanding of complex driving scenarios, significantly improves the accuracy of perception of traffic light status and trajectory prediction, and provides more reliable autonomous driving control decisions.
Smart Images

Figure CN120024353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to an autonomous driving vehicle control method and system based on multimodal data fusion. Background Art
[0002] In the field of autonomous driving, the autonomous driving system needs to make automatic driving control decisions based on the perception of driving traffic scenes. However, driving scenes are complex, highly dynamic, and have dynamic interactions among traffic participants. This makes it difficult to perceive and understand driving traffic scenes for the purpose of making autonomous driving control decisions. Although the existing technology proposes to obtain multimodal data to fully perceive complex driving scenes, the perception of traffic light status is poor and prone to misjudgment or omission; and when dealing with the complex dynamic interaction process of traffic participants in traffic scenes, it is difficult to effectively extract and fuse relevant features, resulting in low trajectory prediction accuracy and failure to provide a reliable basis for autonomous driving control decisions. Summary of the invention
[0003] In order to solve the above problems, the present invention proposes an autonomous driving vehicle control method and system based on multimodal data fusion. By inputting three kinds of perception data, namely three-view camera images, long-range camera images and lidar images, into the multimodal data fusion network MDF-Net for trajectory prediction, the perception of traffic light status is enhanced, and the perception and understanding ability of the autonomous driving system is improved.
[0004] In order to achieve the above object, the present invention adopts the following technical solution:
[0005] In a first aspect, the present invention provides an automatic driving vehicle control method based on multimodal data fusion, comprising:
[0006] Acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image;
[0007] The multimodal image set is input into a data fusion network to obtain a vehicle driving control instruction; the data fusion network includes a signal light perception module, a scene perception module and a trajectory prediction module; wherein, the three-view camera image and the long-range camera image are input into the signal light perception module to extract the signal light perception feature; the three-view camera image and the lidar image are input into the scene perception module to extract the scene perception feature; the signal light perception feature is fused with the scene perception feature to obtain a multi-model fusion feature, the multi-model fusion feature is input into the trajectory prediction module, the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instruction is obtained based on the displacement vector.
[0008] Preferably, the three-view camera images are used to obtain information about vehicles, pedestrians and lane lines; the long-range camera images are used to obtain information about traffic lights; and the lidar images are used to obtain information about other participants in the traffic scene.
[0009] Preferably, the step of inputting the three-view camera image and the long-range camera image into the traffic light perception module to extract the signal perception features specifically includes:
[0010] The traffic light perception module includes a grouped convolutional residual unit and an attention fusion unit;
[0011] The three-view camera images and the long-range camera images acquired at the same time are respectively input into the grouped convolution residual unit to extract the three-view initial features and the long-range initial features; the K, Q, and V values of the two initial features are extracted based on the linear weight layer and input into the attention fusion unit to obtain the signal perception features.
[0012] Preferably, the attention fusion unit includes 6 parallel attention layers.
[0013] Preferably, the step of inputting the three-view camera image and the laser radar image into a scene perception module to extract scene features specifically includes:
[0014] The scene perception module includes a grouped convolutional residual unit, a cascaded attention unit and an average pooling layer;
[0015] The three-view camera image and lidar image acquired at the same time are respectively input into the grouped convolution residual unit to extract the three-view initial features and the lidar initial features; the two initial features are input into the cascade attention unit, and the three-view fusion features and the lidar fusion features are obtained after three-level attention fusion, and the scene features are obtained after the average pooling layer.
[0016] Preferably, the multi-model fusion features are input into the trajectory prediction module, the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instruction is obtained based on the displacement vector, which specifically includes:
[0017] The trajectory prediction module is built based on the GRU network;
[0018] The signal perception features and scene features are used as the initial hidden state vector h of the trajectory prediction module 0 The input is the current position of the vehicle and the target position v 0 ; Based on the initial hidden state h 0 and the initial vehicle position and target position v 0 The output features of the trajectory prediction module will be used as the hidden state vector h at the next moment. 1The output feature is then reduced to 2D through the fully connected layer, and the 2D data represents the displacement vector w of the vehicle's position coordinates at the next moment 0 .
[0019] Preferably, it also includes controlling vehicle driving based on trajectory prediction results; specifically including: calculating the size and direction of the displacement vector based on the predicted displacement data, using the size of the displacement vector as input data for longitudinal control of the vehicle, and using the direction of the displacement vector as input data for lateral control of the vehicle; generating vehicle throttle, direction or brake light instructions based on the displacement vector.
[0020] In a second aspect, the present invention provides an automatic driving vehicle control system based on multimodal data fusion, comprising:
[0021] A data acquisition module, used to acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image;
[0022] A trajectory prediction module is used to input the multimodal image set into a data fusion network to obtain a vehicle driving control instruction; the data fusion network includes a signal light perception module, a scene perception module and a trajectory prediction module; wherein the three-view camera image and the long-range camera image are input into the signal light perception module to extract the signal light perception feature; the three-view camera image and the lidar image are input into the scene perception module to extract the scene perception feature; the signal light perception feature is fused with the scene perception feature to obtain a multi-model fusion feature, the multi-model fusion feature is input into the trajectory prediction module, the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instruction is obtained based on the displacement vector.
[0023] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for controlling an autonomous driving vehicle based on multimodal data fusion described in the first aspect.
[0024] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for controlling an autonomous driving vehicle based on multimodal data fusion as described in the first aspect are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] (1) Aiming at the difficulties of complex, highly dynamic and dynamic interaction of traffic participants in traffic scenes, the present invention designs a multimodal data fusion network based on multimodal perception data obtained by three different perception devices (three-view camera, long-range camera, laser radar) to understand driving scenes and output vehicle control commands. Multimodal data fusion improves the perception ability of complex driving scenes, and improves the accuracy of scene understanding based on the designed multimodal data fusion network.
[0027] (2) The multimodal data fusion network designed by the present invention includes an attention-based three-view camera and a long-range camera fusion module, which improves the network's ability to perceive the state of traffic lights. The multimodal feature fusion module designed based on temporal attention of images and radar data can better extract and fuse the features of the complex dynamic interaction process of traffic participants in traffic scenes, thereby improving the accuracy of trajectory prediction.
[0028] (3) The present invention trains and tests the network based on the CARLA autonomous driving simulator and conducts tests on the public dataset Town05 benchmark. The experimental results show that the designed multimodal data fusion network MDF-Net can make full use of the multimodal data obtained by the three perception devices, thereby improving the autonomous driving system's ability to perceive and understand driving traffic scenes.
[0029] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0031] Figure 1 A main flow chart of an autonomous driving vehicle control method based on multimodal data fusion provided by an embodiment of the present invention;
[0032] Figure 2 This is a data fusion network structure diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] Embodiment 1
[0035] like Figure 1 As shown, this embodiment discloses an automatic driving vehicle control method based on multimodal data fusion, comprising the following steps:
[0036] S1: Acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image;
[0037] S2: Input the multimodal image set into a data fusion network to obtain a vehicle driving control instruction.
[0038] Next, combine Figure 2 , a method for controlling an autonomous driving vehicle based on multimodal data fusion disclosed in this embodiment is described in detail.
[0039] First, the data acquisition device is deployed. A three-view RGB camera, a long-range camera, and a laser radar are installed on the vehicle body. Among these three devices, the three-view camera is used to collect visual information of large-scale driving traffic scenes; the long-range camera is used to detect distant traffic signals in driving tasks as early as possible to facilitate control decisions by the autonomous driving system; and the laser radar is used to collect deep point cloud data in the scene to improve the accuracy of scene perception.
[0040] Considering safety, actual road testing of autonomous driving systems is difficult to implement. Therefore, this embodiment draws on the solutions for most autonomous driving system R&D and testing to collect relevant data sets and test models in an autonomous driving simulator. Based on the model testing requirements proposed in this embodiment, the CARLA autonomous driving simulator is used. Based on the scenes provided by the simulator, a part is selected as the training set to produce scenes, and the other part is selected as the test set to produce scenes. The data collection process for the three different sensors is as follows:
[0041] (1) Three-view RGB images: In the simulator, the shooting angles of the three cameras are set at intervals of 60° from -60° to 60°, which can cover the front and left and right side scenes of the vehicle. The coverage range of each camera angle is set to 65°. The camera height is 2.2 meters and 1.6 meters away from the center of the vehicle. In order to reduce the amount of model calculation, only three types of targets in the driving scene, namely vehicles, pedestrians, and lane lines, are retained for annotation as the true value of the three-view image part in the auxiliary training branch of partial target segmentation.
[0042] (2) Long-range camera image: In the simulator, the long-range camera is set in the front-facing position, and the camera angle coverage is set to 50°. The camera height and position with the vehicle are consistent with the three-view camera. Since the long-range camera is mainly responsible for paying attention to the status of road signals, only the three different colors of traffic lights (red, yellow, and green) in the driving scenes it captures are annotated as the true value of the long-range image part in the auxiliary training branch of partial object segmentation.
[0043] (3) LiDAR point cloud: In the simulator, a LiDAR with a rotation frequency of 10 frames per second is used to acquire simulated point cloud data. The measurement range is set to 90 meters, the radar position is set at the center of the vehicle, and the height is set to 2.8 meters.
[0044] Based on the data perceived by the above three devices, a deep learning network model MDF-Net based on multimodal data fusion is designed. The network is divided into traffic light perception module, scene perception module and trajectory prediction module.
[0045] (1) Traffic light sensing module
[0046] The traffic light perception module is an attention-based multi-image fusion module. Its input data is the three-view camera images and the long-range camera data. The two are fused based on the attention of the traffic signals in the long-range camera to improve the perception accuracy of traffic lights.
[0047] Specifically, the pictures taken by the three-view cameras at the same time are spliced according to the shooting angles to form a three-view camera image with a wide-angle traffic scene vision.
[0048] The telephoto camera image and the three-view camera image obtained by the telephoto camera are respectively input into the group convolution residual network ResNeXt for feature extraction to obtain the multi-view RGB feature F d and the long-range feature F w .
[0049] The feature map F d and F w After a layer of linear weight layer, three values K vector, Q vector and V vector are obtained, specifically {K d ,V d ,Q d} and {K w ,V w ,Q w}.
[0050] K d 、V d and Q d is the feature map F d Key, value and query features obtained through linear transformation; K w 、V w and Q w They are feature maps F w The key, value and query features are obtained through linear transformation; the two pairs of features are subjected to attention calculation of the dot product model. The attention calculation process goes through 6 layers of attention. The results of the 6 parallel attention calculations are connected to obtain the traffic light perception feature F I, so as to fuse the image features of wide-angle traffic scenes and the image features of long-range traffic scenes.
[0051] The traffic light perception module is used to increase the perception weight of distant traffic lights and make a decision on whether to stop the car.
[0052] (2) Scene Perception Module
[0053] The scene perception module is a multimodal feature fusion module of images and radar data based on cascade attention. The input data is three-view camera images and lidar point cloud data. The two perform feature extraction based on convolution operations and multimodal feature fusion process of temporal attention to perceive the overall traffic scene.
[0054] First, the original radar point cloud data is usually messy and high-dimensional. Directly inputting the data into the network may lead to excessive computation and difficulty for the model to learn effective features. Therefore, in order to reduce the computational complexity, the radar data is preprocessed by the following steps:
[0055] 1) Select a point cloud within a certain range in front of the vehicle. In this embodiment, the point cloud within a range of 32 meters is preferably selected. Points are taken at certain intervals to form a matrix point cloud data of size 256×256.
[0056] 2) It is divided into two intervals in height, and forms tensor data to be input into the network with the horizontal plane data.
[0057] 3) After the tensor data is further processed by convolution, pooling and other operations to extract features, it is input into the residual network ResNet 34 for feature extraction to obtain the radar data feature F l .
[0058] Among them, dividing the height into two intervals can reduce the dimension and complexity of the data, making the model easier to process and learn. This avoids the unnecessary computational burden and model learning difficulty caused by the height information being too detailed and complex.
[0059] In traffic scenes, different height intervals in front of the vehicle have different target distributions and characteristics. In this embodiment, the low-height interval is defined as the ground and objects close to the ground, and the high-height interval is defined as vehicles, pedestrians, buildings, etc. This division can more specifically extract key features at different height levels, which helps the model better perceive and understand traffic scenes. It should be understood that the interval division can be set by those skilled in the art according to actual conditions.
[0060] Combining the data of the height interval and the horizontal plane, the multi-dimensional information of the radar point cloud in space can be integrated to form a more comprehensive representation. In this way, the model can consider information in both the horizontal and vertical directions at the same time, so as to more accurately perceive and understand the location, shape and other characteristics of the target in the traffic scene, and improve the accuracy and reliability of scene perception.
[0061] The radar data feature F l and multi-view RGB features F d After three levels of cascaded attention fusion operation, in each level of attention fusion process, the radar data feature F l and multi-view RGB features F d The convolution and pooling dimensionality reduction operations are performed separately, and then the self-attention module is input for fusion. The output preliminary fusion features are divided into two feature models F l2 and F w2 , as the input data for the next level of attention fusion operation.
[0062] After the two modal data are fused through three levels of attention, two output features F are obtained: l3 and F w3 , and then pass through the average pooling layer and then the concatenation layer to obtain a final one-dimensional vector, the scene perception feature F L .
[0063] As a specific implementation method, in order to improve the network's ability to extract and pay attention to the features of traffic lights and traffic participants in driving scenes during the training phase, auxiliary training based on traffic participant and traffic light target segmentation is adopted during the training phase of the traffic light perception module and the scene perception module.
[0064] Specifically, for the three-view camera image, after feature extraction through the residual network ResNeXt, F d After that, a training branch is derived, in which the feature F d First, it passes through two convolutional layers, and then passes through the average pooling layer, the maximum pooling layer, the convolution layer, and the sigmoid layer in turn to obtain the attention weight map, which is fused with the previous convolution layer and then input into the deconvolution layer to obtain the segmentation result map S of the three-view camera image. d .
[0065] For the long-range camera image, after feature extraction through the residual network ResNeXt, F w After that, a training branch is derived, and the subsequent process of this branch is the same as F d The operation of the feature map is the same, and the segmentation result map S of the long-range camera image is obtained. w .
[0066] Segmentation results of three-view camera images d The segmentation result map of vehicles, pedestrians, and lane line targets in the three-view image is used for auxiliary training; the segmentation result map S of the long-range camera image w The segmentation results of the three types of lane line targets in the long-range view are used for auxiliary training to enhance the F d and F w feature representation capability.
[0067] (3) Trajectory prediction module
[0068] The trajectory prediction module is a trajectory prediction and control module based on multimodal fusion features. The input data is the signal light perception feature F output by the signal light perception module. I , and the scene perception feature F output by the scene perception module L , based on GRU to predict trajectory points. Among them, the traffic light perception feature F I The input of is mainly responsible for introducing the state of traffic lights to guide the turning, stopping and starting of the vehicle trajectory; the scene perception feature F L The input is mainly responsible for trajectory prediction based on other participants in the traffic scene; finally, the vehicle driving control output is based on the traffic light status and the predicted trajectory.
[0069] First, the signal light perception feature F I and scene perception features F L Splice to get the multi-model fusion feature F R .
[0070] Then, based on the multi-model fusion feature F R The GRU model is used for trajectory prediction. Specifically:
[0071] The multi-model fusion feature F R As the initial hidden state vector h of the GRU model 0 , and input the vehicle position at the current moment (initial moment) and the target position v 0 ; Based on the initial hidden state h 0 and the initial vehicle position and target position v 0 , the obtained GRU output feature will be used as the hidden state vector h at the next moment 1 At the same time, the output feature is further reduced to 2D through the fully connected layer. The 2D data represents the displacement w of the vehicle's position coordinates at the next moment. 0 ;
[0072] Based on the same process, continue to predict the displacement w of the position coordinates at subsequent times tFinally, based on the predicted displacement data, the magnitude and direction of the displacement vector are calculated, and its magnitude is used as the input data for the longitudinal control of the vehicle, and its direction is used as the input data for the lateral control of the vehicle. The subsequent vehicle controller can generate vehicle throttle, direction or brake light instructions based on the displacement vector, thereby achieving vehicle control.
[0073] In order to verify the effectiveness of this embodiment, the following embodiments are given.
[0074] Based on the above simulation data acquisition method, different weather parameters (sunny, rainy, noon, evening) are set in different scenes to obtain the input data of three modes and the corresponding output data true value, and then the corresponding driving route data is set. The collected data is divided into training set and test set. The loss function of the network training stage includes trajectory prediction loss function and auxiliary training segmentation loss. The vehicle trajectory prediction uses L1 loss function; the auxiliary training segmentation result uses cross entropy loss function.
[0075] After testing the trained model, we use the following two indicators (route completion rate and driving score) to evaluate the test results, referring to the commonly used test evaluation indicators of autonomous driving systems:
[0076] (1) Route completion rate (RC): refers to the percentage of the total length of the test road section that the vehicle completes under the control of the automated driving system based on the designed model.
[0077] (2) Driving Score (DS): refers to the weighted average of the penalty coefficients automatically derived based on the route completion rate and the errors in the driving process. The penalty coefficients are automatically obtained using the built-in statistical method in the simulator, and mainly record the number of times the vehicle runs a red light and collides with other road users.
[0078] The test results of the MDF-Net model designed in this embodiment on two short routes and a long route are as follows: the average driving score DS on the short route is 73.40, and the average route completion rate RC is 87.82. The average driving score DS on the long route is 48.51, and the average route completion rate RC is 88.12. According to the two evaluation index data of DS and RC, it can be seen that the MDF-Net model has achieved good results in the test of simulated driving scenarios.
[0079] In the MDF-Net network provided by the present invention, the traffic light perception module fuses the three-view camera image and the long-range camera image with the traffic signal as the focus of attention through the attention-based multi-image fusion, which significantly improves the perception accuracy of the traffic light, effectively solves the problem of traffic light recognition in automatic driving, and provides a key basis for vehicle driving decision-making; the scene perception module first pre-processes the radar point cloud data according to the characteristics of the radar point cloud data to reduce the calculation complexity, and then uses the multi-modal feature fusion based on cascade attention to make the image and radar data complement each other and fully perceive the traffic scene, and divides the height interval to process the point cloud data, accurately extracts the key features of different height levels, and further improves the accuracy of scene perception; the trajectory prediction module integrates the signal light perception and scene perception features, introduces the traffic light status and the information of other participants in the traffic scene, and performs trajectory prediction based on the GRU model, so that the vehicle driving control is more in line with the actual traffic conditions, and realizes safe and intelligent automatic driving decisions. The overall network design overcomes the technical difficulties of automatic driving from multiple dimensions and enhances the stability and reliability of automatic driving control.
[0080] Embodiment 2
[0081] This embodiment provides an automatic driving vehicle control system based on multimodal data fusion, including:
[0082] A data acquisition module, used to acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image;
[0083] A trajectory prediction module is used to input the multimodal image set into a data fusion network to obtain a vehicle driving control instruction; the data fusion network includes a signal light perception module, a scene perception module and a trajectory prediction module; wherein the three-view camera image and the long-range camera image are input into the signal light perception module to extract the signal light perception feature; the three-view camera image and the lidar image are input into the scene perception module to extract the scene perception feature; the signal light perception feature is fused with the scene perception feature to obtain a multi-model fusion feature, the multi-model fusion feature is input into the trajectory prediction module, the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instruction is obtained based on the displacement vector.
[0084] Embodiment 3
[0085] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the program implements the steps in the method for controlling an autonomous driving vehicle based on multimodal data fusion as described in the first embodiment above.
[0086] Embodiment 4
[0087] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for controlling an autonomous driving vehicle based on multimodal data fusion as described in the first embodiment above are implemented.
[0088] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for controlling an autonomous driving vehicle based on multimodal data fusion, characterized in that: include: Acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image; The multimodal image set is input into a data fusion network to obtain a vehicle driving control instruction; the data fusion network includes a signal light perception module, a scene perception module and a trajectory prediction module; wherein the three-view camera image and the long-range camera image are input into the signal light perception module to extract the signal light perception feature; the three-view camera image and the laser radar image are input into the scene perception module to extract the scene perception feature; The traffic light perception features are fused with the scene perception features to obtain multi-model fusion features, which are then input into the trajectory prediction module. The vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instructions are obtained based on the displacement vector.
2. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 1, characterized in that: The three-view camera images are used to obtain information about vehicles, pedestrians and lane lines; the long-range camera images are used to obtain information about traffic lights; and the lidar images are used to obtain information about other participants in the traffic scene.
3. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 1, characterized in that: The three-view camera images and the long-range camera images are input into the traffic light perception module to extract signal perception features; Specifically include: The traffic light perception module includes a grouped convolutional residual unit and an attention fusion unit; The three-view camera images and the long-range camera images acquired at the same time are respectively input into the grouped convolution residual unit to extract the three-view initial features and the long-range initial features; the K, Q, and V values of the two initial features are extracted based on the linear weight layer and input into the attention fusion unit to obtain the signal perception features.
4. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 3, characterized in that: The attention fusion unit includes 6 parallel attention layers.
5. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 1, characterized in that: The inputting of the three-view camera image and the laser radar image into the scene perception module to extract scene features specifically includes: The scene perception module includes a grouped convolutional residual unit, a cascaded attention unit and an average pooling layer; The three-view camera image and lidar image acquired at the same time are respectively input into the grouped convolution residual unit to extract the three-view initial features and the lidar initial features; the two initial features are input into the cascade attention unit, and the three-view fusion features and the lidar fusion features are obtained after three-level attention fusion, and the scene features are obtained after the average pooling layer.
6. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 1, characterized in that: The multi-model fusion features are input into the trajectory prediction module, the vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instruction is obtained based on the displacement vector, specifically including: The trajectory prediction module is built based on the GRU network; The signal perception features and scene features are used as the initial hidden state vector h0 of the trajectory prediction module, and the input is the current position of the vehicle and the target position v0; the output features of the trajectory prediction module obtained based on the initial hidden state h0 and the vehicle position and target position v0 at the initial moment will be used as the hidden state vector h1 at the next moment; the output features are then reduced to 2 dimensions through a fully connected layer, and the 2D data represents the displacement vector w0 of the vehicle's position coordinates at the next moment.
7. The method for controlling an autonomous driving vehicle based on multimodal data fusion according to claim 6, characterized in that: It also includes controlling vehicle driving based on trajectory prediction results; specifically including: calculating the size and direction of the displacement vector based on the predicted displacement data, using the size of the displacement vector as input data for longitudinal control of the vehicle, and using the direction of the displacement vector as input data for lateral control of the vehicle; generating vehicle throttle, direction or brake light instructions based on the displacement vector.
8. An autonomous driving vehicle control system based on multimodal data fusion, characterized in that: include: A data acquisition module, used to acquire a multimodal image set in a vehicle driving environment; the multimodal image set includes a three-view camera image, a long-range camera image, and a laser radar image; A trajectory prediction module is used to input the multimodal image set into a data fusion network to obtain a vehicle driving control instruction; the data fusion network includes a signal light perception module, a scene perception module and a trajectory prediction module; wherein the three-view camera image and the long-range camera image are input into the signal light perception module to extract the signal light perception feature; the three-view camera image and the laser radar image are input into the scene perception module to extract the scene perception feature; The traffic light perception features are fused with the scene perception features to obtain multi-model fusion features, which are then input into the trajectory prediction module. The vehicle displacement vector is predicted based on the GRU model, and the vehicle driving control instructions are obtained based on the displacement vector.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the steps in the method for controlling an autonomous driving vehicle based on multimodal data fusion as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the automatic driving vehicle control method based on multimodal data fusion as described in any one of claims 1-7 are implemented.
Citation Information
Cited By
Train auxiliary driving method and device
CN121671685A