Intelligent driving environment perception method and system based on multimodal data fusion

By adopting a multimodal data fusion intelligent driving environment perception method in autonomous driving vehicles, combined with deep convolutional neural networks and regional suggestions networks, the problem of sensor signal interference in complex environments is solved, and more accurate environmental perception and safe driving path planning is achieved.

CN119665998BActive Publication Date: 2025-06-06JINCHENG COLLEGE NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510186519.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In complex environments, such as in dense urban high-rise areas or tunnels, sensor signals may be disturbed or distorted, making it difficult for autonomous vehicles to accurately identify road boundaries or navigation signs, increasing the possibility of decision-making errors and safety risks.

Method used

Using an intelligent driving environment perception method based on multimodal data fusion, through the comprehensive use of GPS, LiDAR, camera and IMU sensors, combined with deep convolutional neural network and regional suggestion network, data preprocessing, object detection and environment perception are carried out to achieve fusion and optimization of multimodal data to obtain more comprehensive environmental perception results.

Benefits of technology

In complex environments, the accuracy of autonomous driving vehicles' identification of the surrounding environment is improved, the risks of misjudgment and wrong path selection are reduced, and the safety and reliability of vehicles in complex environments are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119665998B_ABST
    Figure CN119665998B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent driving environment perception method and system based on multimodal data fusion, which relates to the field of autonomous driving technology. The method is as follows: first, the vehicle's location information is obtained through GPS. If the GPS signal is weak, the auxiliary positioning mechanism is enabled; multimodal data acquisition and preprocessing: after the auxiliary positioning mechanism is started, the vehicle collects surrounding environmental information through RGB cameras, laser radars, and IMU sensors; an improved deep convolutional neural network is used to extract features from RGB images, and a high-precision 3D bounding box is generated through a region proposal network; and multimodal information is fused through an efficient image reasoning framework. The present invention is based on multimodal data fusion and environmental perception results. The present invention enables not only accurate target recognition of the surrounding environment during autonomous driving, but also global and local path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and in particular to an intelligent driving environment perception method and system based on multimodal data fusion. Background Art

[0002] In recent years, environmental perception is a crucial link in the application of autonomous driving technology. This function captures the environmental information around the vehicle in real time by deploying a variety of sensor devices (RGB cameras, lidar, GPS locators and IMU sensors). Subsequently, artificial intelligence algorithms are used to deeply process and analyze the captured data to achieve the recognition of moving and stationary targets in the surrounding environment. This includes pedestrians, neighboring vehicles, road facilities, and possible obstacles. In this way, autonomous vehicles are able to build a comprehensive and detailed understanding of the environment and generate an accurate environmental model to provide reliable support for driving decisions.

[0003] Although multi-sensor fusion provides redundancy and complementary advantages in perception, the sensor signals may be interfered with or distorted under certain complex environmental conditions, such as reflective areas of urban high-rise buildings, tunnels, underground garages or other obstructions. This interference may cause the system to receive erroneous data, thereby affecting the algorithm's judgment of the existence of objects and the accurate identification of their positions and movement trajectories, increasing the risk of decision-making errors. In addition, due to sensor misjudgment, the vehicle may not be able to accurately plan its driving path, especially in tunnels or areas with dense urban high-rise buildings. Radar reflections may cause distortion of the environmental model, making it difficult for the vehicle to accurately identify road boundaries or navigation signs, which may cause the vehicle to deviate from the intended route. These problems not only increase the risk of traffic accidents involving autonomous vehicles, but also pose a potential threat to the safety of passengers. Summary of the invention

[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] In view of the problems existing in the intelligent driving environment perception method in the above-mentioned prior art, the present invention is proposed.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: an intelligent driving environment perception method based on multimodal data fusion, the method is:

[0007] Vehicle positioning and signal detection: First, the vehicle's location information is obtained through GPS, and the GPS signal strength is detected. If the GPS signal is weak, the auxiliary positioning mechanism is enabled, that is, the IMU sensor and ground features are used to supplement the positioning information;

[0008] Multimodal data collection and preprocessing: After the auxiliary positioning mechanism is activated, the vehicle collects surrounding environment information through RGB cameras, LiDAR, and IMU sensors, and converts the collected LiDAR point cloud data into pseudo images after preprocessing the data of different sensors;

[0009] Target detection and environmental perception: Use an improved deep convolutional neural network to extract features from RGB images, and combine it with pseudo-image processing to obtain key features. At the same time, the region proposal network GA-RPN generates a high-precision 3D bounding box to determine the precise position and shape of the target in the environment;

[0010] Multimodal data fusion and perception optimization: Through the construction of an efficient image inference framework AF-Infer, multimodal information of RGB images, LiDAR point cloud data, and IMU sensor data is fused to obtain more comprehensive environmental perception results;

[0011] Path planning and motion control: Perform global and local path planning based on environmental perception results and target detection information, and dynamically adjust the planning strategy based on different positioning signal accuracies.

[0012] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, the method also includes: real-time feedback and optimization, continuously collecting data from sensors, performing real-time feedback and adjustment. If the environment changes, the system will readjust the target detection results and path planning according to the current multimodal data.

[0013] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, the method of converting the LiDAR point cloud data into a pseudo image is:

[0014] LiDAR point cloud data consists of a series of points, each of which contains three-dimensional coordinates (x, y, z);

[0015] First, the depth d of each point to the camera plane is calculated, and the point is mapped to the two-dimensional image plane according to the depth;

[0016] The calculation formula is: , ;

[0017] Where u and v are the coordinates of the point on the pseudo image, and f is the focal length of the camera.

[0018] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, the method for extracting the key features is:

[0019] The improved deep convolutional neural network YOLOv5s_Swin Transformer is used to extract features from RGB images; the feature extraction process is expressed as: ;

[0020] in, is the input RGB image, F is the extracted feature map;

[0021] Splice the pseudo image and the feature map of the RGB image to form a new feature map F';

[0022] The process is expressed as: ,in, is the feature map extracted from the pseudo image;

[0023] The output of GA-RPN is expressed as: ,in is the predicted 3D bounding box; then through the point cloud Figure 3 The dimensional box maps the predicted 3D bounding box back to the point cloud space to generate the final 3D bounding box. The mapping process can be expressed as: ,in, is the 3D bounding box in the point cloud.

[0024] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, wherein: the improved deep convolutional neural network YOLOv5s_Swin Transformer is a target detection framework constructed by integrating YOLOv5s and Swin Transformer, including: an image block embedding module and a Swin Transformer module;

[0025] The detection framework takes the RGB image of the road scene captured by the camera of the autonomous vehicle as input, and divides the input image into small blocks through the image block embedding module. Each small block is of size p. Each image block passes through an embedding layer to adjust the image block to an embedding vector of fixed size.

[0026] The formula is: ;

[0027] in, is the input RGB image, is the output feature map, kernel_size is the convolution kernel size, stride is the step size of the convolution kernel moving on the RGB image, and out_channel is the number of output channels;

[0028] The image features are downsampled in four stages through the Swin Transformer module to gradually reduce the resolution of the feature map and expand the receptive field;

[0029] In the first stage, the input image is processed through preliminary convolution and embedding operations to obtain the first feature map, and local window attention is applied. Each stage from the second to the fourth stage uses windows of different sizes to process the feature map and gradually reduces the image resolution to enhance the model's ability to process high-resolution images.

[0030] In each stage, the Swin Transformer module uses the window self-attention mechanism to divide the feature map into non-overlapping windows and performs self-attention calculations independently within each window to calculate the correlation between features based on the query Q, key K, and value V matrices, expressed as:

[0031] ;

[0032] Where d is the dimension of the key, B is the relative position offset, and the symbol T represents the transpose operation of the matrix K; thereby selecting the most relevant features for subsequent processing;

[0033] The Swin Transformer module adopts a window shift strategy to periodically move the window position at the beginning of each stage to capture a wider range of contextual information;

[0034] At the end of each stage, the features are further processed by the multi-layer perceptron MLP to enhance the model's ability to recognize environmental details. The processing formula is:

[0035] ;

[0036] in, represents the output after feature processing of input x, , represents the weight matrix, , It indicates the bias term.

[0037] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, the core of the region proposal network GA-RPN generating a high-precision 3D bounding box is to use the semantic information of the image to guide the generation of anchor points, specifically including: anchor point position prediction and anchor point shape prediction;

[0038] The anchor point position prediction branch Generate a probability map to indicate where the center of the object may be. The size of the probability map is similar to the input feature map. Same, every element Indicated in coordinates , The probability that an object center exists on , where s is the step size of the feature map, and the generation of the probability map is expressed by the following formula: in, is the sigmoid function;

[0039] The anchor point shape prediction branch Predict the width w and height h of the best anchor shape at each position. The shape prediction formula is as follows: , Among them, dw and dh are The output of the network, is an empirical scaling factor, s is the step size;

[0040] Since the anchor shape generated by GA-RPN is variable at different locations, the feature adaptation module is used to adjust the feature map to match the anchor shape, and a 3x3 deformable convolution layer is used to The offset is determined by the output of the shape prediction branch, and the formula for feature adaptation is: in, is the feature of the ith position, is the corresponding anchor point shape.

[0041] As a preferred solution of the intelligent driving environment perception method based on multimodal data fusion described in the present invention, the efficient image inference framework AF-Infer includes three core processing stages: multi-camera data priority sorting and short-term memory enhancement, image preprocessing and thread dispatching, and multimodal adaptive fusion inference;

[0042] The first core stage is: multi-camera data priority sorting and short-term memory enhancement. The data of different cameras are dynamically sorted through a priority queue based on time tags. The priority sorting algorithm focuses on the key tasks of the main camera, supplemented by the auxiliary tasks of the secondary camera, to ensure that key visual targets can be processed first. A target memory module based on time series is introduced to record temporarily obscured target information, such as the short-term trajectory of pedestrians or obstacles.

[0043] The second core stage is: image preprocessing and thread dispatching, which performs unified correction processing on the input image to eliminate the interference of camera distortion and complex lighting conditions. The corrected image is dispatched to the CPU thread pool and GPU CUDA stream through the scheduling module for parallel reasoning;

[0044] The third core stage is: multimodal adaptive fusion reasoning, which integrates IMU inertial navigation data, LiDAR point cloud data and image reasoning results, and completes the deep fusion of multimodal data through an adaptive fusion network. The network achieves the optimal fusion of multimodal information through multi-layer feature extraction and weight adaptive adjustment.

[0045] The present invention also discloses a driving system applied to the above-mentioned intelligent driving environment perception method based on multimodal data fusion, the system comprising:

[0046] Sensing subsystem: integrates multiple sensors to achieve accurate perception of the vehicle's surroundings;

[0047] The path planning subsystem is responsible for calculating a safe and efficient driving path based on the vehicle's real-time location, surrounding environment information, and predetermined driving goals;

[0048] The positioning subsystem measures the speed, distance, acceleration, braking distance, position, heading angle, and sideslip angle of moving objects with high precision by fusing multi-sensor data from various sensor subsystems.

[0049] The motion subsystem achieves precise control of the autonomous vehicle by comprehensively controlling the vehicle's acceleration, braking, steering and suspension functions;

[0050] The exception handling module uses multi-sensor fusion technology to achieve redundant design, ensuring that even if a sensor fails, other sensors can still provide necessary data support.

[0051] The present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the above-mentioned intelligent driving environment perception method based on multimodal data fusion are implemented.

[0052] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned intelligent driving environment perception method based on multimodal data fusion are implemented.

[0053] Beneficial effects of the present invention:

[0054] The present invention combines the output information of multiple sensors such as GPS, LiDAR, camera, IMU sensor, etc., to maximize the advantages of each sensor, while avoiding the limitations of a single sensor in an interference environment. In particular, when the signal is interfered or distorted, other sensors can make up for the lack or deviation of information and reduce misjudgment.

[0055] The present invention processes sensor data (including RGB images and LiDAR point clouds) through an improved deep convolutional neural network YOLOv5s_Swin Transformer to extract accurate target features. Especially in complex environments, the system can automatically optimize target detection results based on the deep learning model and identify key targets such as vehicles, pedestrians, obstacles, and road signs.

[0056] The present invention generates an accurate three-dimensional target bounding box through GA-RPN, which can accurately obtain the spatial position of the target and avoid the problem of environmental model distortion caused by radar reflection.

[0057] In summary, based on the results of multimodal data fusion and environmental perception, the present invention enables not only accurate target recognition of the surrounding environment during autonomous driving, but also global and local path planning. Through accurate target positioning and environmental modeling, the vehicle can plan the best driving route in a complex environment to avoid misjudgment and wrong path selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0059] Figure 1 This is a schematic diagram of the principle of the intelligent driving environment perception method based on multimodal data fusion proposed in the present invention;

[0060] Figure 2 This is a network structure diagram for fusing point cloud data and RGB image data in the intelligent driving environment perception method based on multimodal data fusion proposed in the present invention;

[0061] Figure 3 It is a schematic diagram of the overall structure of the YOLOv5s_Swin Transformer target detection framework in the intelligent driving environment perception method based on multimodal data fusion proposed in the present invention;

[0062] Figure 4This is a schematic diagram of the structure of the GA-RPN network in the intelligent driving environment perception method based on multimodal data fusion proposed in the present invention;

[0063] Figure 5 This is a schematic diagram of the overall structure of the efficient image reasoning framework in the intelligent driving environment perception method based on multimodal data fusion proposed in the present invention. DETAILED DESCRIPTION

[0064] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0066] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0067] Reference Figure 1-Figure 5 , is an embodiment of the present invention, and provides an intelligent driving environment perception method based on multimodal data fusion, the method comprising the following steps:

[0068] Step 1: Vehicle positioning and signal detection: First, obtain the vehicle's location information through GPS and detect the GPS signal strength. If the GPS signal is weak, enable the auxiliary positioning mechanism, that is, use the IMU sensor and ground features to supplement the positioning information;

[0069] Step 2: Multimodal data collection and preprocessing: After the auxiliary positioning mechanism is activated, the vehicle collects surrounding environment information through RGB cameras, LiDAR, and IMU sensors, and after preprocessing the data of different sensors, the collected LiDAR point cloud data is converted into a pseudo image;

[0070] Specifically, LiDAR point cloud data consists of a series of points, each of which contains three-dimensional coordinates (x, y, z);

[0071] First, calculate the depth d of each point to the camera plane. The depth d is the vertical distance from the three-dimensional coordinates (x, y, z) to the camera plane. When the optical axis of the camera coincides with the z-axis, the depth d is the z-coordinate value of the point. Map the point to the two-dimensional image plane according to the depth;

[0072] The calculation formula is: , ;

[0073] Where u and v are the coordinates of the point on the pseudo image, and f is the focal length of the camera.

[0074] Step 3: Target detection and environment perception: Use the improved deep convolutional neural network (YOLOv5s_SwinTransformer) to extract features from RGB images, and combine and process pseudo images to obtain key features;

[0075] Specifically, the improved deep convolutional neural network YOLOv5s_Swin Transformer is used to extract features from RGB images;

[0076] The feature extraction process is expressed as: ;

[0077] in, is the input RGB image, F is the extracted feature map;

[0078] Splice the pseudo image and the feature map of the RGB image to form a new feature map F';

[0079] The process is expressed as: ,in, is the feature map extracted from the fake image.

[0080] Further, the improved deep convolutional neural network YOLOv5s_Swin Transformer is an object detection framework built by integrating YOLOv5s and Swin Transformer, including: image block embedding module and Swin Transformer module;

[0081] The detection framework takes the RGB image of the road scene captured by the camera of the autonomous vehicle as input, and divides the input image into small blocks through the image block embedding module. Each small block is of size p. Each image block passes through an embedding layer to adjust the image block to an embedding vector of fixed size.

[0082] The formula is: ;

[0083] in, is the input RGB image, is the output feature map, kernel_size indicates the convolution kernel size is (p, p), stride is the step size of the convolution kernel moving on the RGB image, and out_channel indicates the number of output channels is C;

[0084] The image features are downsampled in four stages through the Swin Transformer module to gradually reduce the resolution of the feature map and expand the receptive field;

[0085] In the first stage, the input image is processed through preliminary convolution and embedding operations to obtain the first feature map, and local window attention is applied. Each stage from the second to the fourth stage uses windows of different sizes to process the feature map and gradually reduces the image resolution to enhance the model's ability to process high-resolution images.

[0086] In each stage, the Swin Transformer module uses the window self-attention mechanism to divide the feature map into non-overlapping windows and performs self-attention calculations independently within each window to calculate the correlation between features based on the query Q, key K, and value V matrices, expressed as:

[0087] ;

[0088] Where d is the dimension of the key, B is the relative position offset, and the symbol T represents the transpose operation of the matrix K; thereby selecting the most relevant features for subsequent processing;

[0089] The Swin Transformer module adopts a window shift strategy to periodically move the window position at the beginning of each stage to capture a wider range of contextual information;

[0090] At the end of each stage, the features are further processed by the multi-layer perceptron MLP to enhance the model's ability to recognize environmental details. The processing formula is:

[0091] ;

[0092] in, represents the output after feature processing of input x, , represents the weight matrix, , It indicates the bias term.

[0093] At the same time, the region proposal network GA-RPN generates a high-precision 3D bounding box to determine the precise position and shape of the target in the environment. Specifically, the output of GA-RPN is expressed as: ,in is the predicted 3D bounding box; then through the point cloud Figure 3 The 3D Bounding Box in Point Cloud maps the predicted 3D bounding box back to the point cloud space to generate the final 3D bounding box. The mapping process can be expressed as: ,in, is the 3D bounding box in the point cloud.

[0094] The design of Swin Transformer enables it to demonstrate high performance in intelligent driving tasks. By reducing the resolution of feature maps in stages, the model can efficiently process multi-scale features while maintaining computational efficiency. Combined with the target detection framework of YOLOv5s, the network can quickly and accurately detect key elements such as vehicles, pedestrians, and road signs, providing powerful environmental perception capabilities for autonomous vehicles. This network architecture that combines Swin Transformer and YOLOv5s not only improves the accuracy of target detection, but also enhances the robustness of the model in complex environments, providing solid technical support for the safe driving of autonomous vehicles.

[0095] These mathematical expressions summarize the core computational processes of the Swin Transformer module and the YOLOv5s framework in the autonomous driving RGB image feature extraction network, providing accurate and efficient technical support for the environmental perception of autonomous driving vehicles.

[0096] Furthermore, the core of the region proposal network GA-RPN to generate high-precision 3D bounding boxes is to use the semantic information of the image to guide the generation of anchor points, including: anchor point position prediction and anchor point shape prediction;

[0097] The anchor point position prediction branch Generate a probability map to indicate where the center of the object may be. The size of the probability map is similar to the input feature map. Same, every element Indicated in coordinates , The probability that an object center exists on , where i and j are usually used to represent the index position in the two-dimensional space, s is the step size of the feature map, and the generation of the probability map is expressed by the following formula: in, is the sigmoid function;

[0098] The anchor point shape prediction branch Predict the width w and height h of the best anchor shape at each position. The shape prediction formula is as follows: , Among them, dw and dh are The output of the network, is an empirical scaling factor, s is the step size

[0099] Since the anchor shape generated by GA-RPN is variable at different locations, the feature adaptation module is used to adjust the feature map to match the anchor shape, and a 3x3 deformable convolution layer is used to The offset is determined by the output of the shape prediction branch, and the formula for feature adaptation is: in, is the feature of the ith position, is the corresponding anchor point shape.

[0100] In the autonomous driving scenario, GA-RPN has obvious advantages. Through more effective anchor point generation and feature adaptation mechanisms, the detection capability of targets of different scales and shapes can be improved. In particular, for small targets in the distance, occluded targets, and targets in complex backgrounds, GA-RPN can provide more accurate regional suggestions. This helps to improve the perception accuracy of the autonomous driving system for various objects on the road (such as vehicles, pedestrians, traffic signs, etc.), thereby providing a more reliable basis for subsequent decision-making and planning, and enhancing the safety and reliability of the autonomous driving system. At the same time, since the number of unnecessary anchor points can be reduced (such as reducing the number of anchor points by 90% to maintain a high recall rate), the computing cost is reduced and the operating efficiency of the system is improved, making it more suitable for running on resource-constrained on-board computing devices, meeting the strict real-time requirements of the autonomous driving system.

[0101] Step 4: Multimodal data fusion and perception optimization: Through the constructed efficient image inference framework AF-Infer, the multimodal information of RGB images, LiDAR point cloud data, and IMU sensor data is fused to obtain more comprehensive environmental perception results.

[0102] To further explain, the efficient image inference framework AF-Infer includes three core processing stages: multi-camera data prioritization and short-term memory enhancement, image preprocessing and thread dispatching, and multi-modal adaptive fusion inference;

[0103] The first core stage is: multi-camera data priority sorting and short-term memory enhancement. The data of different cameras are dynamically sorted through a priority queue based on time tags. The priority sorting algorithm focuses on the key tasks of the main camera, supplemented by the auxiliary tasks of the secondary camera, to ensure that key visual targets can be processed first. A target memory module based on time series is introduced to record temporarily obscured target information, such as the short-term trajectory of pedestrians or obstacles.

[0104] The second core stage is: image preprocessing and thread dispatching, which performs unified correction processing on the input image to eliminate the interference of camera distortion and complex lighting conditions. The corrected image is dispatched to the CPU thread pool and GPU CUDA stream through the scheduling module for parallel reasoning;

[0105] The third core stage is: multimodal adaptive fusion reasoning, which integrates IMU inertial navigation data, LiDAR point cloud data and image reasoning results, and completes the deep fusion of multimodal data through an adaptive fusion network. The network achieves the optimal fusion of multimodal information through multi-layer feature extraction and weight adaptive adjustment.

[0106] The core design goals of AF-Infer include: priority task processing, short-term memory enhancement, and efficient multimodal data fusion. Through the combination of time tag management and LSTM modules, the system can maintain the consistency of key target perception in a short period of time. In addition, multi-threaded parallelism and GPU optimization ensure that the system can meet real-time requirements, and the inference delay is always lower than the safety threshold of the autonomous driving system (33 milliseconds). In terms of multimodal fusion, the Adaptive FusionNet realizes end-to-end joint inference of IMU, LiDAR and image data to ensure the spatiotemporal consistency of multi-source information.

[0107] In order to achieve the best balance between inference performance and system resources, AF-Infer introduces targeted optimization strategies in multiple processing stages. In the image preprocessing stage, a deep learning-based image correction algorithm is used to significantly reduce the impact of distortion on inference accuracy; in the multi-thread scheduling stage, a dynamic thread management strategy is used to achieve efficient use of CPU and GPU resources; in the multi-modal fusion stage, the system uses an adaptive weight learning optimization module to dynamically adjust the fusion weights of different modal inputs, thereby enhancing the robustness and adaptability of the fusion results.

[0108] AF-Infer combines priority sorting, short-term memory enhancement, and multimodal adaptive fusion to provide an efficient image inference solution for complex scenarios of autonomous driving. The framework has significant advantages in real-time performance, target perception consistency, and multimodal information fusion accuracy.

[0109] Step 5: Path planning and motion control: Perform global path planning and local path planning based on the environmental perception results and target detection information, and dynamically adjust the planning strategy according to the different positioning signal accuracies.

[0110] Step 6: Real-time feedback and optimization. Continuously collect data from sensors, provide real-time feedback and make adjustments. If the environment changes, the system will readjust the target detection results and path planning based on the current multimodal data.

[0111] This embodiment also provides a driving system applied to the above-mentioned intelligent driving environment perception method based on multimodal data fusion, and the system includes: a sensing subsystem, a path planning subsystem, a positioning subsystem, a motion subsystem and an exception handling module.

[0112] Specifically:

[0113] Sensing subsystem: The sensing subsystem is a core component of autonomous driving technology. It integrates multiple sensors to achieve accurate perception of the vehicle's surrounding environment. The system mainly includes RGB cameras, lidar, IMU sensors and Adaptive Fusion Net, each of which performs its own duties and jointly provides necessary data support for autonomous driving vehicles. Among them, Adaptive Fusion Net, in the field of autonomous driving, navigation systems usually rely on the fusion of multiple sensor data, including inertial measurement units (IMUs), global positioning systems (GPS), video cameras and lidar point clouds. However, most traditional methods use fixed-weight fusion strategies, which cannot effectively cope with the challenges brought about by environmental changes. To this end, AdaptiveFusion Net achieves adaptive fusion of multi-source data through a deep convolutional neural network (CNN) architecture. When signal conditions are good, the system prioritizes GPS data for navigation and uses IMU data for auxiliary correction. This combination takes advantage of the global positioning advantages of GPS and the local motion detection capabilities of IMU. When the environment is complex or the signal conditions deteriorate, the GPS signal weakens. At this time, the system automatically reduces the weights of GPS and IMU and relies on video cameras and lidar point cloud data instead. This strategy uses high-definition map data downloaded to the local area for positioning, combined with the environmental perception capabilities of video and lidar to achieve accurate navigation decisions.

[0114] The RGB camera is responsible for capturing color images around the vehicle and providing visual information to the system for identifying obstacles such as road signs, traffic lights, pedestrians, and vehicles. The LiDAR measures the distance between objects and the vehicle by emitting laser beams and receiving reflected signals, generating a high-precision three-dimensional environmental map. The IMU sensor is responsible for providing the vehicle's acceleration and angular velocity data to enable real-time monitoring of the vehicle's dynamic state.

[0115] When the GPS positioning signal is strong, the data of the sensor subsystem will be directly used for navigation planning and vehicle control to ensure that the vehicle drives stably along the predetermined path. When the GPS positioning signal is weak, the sensor subsystem will play a more important role, using the data collected by its sensors to assist positioning, and combining it with map information to perform path planning and update local paths to ensure that the vehicle can drive safely in complex environments.

[0116] In addition, the sensor subsystem also has a redundant design. Even if a sensor fails, other sensors can continue to provide necessary data to ensure the reliability and safety of the system. This design enables the autonomous driving system to operate stably in various environments, effectively improving the safety and reliability of autonomous driving.

[0117] Path planning subsystem: The path planning subsystem is a key module in autonomous driving technology. It is responsible for calculating a safe and efficient driving path based on the vehicle's real-time position, surrounding environment information, and predetermined driving goals. The system uses advanced algorithms and data processing technologies to ensure that the optimal path planning solution can be provided under various traffic conditions and environmental conditions.

[0118] The path planning subsystem first receives real-time data from the sensing subsystem, including information provided by RGB cameras, lidar, and IMU sensors. This data is used to build or update the environment model around the vehicle to provide accurate map information for path planning. In cases where GPS signals are weak or unavailable, the path planning subsystem will rely on these sensor data to obtain detailed information about the surrounding environment, including the location, size, and shape of obstacles, as well as the dynamics of other vehicles and pedestrians. Using this information, the path planning subsystem evaluates different path options and selects the best path.

[0119] The path planning subsystem also includes a feedback mechanism that can dynamically adjust the planned path according to the actual driving conditions of the vehicle and the real-time changes in the environment. This means that even if there are unexpected situations during driving, such as road construction, traffic accidents or weather changes, the system can quickly re-plan the path to ensure that the vehicle can reach its destination safely and smoothly.

[0120] Positioning subsystem: It integrates the RGB camera, IMU sensor and lidar data of multiple sensor subsystems to achieve high-precision measurement of the speed, distance, acceleration, braking distance, position, heading angle, sideslip angle and other parameter information of moving objects.

[0121] Motion subsystem: The motion subsystem is a key technology for achieving precise control of autonomous vehicles. The system integrates multiple sensors, controllers and actuators to comprehensively control the vehicle's acceleration, braking, steering and suspension functions to improve the vehicle's handling, safety and comfort. With the development of autonomous driving technology, the vehicle motion control system maintains vehicle body stability in different driving modes and various road conditions to ensure that the vehicle can operate safely and stably.

[0122] Abnormal processing module, the system adopts multi-sensor fusion technology to achieve redundant design, ensuring that even if a sensor fails, other sensors can still provide necessary data support.

[0123] This embodiment also provides a computer device, which is suitable for the case of an intelligent driving environment perception method based on multimodal data fusion, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the intelligent driving environment perception method based on multimodal data fusion as proposed in the above embodiment.

[0124] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0125] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the intelligent driving environment perception method based on multimodal data fusion proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0126] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An intelligent driving environment perception method based on multimodal data fusion, characterized in that: The method is: Vehicle positioning and signal detection: First, the vehicle's location information is obtained through GPS, and the GPS signal strength is detected. If the GPS signal is weak, the auxiliary positioning mechanism is enabled, that is, the IMU sensor and ground features are used to supplement the positioning information; Multimodal data collection and preprocessing: After the auxiliary positioning mechanism is activated, the vehicle collects surrounding environment information through RGB cameras, LiDAR, and IMU sensors, and converts the collected LiDAR point cloud data into pseudo images after preprocessing the data of different sensors; Target detection and environmental perception: Use an improved deep convolutional neural network to extract features from RGB images, and combine it with pseudo-image processing to obtain key features. At the same time, the region proposal network GA-RPN generates a high-precision 3D bounding box to determine the precise position and shape of the target in the environment; Multimodal data fusion and perception optimization: Through the construction of an efficient image inference framework, multimodal information of RGB images, LiDAR point cloud data, and IMU sensor data is fused to obtain more comprehensive environmental perception results; Path planning and motion control: Perform global and local path planning based on environmental perception results and target detection information, and dynamically adjust planning strategies based on different positioning signal accuracies; Wherein: the method for extracting the key features is: The improved deep convolutional neural network is used to extract features from RGB images; the feature extraction process is expressed as: F = YOLOv5s_Swin Transformer (I RGB ); Among them, I RGB is the input RGB image, F is the extracted feature map; Splice the pseudo image and the feature map of the RGB image to form a new feature map F'; The process is expressed as: F' = Concat (F pesudo ,F'), where,F pesudo is the feature map extracted from the pseudo image; The output of GA-RPN is expressed as: B 3D =GA-RPN(F fusion ), where B 3D is the predicted 3D bounding box; then the predicted 3D bounding box is mapped back to the point cloud space through the point cloud 3D box to generate the final 3D bounding box. The mapping process can be expressed as: point_cloud =Map_to_Point_Cloud(B 3D ), where B point_cloud is the 3D bounding box in the point cloud.

2. The intelligent driving environment perception method based on multimodal data fusion according to claim 1, characterized in that: The method also includes: real-time feedback and optimization, continuously collecting data from sensors, performing real-time feedback and adjustments. If the environment changes, the system will readjust the target detection results and path planning based on the current multimodal data.

3. The intelligent driving environment perception method based on multimodal data fusion according to claim 1, characterized in that: The method of converting the LiDAR point cloud data into a pseudo image is: LiDAR point cloud data consists of a series of points, each of which contains three-dimensional coordinates (x, y, z); First, calculate the depth d from each point to the camera plane. The depth d is the vertical distance from the three-dimensional coordinates (x, y, z) to the camera plane. When the optical axis of the camera coincides with the z-axis, the depth d is the z-coordinate value of the point. Map the point to the two-dimensional image plane according to the depth. The calculation formula is: Where, u and v are the coordinates of the point on the pseudo image, and f is the focal length of the camera.

4. The intelligent driving environment perception method based on multimodal data fusion according to claim 3 is characterized by: The improved deep convolutional neural network is a target detection framework constructed by integrating YOLOv5s and Swin Transformer, including: an image block embedding module and a Swin Transformer module; The detection framework takes the RGB image of the road scene captured by the camera of the autonomous vehicle as input, and divides the input image into small blocks through the image block embedding module. Each small block is of size p. Each image block passes through an embedding layer to adjust the image block to an embedding vector of fixed size. The formula is: X patch =Conv2d(X img ,kernel_size=(p,p), stride=(p,p), out_channel=C); Among them, X img is the input RGB image, X patch is the output feature map, kernel_size is the convolution kernel size, stride is the step size of the convolution kernel moving on the RGB image, and out_channel is the number of output channels; The image features are downsampled in four stages through the SwinTransformer module to gradually reduce the resolution of the feature map and expand the receptive field; In the first stage, the input image is processed through preliminary convolution and embedding operations to obtain the first feature map, and local window attention is applied. Each stage from the second to the fourth stage uses windows of different sizes to process the feature map and gradually reduces the image resolution to enhance the model's ability to process high-resolution images. In each stage, the SwinTransformer module uses the window self-attention mechanism to divide the feature map into non-overlapping windows and performs self-attention calculations independently within each window to calculate the correlation between features based on the query Q, key K, and value V matrices, expressed as: Where d is the dimension of the key, B is the relative position offset, and the symbol T represents the transpose operation of the matrix K; thereby selecting the most relevant features for subsequent processing; The Swin Transformer module adopts a window shift strategy to periodically move the window position at the beginning of each stage to capture a wider range of contextual information; At the end of each stage, the features are further processed by the multi-layer perceptron MLP to enhance the model's ability to recognize environmental details. The processing formula is: MLP(x)=GELU(xW1+b1)W2+b2 Among them, MLP(x) represents the output after feature processing of input x, W1 and W2 represent weight matrices, and b1 and b2 represent bias terms.

5. The intelligent driving environment perception method based on multimodal data fusion according to claim 1, characterized in that: The core of the region proposal network GA-RPN to generate high-precision 3D bounding boxes is to use the semantic information of the image to guide the generation of anchor points, including: anchor point position prediction and anchor point shape prediction; The anchor point position prediction, anchor point position prediction branch N L Generate a probability map to indicate the possible location of the object center. The size of the probability map is the same as the input feature map F'. Each element p(i,j|F') represents the probability of the object center existing at the coordinates (i+1 / 2)s, (j+1 / 2)s, where i and j are used to indicate the index position in the two-dimensional space, and s is the step size of the feature map. The generation of the probability map is expressed by the following formula: p(i,j|F') = σ(N L (F')) where σ is the sigmoid function; The anchor point shape prediction branch N S Predict the width w and height h of the best anchor point shape at each position. The shape prediction formula is as follows: w = σ·s·e dw ,h=γ·s·e dh Where dw and dh are N S The output of the network,γ,is an empirical scaling factor; Since the anchor shape generated by GA-RPN is variable at different positions, the feature adaptation module is used to adjust the feature map to match the anchor shape, and a 3x3 deformable convolution layer N is used. T The offset is determined by the output of the shape prediction branch, and the feature adaptation formula is: i =N T (f i ,w i ,h i ) where f i is the feature of the i-th position, (w i ,h i ) is the corresponding anchor point shape.

6. The intelligent driving environment perception method based on multimodal data fusion according to claim 5 is characterized in that: The efficient image inference framework includes three core processing stages: multi-camera data prioritization and short-term memory enhancement, image preprocessing and thread dispatching, and multi-modal adaptive fusion inference; The first core stage is: multi-camera data priority sorting and short-term memory enhancement. The data of different cameras are dynamically sorted through a priority queue based on time tags. The priority sorting algorithm focuses on the key tasks of the main camera, supplemented by the auxiliary tasks of the secondary camera, to ensure that key visual targets can be processed first. A target memory module based on time series is introduced to record the information of temporarily obscured targets: the short-term trajectory of pedestrians or obstacles. The second core stage is: image preprocessing and thread dispatching, which performs unified correction processing on the input image to eliminate the interference of camera distortion and complex lighting conditions. The corrected image is dispatched to the CPU thread pool and GPU CUDA stream through the scheduling module for parallel reasoning; The third core stage is: multimodal adaptive fusion reasoning, which integrates IMU inertial navigation data, LiDAR point cloud data and image reasoning results, and completes the deep fusion of multimodal data through an adaptive fusion network. The network achieves the optimal fusion of multimodal information through multi-layer feature extraction and weight adaptive adjustment.

7. The driving system of the intelligent driving environment perception method based on multimodal data fusion according to claim 6 is characterized in that: The system includes: Sensing subsystem: integrates multiple sensors to achieve accurate perception of the vehicle's surroundings; The path planning subsystem is responsible for calculating a safe and efficient driving path based on the vehicle's real-time location, surrounding environment information, and predetermined driving goals; The positioning subsystem measures the speed, distance, acceleration, braking distance, position, heading angle, and sideslip angle of moving objects with high precision by fusing multi-sensor data from various sensor subsystems. The motion subsystem enables precise control of the autonomous vehicle by comprehensively controlling the vehicle's acceleration, braking, steering and suspension functions; The exception handling module uses multi-sensor fusion technology to achieve redundant design, ensuring that even if a sensor fails, other sensors can still provide data support.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent driving environment perception method based on multimodal data fusion described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent driving environment perception method based on multimodal data fusion described in any one of claims 1 to 7 are implemented.