Integrated fusion multi-sensor automatic driving intelligent perception device
Through the multi-sensor fusion system and intelligent perception model, the problem of insufficient sensor perception capabilities of autonomous vehicles in complex urban environments has been solved, and high-precision traffic element identification and path planning are achieved around the clock, reducing the risk of accidents.
Patent Information
- Application Number
- CN202110551961.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing autonomous vehicles have difficulty accurately identifying traffic elements and obstacles in complex urban environments, especially in extreme weather or harsh conditions where the sensor perception capability decreases, resulting in a high accident rate.
A multi-sensor system integrating vehicle-mounted lidar, millimeter-wave radar, binocular camera and infrared camera, combined with Beidou short message emergency communication and high-precision positioning interface, realizes multi-source data synchronization and intelligent perception model, and performs real-time 3D target tracking and path planning.
It improves the perception accuracy and safety of vehicles in urban traffic environments, can identify and predict the behavior of traffic participants in real time, reduce the risk of accidents, and has all-weather high-precision perception capabilities.
Smart Images

Figure CN113313154B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of automatic driving intelligent perception, and further relates to an integrated portable automatic driving intelligent perception device based on multi-sensor fusion. BACKGROUND
[0002] Unmanned driving is a product of the high development of artificial intelligence, computer science and automation technology. At present, artificial intelligence is divided into three stages: computational intelligence, perceptual intelligence and cognitive intelligence. However, there is a big gap between machines and humans in terms of cognitive abilities such as understanding, thinking and reasoning. An automatic driving car is a product of the integration of technologies such as automotive electronics, intelligent control and the Internet. Its principle is that an automatic driving system uses a perception system to obtain information about the vehicle itself and the external environment, analyzes the information through a computing system, makes decisions, and controls the execution system to realize vehicle acceleration, deceleration or turning, so as to complete automatic driving without the intervention of a driver. An automatic driving car uses cameras, laser radars, millimeter wave radars and other perception devices and artificial intelligence algorithms to realize perception of complex road environments. Due to the lack of cognitive ability of the current intelligent driving car, the existing intelligent car often has difficulty in making correct understanding when facing decision information, which leads to the fact that the car cannot quickly adapt to the driving environment and perceive the surrounding environmental information in time and quickly, like a human driver. Intelligent driving sensors and computing decision platforms are the basis for realizing the intelligentization of cars. Sensors can replace the eyes of humans to perceive the external environment. Advanced vehicle-mounted visual sensors, radars and other perception devices can obtain and perceive the distance, speed, direction of surrounding vehicles / pedestrians, traffic signs and other information, and then transmit the perception information to the computing decision platform. This module supports fine-grained, structured semantic perception for complex scenarios, can realize highly scalable, modular three-dimensional semantic environment reconstruction, and transparent, traceable and reasoning decision and path planning.
[0003] Complex environment perception technology is a key component of autonomous driving technology and a key breakthrough and core technology for the application of artificial intelligence vision technology in the autonomous driving field. According to statistics on autonomous vehicle accidents, urban streets are more prone to traffic accidents than highways. This is primarily due to the greater complexity of traffic elements, the greater volume of traffic participants, and the more random and diverse behavior of traffic participants in urban environments. These characteristics pose greater perception challenges for autonomous vehicles when operating in urban environments. Since current intelligent driving tests mostly occur on urban roads or highways, accurate perception and recognition in unusual (rare) scenarios is a key concern. For example, extreme weather (such as heavy snow and fog) can reduce the maximum range and signal quality (sensitivity, contrast, and excessive visual clutter) of human vision, AV vision systems (cameras, lidar), and DSRC transmissions. Excessive dust or physical obstructions (such as snow or ice) on the vehicle can interfere with or reduce the maximum perception range and signal quality (sensitivity, contrast, and physical obstruction of the field of view) of all essential AV sensors (cameras, lidar, and millimeter-wave radar). Darkness or low lighting (such as in tunnels) can reduce the maximum range and signal quality (sensitivity and contrast) of AV camera systems. The problem of reduced sensor perception capabilities caused by these common limitations has always been a difficult problem to overcome in the industry.
[0004] The perception capability of intelligent driving mainly includes internal perception, driver perception and environment perception. Internal perception mainly obtains the vehicle state through CAN bus to collect the information of each electronic control unit in the vehicle and the data information generated by various sensors loaded on the vehicle, including vehicle body (temperature inside and outside the vehicle, air flow, tire pressure), power (oil pressure, speed, oil), vehicle safety (seat belt, airbag, door and window lock), etc. Driver perception mainly refers to the realization of fatigue monitoring, attention monitoring, overspeed monitoring, driving posture monitoring of the driver through the camera, face recognition, prompter and other intelligent kits, so as to realize active safety from the driver driving level and reduce accidents. Environment perception mainly uses sensors, positioning navigation and vehicle-to-vehicle communication (V2X) to realize the identification, perception and prediction of the environment. The mainstream sensor perception technology includes visual perception, laser perception and millimeter wave perception. Visual perception is based on the image information collected by the camera, which uses visual related algorithms for processing to recognize the surrounding environment; laser perception is based on the point cloud data collected by the laser radar, which uses filtering, clustering and other technologies to perceive the environment; millimeter wave perception is based on the distance information collected by the millimeter wave radar, which uses distance related algorithms for processing to recognize the surrounding environment. However, the data collected by a single sensor does not have integrity and intelligence, 3D environment modeling makes laser radar become the core sensor, but it cannot identify images and colors, and its performance is significantly reduced in bad weather. Millimeter wave radar can realize all-weather perception, but its resolution is low and it is difficult to image. The camera is cheap and can identify traffic participants and traffic signs, but it cannot realize point array modeling and long-distance ranging. Therefore, multi-sensor fusion is the only way to realize high-level automatic driving and an important trend for the future development of automatic driving. Traditional object detection models usually only perform subsequent operations on the last feature map of the deep convolutional network, and the corresponding down-sampling rate (image reduction multiple) of this layer is usually large, such as 16, 32, which causes less effective information of small objects on the feature map, and the detection performance of small objects will decrease sharply. Since the anchor of RPN is uniformly distributed, its variance is very large and difficult to learn, which needs to be iteratively regressed. However, RPN does not have means such as RoIPool or RoIAlign for feature alignment, because the input of RPN is very important, and only regular sliding convolution can be performed for output, which causes the symmetry problem of anchor and feature. In order to alleviate the alignment problem, some researches use deformable convolution to perform spatial transformation on the feature map, hoping to make the fine-tuned anchor and the transformed feature aligned. However, this method does not have strict constraints to ensure that the feature and the transformed anchor are aligned, and it is also difficult to determine whether the transformed feature and the anchor are aligned.When determining whether an anchor is positive or negative, simply using the anchor-free or anchor-base method is not enough, because using the anchor-free standard will cause the stage 2 requirements to be too low, while using the anchor-base will cause the stage 1 to fail to regress enough positive samples. Summary of the Invention
[0005] In order to improve the environmental perception level of autonomous driving and reduce the accident rate of smart cars in urban traffic environments, the present invention proposes an integrated fusion multi-sensor autonomous driving intelligent perception device that can recognize the traffic environment in real time, reliably and accurately and respond promptly.
[0006] The above-mentioned purpose of the present invention can be achieved through the following measures. An integrated fusion multi-sensor automatic driving intelligent perception device includes: a sensor data acquisition module, a sensor data synchronization module, a vehicle-mounted laser radar, a millimeter-wave radar, a binocular camera and an infrared camera integrated with the vehicle-mounted laser radar on the automatic driving car through an expandable Beidou short message emergency communication and Beidou high-precision positioning interface, a sensor data synchronization module, a vehicle-mounted computing unit, a sensor data processing module, a planning control module, a sensor fusion execution control module and a power supply module. It is characterized in that: the vehicle-mounted laser radar, the millimeter-wave radar, the binocular camera and the infrared camera are integrated, and artificial intelligence technology is used to simulate the human perception process of the external environment, establish a sensor, positioning, AI perception, path planning and decision-making, and control vehicle intelligent perception model module, the intelligent perception model module performs real-time 3D target tracking and detection of static and dynamic obstacles outside the vehicle, detects the environment in which the vehicle is located, obtains information and behavior information of each element in the scene, sends and receives vehicle positioning information, status information and control information to the sensor data acquisition module in real time, and adopts a multi-sensor fusion method for intelligent driving perception. Knowing the longitude, latitude, altitude, speed, heading angle, pitch angle, roll angle, Lidar point cloud information and high-definition video, it obtains and outputs the vehicle's status information. After completing the data reception of all sensors, it transmits the sensor collected data to the sensor data synchronization module. The sensor data synchronization module automatically conducts comprehensive analysis of information and data from multiple intelligent sensors or multiple sources, and realizes millimeter-level spatial synchronization and nanometer-level time synchronization of multi-source heterogeneous sensors. It also sends the spatiotemporal registration of target-level detection information to the time and space synchronization calibration module, performs time and space synchronization calibration on the sensor data, and sends the synchronized and calibrated data to the sensor data processing module in real time. Data is processed by the on-board computing unit on the on-board computing platform, and the processed data is sent to the planning and control module. The planning and control module completes the planning, control and decision-making of the vehicle's driving path, obstacle avoidance, and perception based on the perception results of the intelligent perception model module, determines the optimal path and decision of the vehicle, makes real-time trajectory predictions, and guides the vehicle to complete path planning. The vehicle's decision-making information is sent to the artificial intelligence algorithm module to determine the weight of each sample, and the new data set with modified weights is sent to the lower-level classifier for training. The classifiers obtained from each training are finally fused together and reach the sensor fusion control vehicle execution module as the final decision classifier. The sensor fusion control vehicle execution module plans the control path according to the planning control module. Under the control of the vehicle decision system, it connects to the vehicle's central control module to control the vehicle in real time, executes perception data, identifies the trafficability, static and dynamic objects within the full field of view around the vehicle body, and makes decisions on braking and obstacle avoidance.
[0007] Compared with existing equipment, the present invention has the following advantages and beneficial effects:
[0008] The present invention uses an expandable Beidou short message emergency communication and Beidou high-precision positioning interface to connect an on-board laser radar, millimeter-wave radar, binocular camera, and a sensor data acquisition module, a sensor data synchronization module, a sensor data processing module, a planning control module, an on-board computing unit, a sensor fusion execution control module, and a power module on an autonomous vehicle. The on-board laser radar, millimeter-wave radar, binocular camera, and infrared camera are efficiently integrated together, with high integration, small size, miniaturization, and flexible configuration. Through multi-sensor fusion, the vehicle can achieve 3D target detection of external traffic participants for system control and path planning; can achieve target tracking of participants to complete flexible control of the vehicle system; can complete rapid monitoring of traffic lights to make braking or communication decisions for the vehicle; can make real-time trajectory predictions of traffic participants around the vehicle body, allowing the system to make pre-judgments and decisions; can identify lane lines, traffic signs, crosswalks, and other information on the road to keep the vehicle from deviating from the lane and avoid them in time; can perform real-time detection of static and dynamic obstacles, and make braking and avoidance decisions for the system. It realizes functions related to highly automated driving, monitors the vehicle's trajectory in real time, and plans the route. The overall route is: collect multi-sensor data and perform synchronous calibration, perceive and detect all obstacles around the vehicle after front-end fusion, process the fused data through artificial intelligence technology, and track obstacles. Combined with high-precision positioning information such as road objects, it predicts the movement trajectory of all obstacles in the future, and finally outputs obstacle perception information and future situation, and guides the vehicle to complete path planning. It solves the two core problems of understanding urban traffic scenes (roadways, sidewalks, traffic signs, buildings, trees, lawns, etc.) throughout the day, detecting traffic participants (vehicles, pedestrians, etc.), and identifying behavioral intentions and trajectory prediction.
[0009] The present invention uses vehicle-mounted laser radar, millimeter-wave radar, binocular camera and infrared camera fusion, uses artificial intelligence technology to simulate the human perception process of the external environment, and establishes intelligent perception model modules for sensors, positioning, AI perception, path planning and decision-making, and vehicle control. The intelligent perception model module performs real-time 3D target tracking and detection of static and dynamic obstacles outside the vehicle, detects the environment in which the vehicle is located, obtains information and behavior information of each element in the scene, and sends and receives vehicle positioning information, status information and control information to the sensor data acquisition module in real time to complete the data of all sensors. It uses multi-sensor fusion to perform intelligent driving perception, and can obtain and output vehicle status information, including sensor longitude, latitude, altitude, speed, heading angle, pitch angle, roll angle, Lidar point cloud information, high-definition video, etc. In terms of sensor spatiotemporal synchronization, it can achieve millimeter-level spatial synchronization and nanometer-level time synchronization of multi-source heterogeneous sensors, with a time accuracy of 10 -6 m 3 The device can sense all static and dynamic traffic elements within a 360-degree range of the vehicle body, with a stable forward obstacle detection range of 250 meters and a minimum detection distance of 0.1 meter. The traffic element perception angle deviation is less than or equal to 0.015°. It can identify all obstacles around the vehicle, analyze the intentions and behaviors of each traffic participant, and predict their trajectories. It has strong robustness and value-added expansion capabilities.
[0010] The present invention transmits sensor-collected data to a sensor data synchronization module, which automatically analyzes and synthesizes information and data from multiple intelligent sensors or multiple sources, achieving millimeter-level spatial synchronization and nanometer-level temporal synchronization for multi-source heterogeneous sensors. It also performs spatiotemporal registration of target-level detection information, calibrates sensor data for both temporal and spatial synchronization, and feeds the synchronized and calibrated data into a sensor data processing module in real time. The sensor data processing module processes the data via an onboard computing unit on an onboard computing platform, effectively reducing the amount of data required for transmission and improving processing efficiency. The onboard computing unit delivers the processed data to a planning and control module, which, based on the perception results of the intelligent perception model module, completes planning, control, and decision-making for the vehicle's driving path, obstacle avoidance, and perception, determining the optimal path and decision-making for the vehicle and making real-time trajectory predictions.
[0011] The sensor fusion control vehicle execution module according to the control path planning of the planning control module, under the control of the vehicle decision system, connects the automobile central control module to perform real-time control on the vehicle, executes the perception data, identifies the passing ability, static and dynamic objects in the global view range around the vehicle body, and makes the decision of braking and obstacle avoidance. In this way, the information and data from multiple sensors or multiple sources are automatically analyzed and synthesized, so that the vehicle can perceive more abundant and accurate information than a single sensor; through global tracking of all targets detected by multiple sensors, the functions of environment perception, obstacle detection, trajectory prediction, early warning information, navigation function and night driving perception can be effectively realized. The device can identify the front sidewalk, traffic roadside signs, traffic signal lights and road terrain conditions on the driving road perception level; it can accurately identify the objects affecting traffic safety in real time, reliably and accurately plan the driving path that can guarantee the standard, safety and rapid arrival of the destination. On this basis, the behavior and trajectory trend of the detected target objects can be predicted, including the distance of the target object from the vehicle, the driving speed, the moving direction and the motion trajectory. The device can obtain the information of each element in the scene and the intention and behavior information of each traffic participant at the same time, and these results can be directly called by the intelligent driving car decision system to provide the intelligent driving car with full-range high-precision perception data, greatly improving the safety of the intelligent driving car in real road driving. In addition, it can be connected to the automobile power supply or independently powered by an external power supply, and can be externally extended in power supply. It is powerful, flexible and configurable, can be externally connected to multiple interfaces, and has strong expansibility. This technical route can be moved to other unmanned autonomous systems, such as unmanned ship systems, various mobile robot systems, etc. By modifying the environment perception content, it can be applied to target perception in other scenes, such as industrial park intelligent delivery, mine intelligent loading and unloading, automatic garbage collection, etc., and can provide high-quality services for special unmanned vehicles, has strong robustness and value-added expansion capability, and enriches the application of intelligent vehicles. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a schematic diagram of the integrated fusion multi-sensor automatic driving intelligent perception device of the application;
[0013] Figure 2 is Figure 1 a schematic diagram of the fusion implementation path. DETAILED DESCRIPTION
[0014] Reference Figures 1-2In the preferred embodiment described below, an integrated fusion multi-sensor automatic driving intelligent perception device comprises: a vehicle-mounted laser radar, a millimeter wave radar, a binocular camera, and a sensor data acquisition module fused with an infrared camera, a sensor data synchronization module, a vehicle-mounted computing unit, a sensor data processing module, a planning control module, a sensor fusion execution control module, and a power module connected through an extensible Beidou short message emergency communication and Beidou high-precision positioning interface on an automatic driving vehicle. The device is characterized in that the vehicle-mounted laser radar, the millimeter wave radar, the binocular camera, and the infrared camera are fused, the artificial intelligence technology is used to simulate the perception process of people to the external environment, the intelligent perception model of the sensors, positioning, AI perception, path planning and decision, and control of the vehicle is established, the intelligent perception model performs real-time 3D target tracking detection on static and dynamic obstacles outside the vehicle, detects the environment in which the vehicle is located, obtains the behavior information of each element information in the scene, and sends, receives vehicle positioning information, state information, and control information to the sensor data acquisition module in real time. The intelligent driving perception longitude, latitude, height, speed, heading angle, pitch angle, roll angle, Lidar point cloud information, and high-definition video are obtained and output, the state information of the vehicle is obtained, after all the sensor data is received, the sensor data is transmitted to the sensor data synchronization module, the sensor data synchronization module automatically analyzes the information and data from multiple intelligent sensors or multiple sources, realizes the spatial synchronization of millimeter level and the time synchronization of nanometer level of the multi-source heterogeneous sensors, and sends the time and space registration of the target level detection information into the time and space synchronization calibration module. The sensor data is time and space synchronized and calibrated, the synchronized and calibrated data is sent into the sensor data processing module in real time, the data is processed by the vehicle-mounted computing unit on the vehicle-mounted computing platform, the processed data is sent to the planning control module, the planning control module completes the planning control and decision of the vehicle driving path, obstacle avoidance, and perception according to the perception results of the intelligent perception model module, determines the optimal path and decision of the vehicle, makes real-time trajectory prediction, guides the vehicle to complete path planning, and sends the decision processing information of the vehicle to the intelligent algorithm module. The weight of each sample is determined, the new data set with modified weight is sent to the lower classifier for training, the classifier obtained by each training is finally fused as the final decision classifier to reach the sensor fusion control vehicle execution module. According to the control path planning of the planning control module, the sensor fusion control vehicle execution module controls the vehicle in real time under the control of the vehicle decision system, executes the perception data, identifies the passing ability, static and dynamic objects in the global field of view around the vehicle body, and makes the decision of braking and obstacle avoidance.
[0015] The intelligent perception model adopts a multi-sensor fusion manner to fuse the collected data of millimeter wave radar, laser radar, binocular camera and infrared camera on the autonomous vehicle. A sensor data synchronization module performs space-time registration on the target level detection information of the intelligent sensor to realize millimeter-level space synchronization and nanometer-level time synchronization of the multi-source heterogeneous sensors. A vehicle-mounted computing unit provides strong computing power to ensure high energy efficiency and high performance of the device.
[0016] The vehicle-mounted computing unit processes multi-source data from the front-end perception to the back-end, specifically plans vehicle actions, checks whether the abstract strategy is executable or the action that satisfies the strategy is executable, converts the learned abstract strategy into actual control actions for the vehicle, and thus fully guarantees the safety of the system.
[0017] The artificial intelligence computing method module adopts an artificial intelligence algorithm built in an artificial intelligence AI chip, makes vehicle decisions and plans based on the artificial intelligence algorithm, completes preliminary calculation, adopts reinforcement learning to make high-level strategies for driving needs, and implements specific path planning and obstacle avoidance according to these strategies and dynamic planning.
[0018] The artificial intelligence computing method module adopts deep learning to perform three-dimensional point cloud target detection of the laser radar. The three-dimensional point cloud target detection detects the point cloud of the entire scene. An abstract strategy is checked for whether it is executable or an action that satisfies the strategy is executable. The structure includes three parts of a feature extraction stage, a backbone network and an RPN. The detailed steps are as follows:
[0019] 1. The feature extraction stage adopts a feature extraction module to divide the point cloud of the entire scene into three-dimensional grids of the same size. The point cloud data containing the coordinate values and reflection intensity information of each point is input. In order to fix the number of point clouds in each three-dimensional grid, if the number of point clouds is too small, it is directly padded to a fixed number with zero. If the number of point clouds is too large, a fixed number of point clouds are randomly selected. The center of gravity of each grid is calculated to obtain the offset of each point cloud and the center of gravity in the grid. The offset is spliced to the feature, and then a plurality of PointNet networks are used to extract high-dimensional features of the point cloud in the three-dimensional grid. A CNN model based on a classification task (such as ImageNet) is used as a feature extractor. The input is a convolution feature picture in the form of length, width and height DxWxH. After processing by the pre-trained CNN model, a convolution feature map (convolution layer) is obtained. The output of the convolution layer is stretched into a one-dimensional vector.
[0020] 2. The feature extraction module uses a series of convolutions and pooling to extract a feature map from the original image in the convolutional layer. The target location is obtained from the feature map through network training. The target to be classified is extracted from the feature map, and the feature map is divided into multiple small regions. The coordinates of the foreground region are obtained, and the data is pooled into a fixed length as the network input. The data is mapped to an area of the original image centered on the current sliding window, and the point cloud is converted into a pseudo-image structure to be sent to the backbone network for processing. The region proposal network is used to extract the pseudo-image of the detected region. R-CNN uses the Selective Search algorithm to extract (propose) possible RoIs (regions of interest) and then classifies each extracted region using a standard CNN. The Selective Search method sets 2,000 candidate regions of different shapes, sizes, and locations around the target object, and then convolves these regions to find the target object.
[0021] 3. The backbone network consists of two parts: the first part is a top-down network structure, which is mainly used to increase the number of channels in the feature map and reduce the resolution of the feature map; the second part processes the multiple feature maps of the first part through multiple upsampling operations, and then splices the results to form a multi-scale feature map structure, ready to be sent to the final stage of the network; integrating the entire object detection process into a single neural network.
[0022] 4. The RPN part uses the RPN structural module, which receives the results processed by the backbone network. It mainly uses multiple convolutional layers for operation, and finally uses three independent convolutions for object category classification. After the feature map is obtained through four downsampling layers, the feature map undergoes two convolutions: one convolution for foreground and background classification, and the other for boungding box regression. It performs object position regression and object orientation estimation, and estimates the probability of each area being the target or background, resulting in a fixed-length vector.
[0023] 5.RPN structure module The fully convolutional network has two built-in convolutional layers. The first convolutional layer encodes all the information of the convolutional feature map, encodes each sliding window position of the feature map into a feature vector, and maintains the position of the "things" encoded relative to the original image; the second convolutional layer processes the extracted convolutional feature map, searches for a predefined number of regions that may contain the target, and outputs k anchors corresponding to each sliding window position as the probability of the object. The center of the anchor point is located at the center of the convolution kernel sliding window, and a binary category label is assigned to each anchor point. The output value of the RPN network W×H×k anchor points is calculated. The size of the output convolutional layer is the total output length of 2×k anchors corresponding to the probabilities of two output objects and k regressed regions.
[0024] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. An integrated multi-sensor fusion autonomous driving intelligent perception device, comprising: Through the expandable Beidou short message emergency communication and Beidou high-precision positioning interface, the vehicle-mounted laser radar, millimeter-wave radar, binocular camera and infrared camera integrated sensor data acquisition module, sensor data synchronization module, vehicle-mounted computing unit, sensor data processing module, planning control module, sensor fusion execution control module and power supply module are connected to the autonomous driving vehicle. The characteristics are: the vehicle-mounted laser radar, millimeter-wave radar, binocular camera and infrared camera are integrated, and artificial intelligence technology is used to simulate the human perception process of the external environment, establish a sensor, positioning, AI perception, path planning and decision-making, and control vehicle intelligent perception model module. The intelligent perception model module performs real-time 3D target tracking and detection of static and dynamic obstacles outside the vehicle, detects the environment in which the vehicle is located, obtains information and behavior information of each element in the scene, and sends and receives vehicle positioning information, status information and control information to the sensor data acquisition module in real time. It uses multi-sensor fusion to perform intelligent driving perception of longitude, latitude, altitude, speed, heading angle, pitch angle, roll angle, Lidar point cloud information and high-definition video, obtains and outputs vehicle status information, and transmits sensor collected data to the transmission after completing data reception of all sensors. The sensor data synchronization module automatically analyzes information and data from multiple intelligent sensors or multiple sources, realizes millimeter-level spatial synchronization and nanometer-level time synchronization of multi-source heterogeneous sensors, and sends the synchronized and calibrated data to the sensor data processing module in real time. The data is processed by the on-board computing unit on the on-board computing platform and the processed data is sent to the planning and control module. The planning and control module completes the planning, control and decision-making of the vehicle's driving path, obstacle avoidance and perception based on the perception results of the intelligent perception model module, determines the optimal path and decision of the vehicle, and makes real-time trajectory prediction. After the vehicle completes path planning, the decision-making information is processed and sent to the artificial intelligence algorithm module to determine the weight of each sample. The new data set with the modified weight is sent to the lower-level classifier for training. The classifiers obtained from each training are finally fused and sent to the sensor fusion execution control module as the final decision classifier. The sensor fusion execution control module plans the control path according to the planning control module. Under the control of the vehicle decision system, it connects to the car's central control module to control the vehicle in real time, executes perception data, identifies the trafficability and static and dynamic objects in the full field of view around the vehicle body, and makes decisions on braking and obstacle avoidance. The artificial intelligence calculation method module uses deep learning to perform three-dimensional point cloud target detection of lidar. The three-dimensional point cloud target detection detects the point cloud of the entire scene, and uses three parts including the feature extraction stage, the backbone network and the RPN structure to check whether the abstract strategy is executable or to perform actions that meet the strategy.
2. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 1, characterized in that: The intelligent perception model module adopts a multi-sensor fusion approach. By integrating the collected data of the millimeter-wave radar, lidar, binocular camera and infrared camera on the autonomous driving vehicle, the sensor data synchronization module performs spatiotemporal alignment on the target-level detection information of the intelligent sensor, realizing millimeter-level spatial synchronization and nanometer-level time synchronization of multi-source heterogeneous sensors. The computing unit then provides powerful computing power to ensure the high energy efficiency and high performance of the device.
3. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 2, characterized in that: The on-board computing unit processes multi-source data from front-end perception to back-end, makes specific plans for vehicle actions, checks whether abstract strategies are executable or executes actions that satisfy the strategies, and converts the learned abstract strategies into actual control actions for the vehicle, thereby fully ensuring the safety of the system.
4. The integrated fusion multi-sensor autonomous driving intelligent perception device according to claim 1, characterized in that The artificial intelligence calculation module uses the artificial intelligence algorithm built into the artificial intelligence AI chip to make vehicle decisions and planning based on the artificial intelligence algorithm, complete the initial calculation rate, and use reinforcement learning to decide the advanced strategies required for driving. Specific path planning and obstacle avoidance are implemented according to these strategies and dynamic planning.
5. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 1, characterized in that: In the feature extraction stage, the feature extraction module is used to divide the point cloud of the entire scene into three-dimensional grids of the same size, and input point cloud data containing the coordinate value and reflection intensity information of each point; in order to fix the number of point clouds in each three-dimensional grid, if the number of point clouds is too small, it is directly padded with zeros to a fixed number; if the number of point clouds is too large, a fixed number of point clouds are directly randomly selected, and then the center of gravity of each grid is calculated, and the offset of each point cloud and the center of gravity in the grid is obtained and spliced to the feature, and then multiple PointNet networks are used to extract high-dimensional features of the point cloud in the three-dimensional grid; the CNN model based on the classification task is used as the feature extractor, and the convolution feature image in the form of length, width and height D×W×H is input. After processing by the pre-trained CNN model, the convolution feature map is obtained, and the output of the convolution layer is stretched into a one-dimensional vector.
6. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 1, characterized in that: The feature extraction module uses a series of convolutions and pooling to extract feature maps from the original image of the convolution layer. The location of the target is obtained from the feature map through network training. The target to be classified is extracted from the feature map, and the feature map is divided into multiple small areas. The coordinates of the foreground area are obtained. The data is pooled into a fixed length as the input of the network. The center of the current sliding window is used as the center to map to an area of the original image. The point cloud is converted into a pseudo-image structure to be sent to the backbone network for processing. The detected region pseudo image is extracted through the region generation network.
7. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 1, characterized in that: The backbone network consists of two parts: the first part is a top-down network structure that is mainly used to increase the number of channels of the feature map and reduce the resolution of the feature map; the second part processes the multiple feature maps of the first part through multiple upsampling operations, and then splices the results to form a multi-scale feature map structure, which is ready to be sent to the last stage of the network; Integrate the entire object detection process into a neural network.
8. The integrated fusion multi-sensor autonomous driving intelligent perception device according to claim 1, characterized in that: The RPN part adopts the RPN structure module, which receives the results processed by the backbone network and uses multiple convolutional layers for operation. Finally, three independent convolutions are used for object category classification. After the feature map is obtained through 4 downsampling layers, the feature map is convolved twice. One convolution is used for foreground and background classification, and the other is used for boungdingbox regression. The object position is regressed and the object orientation is estimated. The probability of each area being the target or background is estimated to obtain a vector of fixed length.
9. The integrated multi-sensor fusion autonomous driving intelligent perception device according to claim 1, characterized in that: The RPN structure module fully convolutional network has two built-in convolutional layers. The first convolutional layer encodes all the information of the convolutional feature map, encodes each sliding window position of the feature map into a feature vector, and maintains the position of the "things" encoded relative to the original image; the second convolutional layer processes the extracted convolutional feature map, searches for a predefined number of regions that may contain the target, and outputs k anchors corresponding to each sliding window position as the probability of the object. The center of the anchors is located at the center of the convolution kernel sliding window, and a binary category label is assigned to each anchor. The output value of the RPN network W×H×k anchors is calculated. The size of the output convolutional layer, the total output length is 2×k anchors corresponding to the probability of two output objects and k regressed regions.
Citation Information
Patent Citations
Automatic driving system based on enhanced learning and multi-sensor fusion
CN108196535A
An unmanned vehicle target detection method based on multimodal depth learning
CN109543601A
3D target detection method for point cloud screening based on image semantic features
CN111145174A