Automatic obstacle avoidance system for vehicles based on real-time 3D target detection
By using LoGoNet network in the autonomous driving system for multimodal data fusion, the problem of degradation in the performance of the autonomous driving obstacle avoidance system in complex scenarios is solved, and higher target detection accuracy and better obstacle avoidance decision-making effect are achieved.
Patent Information
- Application Number
- CN202411835395.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-13
AI Technical Summary
The existing autonomous driving obstacle avoidance system is difficult to effectively deal with and analyze a large number of obstacles and traffic conditions in complex scenarios, resulting in performance degradation, accuracy and reliability issues, limiting the system's adaptability and decision-making effect.
采用基于实时3D目标检测的汽车自动避障系统,通过LoGoNet网络进行局部到全局的多模态数据融合,结合图像和激光雷达数据,实现更完整的三维环境感知能力。
The target detection accuracy is improved, false detection and missed detection situations are reduced, and the car's automatic obstacle avoidance is realized through real-time 3D target detection. The system's adaptability and decision-making effect are significantly improved.
Smart Images

Figure CN119323777B_ABST
Abstract
Claims
1. An automatic obstacle avoidance system for automobiles based on real-time 3D target detection, characterized in that: The system comprises: Point cloud data collection module: including laser radar, used to collect three-dimensional point cloud data of the surrounding environment; Image data collection module: including a depth camera module for capturing two-dimensional images of the surrounding environment; LoGoNet network: used to realize real-time target detection of obstacles; connected with the point cloud data collection module and the image data collection module, it processes the collected point cloud data and image data in real time, detects whether there are obstacles, and determines the category of obstacles; Early warning and feedback module: obtain the detection result of the LoGoNet network, set the first distance threshold, the second distance threshold and the first height threshold, and the second height threshold. If there is an obstacle in the forward area and it is a dynamic target, when the distance of the obstacle detected in front of the vehicle's driving route is not greater than the first distance threshold, start to avoid the obstacle; if there is an obstacle in the forward area and it is a static target, when the height of the obstacle detected in front of the vehicle's driving route is greater than the first height threshold and the distance is not greater than the first distance threshold, re-plan the route to avoid the obstacle; if there is an obstacle in the forward area and it is a static target, when the height of the obstacle detected in front of the vehicle's driving route is between the second height threshold and the first height threshold and the distance is not greater than the second distance threshold, re-plan the route to avoid the obstacle; The first distance threshold is greater than the second distance threshold, the first height threshold is greater than the second height threshold, the first distance threshold is 20m, the second distance threshold is 10m, the first height threshold is 3m, and the second height threshold is 0.3m; The LoGoNet network includes a global fusion module GoF, a local fusion module LoF, and a feature dynamic aggregation module FDA. It takes the point cloud data P obtained by the laser radar and the image data I obtained by the depth camera module as input, including a point cloud branch and an image branch. For the point cloud branch, given the input point cloud, a 3D voxel-based skeleton network is used to generate voxel features. ; Then use the region proposal network RPN to generate the initial bounding box proposal , n is the number of bounding box proposals; For the image branch, the image data I is processed using the 2D pixel skeleton to obtain dense semantic image features. ; Voxel features and dense semantic image features Input into the global fusion module GoF, dense semantic image features and bounding box proposals Input to the local fusion module LoF; obtain the features of the local grid area of interest through the local fusion module LoF processing and local mesh fusion features , obtain the global fusion feature through the global fusion module GoF ; Will , , The input is fed into the feature dynamic aggregation module FDA for cross-modal fusion, and the target detection result is output. The target detection result includes the category to which the obstacle target belongs and the position and size of the obstacle target in three-dimensional space; The processing process of the feature dynamic aggregation module FDA is: , , Aggregate according to the following formula to obtain the aggregated features : = + + , Will After being input into CBAM attention for processing, it is then combined with Perform residual connection, then input it into the residual connection block G-RCB composed of GLU and residual, and output the final classified image with bounding box; The CBAM attention includes a channel attention module and a spatial attention module. The input of the CBAM attention is used as Query, Key, and Value respectively. The three are arranged into a matrix and input into the channel attention module to emphasize the features in different channels. The formula of the channel attention module is expressed as: , in, It refers to the sigmoid function; , are the parameters obtained through MLP training, , , R is the hidden activation size, r is the reduction rate, C is the number of channels; b is the bias parameter; and They are two different spatial context descriptors generated by compressing features in the spatial dimension through average pooling and maximum pooling respectively; is the channel attention feature; The spatial attention module is used to emphasize the spatial region in the feature, and its formula is expressed as: = , in, Indicates filter 7 7 convolution operations, and They are two 2D maps generated by applying average pooling and maximum pooling in the channel dimension respectively; [;] represents the concatenation operation in the channel dimension; is the spatial attention feature; First, the channel attention feature Multiply it with the input of CBAM attention to get the refined feature F', and use F' as the input of the spatial attention module; then use the spatial attention feature Multiply it with F' to get the final adaptively refined features; The structure of the G-RCB is expressed by the formula: , in, is the output of the linear gating unit GLU, is the input of the linear gating unit GLU, W is the trainable weight, is the overall output of G-RCB; The global fusion module GoF is used to obtain global features. The process is as follows: input voxel features and dense semantic image features , by The spatial positions of all points are averaged to obtain the averaged features of each voxel The point centroid : Next, for each point centroid Assign a voxel grid index and match the relevant voxel features through the hash table; use the camera projection matrix M' and the point centroid Calculate the reference point on the image plane : , Among them, M' is used to project points in three-dimensional space onto the two-dimensional image plane; is matrix multiplication; at the same time, through the voxel feature The sampling offset is obtained by linear projection through the linear layer and attention weights ; Dense semantic image features Learn the offset , according to the formula Generate image features ; At the reference point Based on the reference point A set of surrounding image features According to formula (3), clustering is performed to generate clustered image features. , (3) in, and is the learnable weight, is the number of attention heads, K is the total number of sampling points; and They represent the sampling offset and attention weight of the kth sampling point in the mth attention head respectively, O is a trainable parameter; h is a trainable bias vector; Each voxel feature Represented as a query , the image features will be aggregated Represented as a key K and a value V, where After the input is processed by CBAM attention, it is cascaded with K and V, and then with the averaged feature of each voxel Splicing to get fused voxel features ; Finally, Perform ROI pooling to generate global fusion features ; The local fusion module LoF includes a grid point dynamic fusion module GDF, which is used to dynamically fuse point cloud features and image features and perform local multimodal feature fusion on image features; The process of the grid point dynamic fusion module GDF is: given each bounding box proposal , which is divided into a regular voxel grid , where j is the voxel grid index; take the center point As the corresponding voxel grid The grid points are used to propose bounding boxes using the position information encoder PIE Encode the location information of each bounding box proposal Create a mesh feature ; Perform PIE processing on all bounding box proposals to obtain local grid area of interest features = , is the encoded local original point cloud position index; Features in the region of interest on the local grid Linear projection is performed on the sample to obtain the sampling offset and attention weights , dense semantic image features Learn the offset ,according to Generate image features , According to formula (3), clustering is performed to generate clustered image features. ; (3) in, and is the learnable weight, is the number of attention heads, K is the total number of sampling points; O is a trainable parameter; h is a trainable bias vector; The local grid region of interest features Represented as a query , the image features will be aggregated Represented as a key K and a value V, where After the input is processed by CBAM attention, it is cascaded with K and V, and then combined with the local grid area of interest features. Splicing to get dynamic fusion features of grid points ; Then Perform ROI pooling to obtain local grid fusion features , obtain the output of the local fusion module; The three-dimensional point cloud data collected by the point cloud data collection module undergoes the following enhancement processing: (1.1) The point cloud data is collected in the road scene of daily life. The point cloud data in a collection scene is a point cloud set. All the collected point cloud sets constitute a point cloud data set. At the same time, at least 5 RGB images are randomly taken using the depth camera module in each point cloud scene. The RGB images taken in all point cloud scenes constitute an image data set. The point cloud set in a point cloud scene and the corresponding RGB image constitute a sample data, and all sample data constitute a data set. (1.2) Set the detection distance of the X and Y axes to [-75.2m, 75.2m], the detection distance of the Z axis to [-2m, 4m], and the voxel size to (0.1m, 0.1m, 0.15m); XYZ represents the spatial coordinate axis of the point cloud; (1.3) Normalize the collected point cloud data; (1.4) Then put the point cloud data in [- ] to perform random rotation with the Z axis as the rotation axis to simulate the scene of the vehicle in different directions; (1.5) Then globally scale the point cloud data using a scaling factor in the range of [0.95, 1.05] to simulate the observation effects at different distances; (1.6) Finally, random noise is added to the point cloud data to simulate the sensor noise in the real world and obtain the enhanced point cloud dataset.
2. The automatic obstacle avoidance system for automobiles based on real-time 3D target detection according to claim 1, characterized in that: The position information encoder PIE is implemented using a multi-layer perceptron MLP.
3. The automatic obstacle avoidance system for automobiles based on real-time 3D target detection according to claim 1, characterized in that: The training process of the LoGoNet network is: In the first stage, the point cloud branch is trained 20 times. When training the point cloud branch, the weight of the image branch is frozen, and the overall loss function L is composed of the RPN loss L RPN , confidence loss L conf and box regression loss L reg composition, L=L RPN +L conf +ɑL reg Among them, ɑ is a hyperparameter for balancing different losses; In the second stage, the entire LoGoNet network is trained for 60 epochs, and the batch size is set to 8 batches per GPU; In the third stage, the entire LoGoNet network is trained for 80 epochs, and the batch size is set to 2 batches per GPU; When the training gradient value of the training set is within the range of ±1e-5, the optimal training parameters are obtained. Then, the training is completed when the target detection accuracy is not less than 95% on the test set and the validation set.
4. The automatic obstacle avoidance system for automobiles based on real-time 3D target detection according to claim 3 is characterized in that: The value of ɑ is 0.5-1.
5. The automatic obstacle avoidance system for automobiles based on real-time 3D target detection according to claim 1, characterized in that: The depth camera module includes four cameras: front view, rear view, side view, and side rear view blind spot compensation. The front view is composed of three cameras: one 2-megapixel lens with a long-range field of view FOV of 30°, one 2-megapixel lens with a medium-range field of view FOV of 60°, and a 5-megapixel lens with a short-range field of view FOV of 130°; the side view is composed of two 2-megapixel FOV100° lenses, the side rear view blind spot compensation is composed of two 2-megapixel FOV100° lenses, and the rear view is composed of one 2-megapixel FOV100° lens.