Vehicle three-dimensional information processing method based on multi-scale data fusion

By using multi-scale data fusion and joint loss function methods in the vehicle three-dimensional detection algorithm, the existing vehicle three-dimensional detection algorithm has solved the problems of single information and low accuracy, and efficient and accurate vehicle three-dimensional information detection is achieved to meet the needs of autonomous driving.

CN119991923APending Publication Date: 2025-05-13SHAANXI HIGH SPEED ELECTRONIC ENG CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410241798.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing three-dimensional vehicle detection algorithms have problems such as single information, over-reliance on templates, low accuracy and requiring a large number of GPU resources.

Method used

Using a vehicle three-dimensional information processing method based on multi-scale data fusion, a multi-scale weighted fusion feature map is generated through the ResNet-50 backbone feature extraction network and an autonomously designed weighted multi-scale feature fusion module, and a joint loss function is used for training to achieve fast and accurate detection of vehicle three-dimensional information.

Benefits of technology

The accuracy and speed of vehicle three-dimensional detection are improved, and the vehicle three-dimensional information can be quickly solved from the perspective of a single-sided road camera to meet the requirements of accuracy and speed for autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991923A_ABST
    Figure CN119991923A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle three-dimensional information processing method based on multi-scale data fusion, and the method comprises the steps: firstly, extracting the features of an original image through employing ResNet-50, and obtaining a multi-scale image data feature graph; meanwhile, point cloud data in a road scene are obtained, a pseudo image of the point cloud data is generated by using a Point Pill network to serve as a point cloud data feature map, then the point cloud data feature map passes through ResNet-18 to obtain a multi-scale point cloud data feature map, and the multi-scale point cloud data feature map is subjected to a point cloud data feature map; sending the multi-scale image data feature map and the multi-scale point cloud data feature map into a designed continuous projection convolution fusion module, fusing the multi-scale data feature map, and generating a multi-scale weighted fusion feature map; and carrying out regression on the feature map by using 1 * 1 convolution and an attention mechanism to obtain a vehicle information thermodynamic diagram, and decoding to obtain vehicle three-dimensional information. The network designed by the scheme can be used for quickly realizing feature extraction and fusion on the input image, so that the problem of solving the three-dimensional information of the vehicle under the view angle of the monocular roadside camera is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and artificial intelligence, and in particular to a vehicle three-dimensional information processing method based on multi-scale data fusion. Background Art

[0002] In the Intelligent Transportation System (ITS), vehicle 3D detection is an important foundation for autonomous driving and intelligent obstacle avoidance. Vehicle 3D information includes basic vehicle type, vehicle center point offset, vehicle 3D frame vertices and vehicle 3D frame center point, etc. When an autonomous vehicle is driving on a lane, it needs to perceive the surrounding 3D scene, especially the dynamic scene. As the most obvious moving target in the road scene, it is also particularly important to perceive it. As the basis of the perception system, vehicle 3D detection plays an important guiding role in motion prediction, path planning, etc. Therefore, the problem of solving vehicle 3D detection has been widely discussed and studied.

[0003] Traditional vehicle 3D detection algorithms are mainly based on images or laser point clouds. Image-based methods mainly include methods based on geometric attributes and methods based on template matching: In the method based on geometric attributes, target detection is first used to accurately detect the two-dimensional box, and then the close geometric relationship between the perspective projection of the three-dimensional coordinate point and the two-dimensional box is used to solve the three-dimensional information of the vehicle through linear or nonlinear constraints. However, this method can usually only solve the three-dimensional size or position information of the vehicle, but cannot ensure the accurate positioning of the road scene; in the method based on template matching, known CAD templates or other geometric templates are usually used to accurately match the vehicle in the road scene to obtain the three-dimensional information of the vehicle. However, this method is too dependent on the template, and the result accuracy is not high, which cannot meet the accuracy requirements of autonomous driving. Among the methods based on laser point clouds, there are mainly voxel-based methods and point-based methods: voxel-based methods usually use point cloud target detection algorithms to convert irregular point clouds into compact voxel representations, and then extract the point cloud features of the three-dimensional target through a three-dimensional convolutional neural network. However, due to the discretization of point clouds, information loss will occur, thereby reducing the fine-grained positioning accuracy; point-based methods usually use PointNet and its variants or graph neural networks to train point cloud data. However, compared with the target detection algorithm of two-dimensional images, the algorithm difficulty of this method has increased by one dimension, and the required training resources are also very large. Therefore, the current three-dimensional vehicle detection mainly has the following problems: the image-based algorithm can only solve a single piece of information and is too dependent on templates; the laser point cloud-based method has low accuracy and requires large GPU resources. Summary of the invention

[0004] The purpose of the present invention is to provide a vehicle three-dimensional information processing method based on multi-scale data fusion to solve the following main problems in the current vehicle three-dimensional detection mentioned in the above background technology: the image-based algorithm can only solve a single piece of information and is too dependent on templates; the laser point cloud-based method has low accuracy and requires large GPU resources.

[0005] To achieve the above object, the present invention provides the following technical solution: a vehicle three-dimensional information processing method based on multi-scale data fusion, the information processing method comprising the following steps:

[0006] S1. First, data set and preprocessing are required: the data set used is the vehicle three-dimensional detection data set SVLD-3D. The algorithm verification and data annotation process are completed with the assistance of independent annotation software to ensure that the data maintains a certain numerical resolution. It contains video data of 21 scenes from three perspectives: left, middle, and right under seven highway scenes. At the same time, the data set also contains annotation information such as spatial position and speed collected by Lidar and GPS. During the network input process, the pseudo image generated by the original RGB image and point cloud data is scaled to a certain size, and then the image is downsampled to determine the downsampling factor size, and then the processed image is sent to the ResNet50 backbone feature extraction network.

[0007] S2. Then the vehicle 3D detection network is needed: This paper uses ResNet-50 as the backbone feature extraction network because the network contains a residual structure and has good feature extraction capabilities. The feature map obtained by the network is then sent to the self-designed weighted multi-scale feature fusion module, which is to better adapt to multi-scale vehicle information. At the same time, this paper deconvolves the five feature maps output by the feature extraction network and upsamples them to feature sets of the same size, and then performs weighted feature fusion.

[0008] S3. Finally, it is necessary to design and apply the joint loss function: For a series of problems in vehicle 3D perception, the loss function can be divided into six aspects and summarized into basic loss function and enhanced loss function. The basic loss function includes vehicle body type classification, vehicle body center point offset regression, vehicle body 3D box vertex regression, and vehicle body 3D size regression, while the enhanced loss function is the loss function of vehicle body embedding space constraints. This loss function also includes spatial reprojection loss based on camera calibration and vehicle IoU spatial constraint loss.

[0009] Preferably, the resolution of the vehicle 3D detection dataset is 1920×1080 and 1080×720, and the original RGB image and the pseudo image generated by the point cloud data are scaled to a size of 512×512.

[0010] Preferably, the downsampling factor is set to 4, and the trained vehicle three-dimensional detection network model is loaded, and the two pictures obtained in step 1 are respectively sent to the pseudo image backbone feature extraction module and the image backbone feature extraction module to generate a multi-scale weighted fusion feature map.

[0011] Preferably, the pseudo image is generated for the point cloud data, and combined with the 2D image, the two images are preprocessed into a predetermined size.

[0012] Preferably, the multi-scale weighted fusion feature map is sent to a multi-task detection head, and the three-dimensional information of the vehicle can be regressed through an attention mechanism and multi-layer convolution.

[0013] Preferably, the ResNet-50 is a deep convolutional neural network that can be used for tasks such as image classification, target detection, and image segmentation. It uses residual blocks to solve the gradient vanishing problem in deep neural networks, so that deeper networks can be trained. Feature extraction is an important application of ResNet-50. By inputting an image into the network, a high-level feature representation of the image can be extracted for subsequent tasks.

[0014] Preferably, the joint loss function includes a basic loss and an enhanced loss, and there are branch loss functions under the basic loss and the enhanced loss, the branch loss functions include vehicle type loss, vehicle three-dimensional information regression loss and loss function of embedded space constraints, and the vehicle three-dimensional information regression loss includes the offset rate regression loss of the vehicle body center point, the eight vertices regression loss of the vehicle body three-dimensional box and the three-dimensional size regression loss of the vehicle body.

[0015] Preferably, the loss function of the embedding space constraint is divided into two parts, namely, the reprojection constraint of camera calibration and the vehicle IoU constraint, and the state quantity of the vehicle is solved by optimizing the visual reprojection constraint, and the most critical part is to derive the Jacobian of the error with respect to the state quantity.

[0016] Preferably, the constructed roadside perspective vehicle three-dimensional detection dataset SVLD-3D, labeling tool LabelImg-3D and evaluation system are experimentally verified on the SVLD-3D dataset. When the intersection-over-union ratio threshold is 0.7, the average three-dimensional accuracy of the network is 51.30%, the frame rate reaches 41.18, and the vehicle three-dimensional spatial positioning and three-dimensional size prediction accuracies are 98% and 85% respectively, which can meet the accuracy and speed requirements of actual three-dimensional perception.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] The present invention can overcome the problems of low accuracy, single solution information and over-reliance on templates under the existing technology. At the same time, the method adopted in this solution is used to fuse multi-scale data feature maps to generate multi-scale weighted fusion feature maps, and then use 1*1 convolution and attention mechanism on the feature maps to regress and obtain vehicle information heat maps. After decoding, the three-dimensional information of the vehicle is obtained. The network designed by this solution can quickly extract and fuse the input images, thereby realizing the solution of the three-dimensional information of the vehicle from the perspective of a monocular roadside camera. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The network structure diagram of the present invention;

[0020] Figure 2 It is a pre-processing flow chart of the present invention;

[0021] Figure 3 It is a schematic diagram of weighted feature map fusion of the present invention;

[0022] Figure 4 Schematic diagram of the joint loss function of the present invention. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] See also Figure 1-4 The present invention provides a vehicle three-dimensional information processing method based on multi-scale data fusion, and the information processing method comprises the following steps:

[0025] S1. First, data set and preprocessing are required: the data set used is the vehicle three-dimensional detection data set SVLD-3D. The algorithm verification and data annotation process are completed with the assistance of independent annotation software to ensure that the data maintains a certain numerical resolution. It contains video data of 21 scenes from the left, middle and right perspectives in seven highway scenes. At the same time, the data set also contains annotation information such as spatial position and speed collected by Lidar and GPS. During the network input process, the original RGB image and the pseudo image generated by the point cloud data are scaled to a certain size, and then the image is downsampled to determine the downsampling factor size, and then the processed image is sent to the ResNet50 backbone feature extraction network, such as Figure 1The network structure used in this article is a deep network structure based on data fusion. First, a pseudo image is generated using the Pillar feature extraction network, and the RGB image passes through the pseudo image backbone feature extraction module and the image backbone feature extraction module respectively; then, a multi-scale weighted fusion feature map is output through continuous projection convolution fusion; and the detection result is output after passing through the multi-task detection head.

[0026] S2. Then the vehicle 3D detection network is needed: This paper uses ResNet-50 as the backbone feature extraction network, because the network contains a residual structure and has good feature extraction capabilities. The feature map obtained by the network is then sent to the self-designed weighted multi-scale feature fusion module, which is to better adapt to multi-scale vehicle information. At the same time, this paper deconvolves the five feature maps output by the feature extraction network and upsamples them to feature sets of the same size, and then performs weighted feature fusion. The strategy formula is as follows:

[0027]

[0028] Among them, P i ' is the feature map after deconvolution and upsampling, w i is the weight of the corresponding feature map, and

[0029] The ratio can be set according to the specific experimental task. This paper adopts the average value, that is, 0.2.

[0030] On the basis of multi-scale fusion feature maps, different information is used as the output of the detection head according to the actual needs of vehicle 3D detection. The detection head includes four branches: vehicle type classification, vehicle center point regression, vehicle three-dimensional box vertex regression and vehicle 3D size regression. In the vehicle type classification detection head, the attention mechanism is used to improve the network's ability to distinguish different types of vehicles, while the other three branches are only implemented by direct regression analysis of the full convolution layer;

[0031] S3. Finally, the joint loss function needs to be designed and applied: For a series of problems in vehicle 3D perception, the loss function can be divided into six aspects and summarized as basic loss function and enhanced loss function. The basic loss function includes body type classification, body center point offset regression, body 3D box vertex regression, and body 3D size regression, while the enhanced loss function is the loss function of body embedding space constraint. This loss function also includes spatial reprojection loss based on camera calibration and vehicle IoU space constraint loss. The schematic diagram of the joint loss function is as follows Figure 4 The formula of the joint loss function is as follows:

[0032] L = α(λc L c +λ co L co +λ v L v +λ s L s )+β(λ proj L proj +λ iou L iou )

[0033] Among them, α and β are the weighted ratios of the basic loss and the enhanced loss, usually 1:1; the value of λ is to balance the various losses and add different weights to different loss functions. c =1,λ co =1,λ v =0.1,λ s =0.1,λ proj =0.1,λ iou =1, which is determined by the value ratio of these loss functions. The branch loss function formula is as follows:

[0034] Vehicle Type Loss:

[0035]

[0036] Where N is the number of positive samples, η and ξ are hyperparameters for adjusting the loss weights of positive and negative samples, respectively, usually set to 2 and 4, and p xyz It is the response value of each correctly labeled vehicle target in the feature map represented by a Gaussian kernel function.

[0037] Vehicle 3D information regression loss:

[0038] The loss function includes three types: the regression loss of the offset rate of the center point of the vehicle body, the regression loss of the eight vertices of the three-dimensional box of the vehicle body, and the regression loss of the three-dimensional size of the vehicle body. They can be expressed by L o ,L p ,L s Indicates. All three methods use absolute error loss. The loss function formula is as follows:

[0039]

[0040]

[0041]

[0042] in, Indicates whether there is a vehicle target at i, j in the output feature map, p c and p vrepresents the real coordinates of the center point of the vehicle and the vertices of the 3D box relative to the scaled input image size h,w, Relative to the output feature map size, A feature map representing the true 3D size of the vehicle.

[0043] Loss function for embedding space constraints:

[0044] This part of the loss function is divided into two parts, namely the camera calibration reprojection constraint and the vehicle IoU constraint, respectively. rccc ,L iou It is expressed as follows:

[0045]

[0046]

[0047] in, Represents the 3D feature map of the vehicle obtained by using camera calibration prior knowledge and network prediction The calculated vehicle 3D frame reprojection vertex feature map, represents the feature map of the predicted vehicle 3D box vertices relative to the original unscaled image size, and It represents the predicted value and true value of the minimum bounding rectangle calculated by the vertex feature map of the vehicle 3D box relative to the original unscaled image size, and IoU represents the calculation strategy of the intersection over union loss.

[0048] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A vehicle three-dimensional information processing method based on multi-scale data fusion, characterized in that: The information processing method includes the following steps: S1. First, data set and preprocessing are required: the data set used is the vehicle three-dimensional detection data set SVLD-3D. The algorithm verification and data annotation process are completed with the assistance of independent annotation software, so that the data maintains a certain numerical resolution. It contains video data of 21 scenes from the left, middle and right perspectives in seven highway scenes. At the same time, the data set also contains annotation information such as spatial position and speed collected by Lidar and GPS. During the network input process, the original RGB image and the pseudo image generated by the point cloud data are scaled to a certain size, and then the image is downsampled to determine the downsampling factor size, and then the processed image is sent to the ResNet-50 backbone feature extraction network; S2. Then the vehicle 3D detection network is needed: This paper uses ResNet-50 as the backbone feature extraction network. This is because the network contains a residual structure and has good feature extraction capabilities. The feature map obtained by the network is then sent to the self-designed weighted multi-scale feature fusion module. This is to better adapt to multi-scale vehicle information. At the same time, this paper deconvolves the five feature maps output by the feature extraction network and upsamples them to feature sets of the same size. Then, weighted feature fusion is performed; S3. Finally, it is necessary to design and apply the joint loss function: For a series of problems in vehicle 3D perception, the loss function can be divided into six aspects and summarized into basic loss function and enhanced loss function. The basic loss function includes vehicle body type classification, vehicle body center point offset regression, vehicle body 3D box vertex regression, and vehicle body 3D size regression, while the enhanced loss function is the loss function of vehicle body embedding space constraints. This loss function also includes spatial reprojection loss based on camera calibration and vehicle IoU spatial constraint loss.

2. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1, characterized in that: The resolution of the vehicle 3D detection dataset is 1920×1080 and 1080×720, and the original RGB image and the pseudo image generated by the point cloud data are scaled to 512×512.

3. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1, characterized in that: The downsampling factor is set to 4, and the trained vehicle 3D detection network model is loaded. The two pictures obtained in step 1 are respectively sent to the pseudo image backbone feature extraction module and the image backbone feature extraction module to generate a multi-scale weighted fusion feature map.

4. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1, characterized in that: The pseudo image is generated for the point cloud data, and combined with the 2D image, the two images are preprocessed into a predetermined size.

5. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1, characterized in that: The multi-scale weighted fusion feature map is sent to the multi-task detection head, and the three-dimensional information of the vehicle can be regressed through the attention mechanism and multi-layer convolution.

6. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1 is characterized in that: The ResNet-50 is a deep convolutional neural network that can be used for tasks such as image classification, target detection, and image segmentation. It uses residual blocks to solve the gradient vanishing problem in deep neural networks, so that deeper networks can be trained. Feature extraction is an important application of ResNet-50. By inputting images into the network, high-level feature representations of the images can be extracted for subsequent tasks.

7. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1 is characterized in that: The joint loss function includes a basic loss and an enhanced loss, and there are branch loss functions under the basic loss and the enhanced loss. The branch loss functions include vehicle type loss, vehicle three-dimensional information regression loss and loss function of embedded space constraints. The vehicle three-dimensional information regression loss includes the offset rate regression loss of the vehicle body center point, the eight vertices regression loss of the vehicle body three-dimensional box and the three-dimensional size regression loss of the vehicle body.

8. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1 is characterized by: The loss function of the embedding space constraint is divided into two parts, namely the reprojection constraint of camera calibration and the vehicle IoU constraint, and the state of the vehicle is solved by optimizing the visual reprojection constraint. The most critical part is to derive the Jacobian of the error with respect to the state.

9. The vehicle three-dimensional information processing method based on multi-scale data fusion according to claim 1, characterized in that: The constructed roadside vehicle three-dimensional detection dataset SVLD-3D, the annotation tool Label Img-3D and the evaluation system are experimentally verified on the SVLD-3D dataset. When the intersection-over-union ratio threshold is 0.7, the average three-dimensional accuracy of the network is 51.30%, the frame rate is 41.18, the vehicle three-dimensional spatial positioning and three-dimensional size prediction accuracies are 98% and 85% respectively, which can meet the accuracy and speed requirements of actual three-dimensional perception.

Citation Information

Cited By

  • Semantic grid map generation method and system for autonomous vehicle

    CN120912803A