Vehicle formation 3D target detection method based on license plate information

By using license plate information and deep convolutional networks in the 3D object detection method, combined with point cloud and image information, the problem of decreasing recognition accuracy of new energy vehicles and special vehicles is solved, and stronger generalization capabilities and the safety of autonomous vehicle formations are achieved.

CN119942520APending Publication Date: 2025-05-06BEIJING UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411844164.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-15
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When facing new energy vehicles and special vehicles, the existing 3D target detection methods have decreased recognition accuracy and poor generalization capabilities, making it difficult to accurately identify and track vehicles in the autonomous vehicle fleet.

Method used

The vehicle formation 3D object detection method based on license plate information is adopted, and the vehicle position and identity information provided by license plate recognition is combined with point cloud information and image information through a deep convolution network, and the vehicle position and identity information are used to enhance the generalization ability of the 3D object detection model, and the concealment and identification information are provided through "digital camouflage license plates".

Benefits of technology

It improves the generalization ability of the 3D object detection model, can accurately identify and distinguish different vehicles, reduces the situation of following the wrong vehicles, and provides concealment and identity identification information in special vehicles, enhancing the safety and reliability of the formation of autonomous driving vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942520A_ABST
    Figure CN119942520A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle formation 3D target detection method based on license plate information, and the method comprises the steps: carrying out the preliminary judgment of a point cloud distribution region similar to a vehicle object through the voxelized point cloud information, judging the information of a license plate and the region of a picture through the RGB image information, and obtaining some 2D RoLs through the projection of 3D RoLs obtained in a point cloud space to the picture, and judging which 3DRoLs need to be reserved by comparing the 2D RoLs with the area where the license plate is located, then cutting the point cloud space, and reserving the area where the vehicle exists. Meanwhile, a depth map is generated through the real point cloud and the image, and then a dense point cloud is generated. Point cloud information and license plate recognition information are rested and then connected, and then the information is used for predicting the confidence coefficient and bounding box of the vehicle. Meanwhile, the invention further provides a digital camouflage license plate, and the concealment of the special vehicle cannot be damaged while the functions can be guaranteed to be achieved. The accuracy of vehicle formation tracking can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and computer vision, and in particular relates to a vehicle formation 3D target detection method based on license plate information. Background Art

[0002] With the vigorous development of artificial intelligence technology, the vehicle's autonomous driving technology has gradually matured and is used in various fields including vehicle platooning. 3D object detection technology is one of its key technologies. The current mainstream 3D object detection methods mainly include detection methods based on LiDAR point cloud and detection methods based on binocular vision.

[0003] Among them, the detection method based on LiDAR point cloud is limited by the sparsity of point cloud and is difficult to provide sufficient effective information. The vision-based method has a natural disadvantage in judging the vehicle's posture, and the camera is easily affected by external ambient light, which reduces the reliability of the system. Some cutting-edge research is trying to use multi-sensor fusion methods to perform 3D target detection, but the current methods either do not make full use of information or invest all of it in deep neural network training. Although a large amount of information can be extracted, the use of information is too rough and will be interfered by a large amount of invalid information. In addition, the 3D target detection model trained in this way has poor generalization ability.

[0004] At the same time, in the case of autonomous driving of vehicle platoons, the point cloud-based detection method, when faced with vehicles of similar or even identical shapes on the road, cannot accurately distinguish between different vehicles, and it is easy to follow the wrong vehicle.

[0005] With the development of new energy vehicles, new power structures have been introduced into vehicle design, resulting in significant changes in vehicle appearance in recent years. This phenomenon will intensify with the development of new energy vehicles and the continuous innovation of enterprises. The models trained by traditional 3D object detection solutions for vehicles will continue to decline in accuracy in the face of constantly updated vehicle appearances. It is extremely costly for car companies to continuously learn new data to ensure that they can identify new vehicles, and the learning process is uncertain. The hidden layer of deep learning itself has unknown potential risks, and constantly updating learning data is not easy. Model , which will increase this uncertainty, on the one hand increasing the testing costs of car companies, and on the other hand increasing the safety risks of actual cars on the road.

[0006] And this is the case when facing common civilian vehicles. In fact, when facing scenarios with military vehicles, on the one hand, there is a lack of public data sets to train 3D object detection models to adapt to special vehicles. On the other hand, in actual situations, 3D object detection models need to face the situation of detecting unknown vehicles.

[0007] An existing method to solve similar problems is to enhance the generalization ability of the 3D object detection model. However, if the recognition range of the 3D object detection model is simply relaxed, the model may mistakenly recognize other objects as target objects. Therefore, the constraints on the 3D object detection model during training cannot be simply relaxed. How to enhance the generalization ability of the 3D object detection model while ensuring the accuracy of model recognition is a major difficulty currently faced by the industry. Summary of the invention

[0008] In view of the problems mentioned in the technical background, the present invention proposes a 3D target detection method for vehicle formation based on license plate information. The method utilizes point cloud information and image information and is implemented using a deep convolutional network.

[0009] The specific steps include:

[0010] Step 1, hardware preparation: Install a lidar (Velodyne) and camera on the front of the rear vehicle in the convoy.

[0011] Step 2, device calibration: calibrate the coordinate transformation matrix (external parameters) from the LiDAR to the camera, and calibrate the camera internal parameters, and save the calibration data in the calib file of your own dataset. The calib file can be used to convert the LiDAR coordinate system and the camera coordinate system.

[0012] Step 3: prepare a data set, which includes RGB images captured by the camera and point cloud information captured by the lidar. The data set is a public data set, such as the KITTI data set, which includes point clouds, images, and calib files.

[0013] Step 4: The data set and calibration file calib in step 3 are used for model training and verification. The calibration camera internal and external parameters obtained in step 2 are used for real vehicle testing in vehicle formation and actual use of equipment.

[0014] Step 5: The original point cloud (in the real point cloud stream) provided by the lidar in step 3 is voxelized and used for RPN to perform preliminary region selection to select the area where the object point cloud of the vehicle is located as the preliminary region of interest 3DRoI.

[0015] Step 6: The RGB image captured by the camera in step 3 is used in two parallel lines in this method:

[0016] The first is to use RGB images for efficient license plate recognition, determine the license plate information and the location of the license plate on the RGB image, and based on the logic that an object with a license plate that looks like a vehicle must be a vehicle, preliminarily determine the approximate area of ​​the vehicle on the RGB image, namely the 2D RoI. Using the external parameters obtained in step 2, project the 3D RoI onto the RGB image plane to obtain a set of 2D RoIs. The RoI determined by the license plate and the RoI obtained from the real point cloud will jointly determine the target area for point cloud space cropping with a trained weight.

[0017] Steps 6 and 7 are in parallel and are two routes for using the same image. The purpose of step 6 is to help determine the point cloud space area where the vehicle exists, and the function of step 7 is to generate a "generated point cloud" (also known as a "pseudo-point cloud Pseudo-LiDAR").

[0018] Step 7, another parallel utilization line of the RGB image captured by the camera in step 3: first generate a depth image (RGBD image) with sparse depth information from the real point cloud information in step 3. The depth information of the depth image generated here comes from the point cloud, and the depth information density is low. Then the corresponding RGB image is used to supplement the information, and the foreground and background are used to infer the gaps between the sparse depth information to further generate a depth image with dense depth information. Then, a dense point cloud is generated from the depth image with dense depth information. This point cloud will be referred to as the "generated point cloud" below. The function of step 7 is to generate a "generated point cloud". The "generated point cloud" is a point cloud with a higher density.

[0019] Step 8: voxelize the "generated point cloud" obtained in step 7.

[0020] Step 9: The target area determined in step 6 is used to cut the real point cloud and the generated point cloud processed in step 8 respectively to obtain "real point cloud area A" and "generated point cloud area B".

[0021] Step 10: The corresponding spatial information of "real point cloud area A" and "generated point cloud area B" is consistent and has the same representation; the point cloud information density is different, and the features of each pair of grids of "real point cloud area A" and "generated point cloud area B" are fused to obtain a dense fused point cloud C.

[0022] Step 11, reshape the dense fused point cloud C obtained in step 10 and the output of the RGB image license plate detection mentioned in step 6, unify the processed point cloud information and the output of the license plate detection network into the same dimension, then connect the point cloud information and license plate information of the unified dimension and input them into the fully connected layer FC for predicting the vehicle confidence and bounding box.

[0023] Among them, during the network training process, when predicting the bounding box, an auxiliary detection head can be added to standardize the prediction results. After the training is completed, the auxiliary head can be removed to speed up the target detection speed.

[0024] In addition, the present invention also provides a special license plate, namely a "digital camouflage license plate", which can maintain the concealment of special vehicles under human vision, and like ordinary license plates, can be used for the above-mentioned network-assisted 3D target detection.

[0025] The "digital camouflage license plate" is consistent with the national standard license plate in terms of physical structure and physical size. Its structural size adopts the national standard style for different models, which is convenient for installation on existing vehicles. The difference from the national standard license plate is reflected in the surface color and pattern information of the license plate. The original intention of the design of the digital camouflage license plate is that special vehicles will remove or cover the national standard license plate when performing tasks to ensure their concealment when performing tasks, but this will affect the recognition accuracy of the 3D target detection method and make it difficult to distinguish the vehicle identity information. This is a problem that needs to be solved when the autonomous driving vehicle formation performs tasks. The digital camouflage license plate is designed to provide the computer with license plate information and identity information without destroying the concealment of special vehicles. The characteristic of the "digital camouflage license plate" is that its surface pattern is generated by a set of encryption programs. The encryption program can generate a digital camouflage pattern containing unique characteristics by inputting a known identification number, and the pattern is printed on the digital camouflage license plate. At the same time, the digital camouflage license plate still retains the black frame line required by the national standard license plate. This feature is retained to facilitate the license plate recognition network to quickly identify the license plate position. At the same time, because the frame line is relatively thin, on the digital camouflage pattern, at a certain distance, it hardly affects the concealment brought by the camouflage effect.

[0026] The digital camouflage license plate retains the physical structure, size and national standard black frame of the ordinary license plate. When training the model, the data set contains vehicles equipped with digital camouflage license plates in addition to vehicles equipped with ordinary license plates. When necessary, special vehicles are equipped with digital camouflage license plates. While the digital camouflage license plates do not destroy the concealment of special vehicles, the license plate detection network can detect the digital camouflage license plates. Unlike the ordinary license plate information detection network, when the network detects the digital camouflage license plate, it will use the corresponding decryption program to read out the vehicle identity information contained in the digital camouflage license plate. The rest is the same as processing ordinary license plates. The position of the digital camouflage license plate in the image is detected, and the point cloud area containing the vehicle is determined as in step 6.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1. The present invention has a breakthrough in judgment logic. Based on the judgment logic that objects with license plates that look like vehicles are vehicles, new vehicles and special vehicles with license plates but with certain differences in appearance from traditional vehicles can be detected. Judging vehicles based on this logic can greatly enhance the generalization ability of the 3D object detection model.

[0029] 2. The present invention has strong generalization ability and can accurately distinguish vehicle information. In addition to determining the area of ​​interest, the license plate recognition information also records vehicle information, which can be used to distinguish vehicles with the same or similar appearance, avoiding tracking errors when the front vehicle meets the rear vehicle. This is not possible with the traditional point cloud method.

[0030] 3. The present invention also provides a special license plate, namely a "digital camouflage license plate", which can provide concealment for special vehicles and assist the 3D target detection model in determining the vehicle position range and vehicle identity information, which is of extremely important strategic significance to the autonomous driving special vehicle formation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a network construction structure diagram in the present invention.

[0032] Figure 2 It is a generated point cloud structure diagram in the network structure of the present invention.

[0033] Figure 3 It is a structural diagram of point cloud fusion and feature extraction in the present invention.

[0034] Figure 4 This is an example diagram of the “digital camouflage license plate” in the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention will be described in detail below in conjunction with the drawings of this specification. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments:

[0036] The present invention provides a 3D target detection method for vehicle formations based on license plate information. The application scenario of this method is in the scenario of automatic driving of special vehicle formations. In special vehicle formations, the rear side will track the automatic driving of the front vehicle, and at the same time, it needs to face a no-communication environment in some actual environments, so the rear vehicle needs to be able to accurately detect the front vehicle. The detection method based on laser radar is easy to determine the position and posture of the object, but due to the sparsity of the point cloud, it is difficult to distinguish different objects. The image-based method has rich information and is easy to distinguish objects, but it is limited by the information dimension and it is difficult to determine the position and posture of the object.

[0037] The method of the present invention is based on the information captured by the laser radar and RGB camera installed on the front of the rear vehicle, wherein the camera provides 2D image information and the laser radar provides 3D point cloud information.

[0038] In general, the present invention utilizes point cloud information as follows: the original point cloud is processed by RPN to obtain the region of interest (RoI); the cropped and voxelized real point cloud and the generated point cloud are fused to enrich the point cloud information and obtain a dense point cloud; the vehicle confidence and (pose) bounding box are determined by the dense point cloud.

[0039] In general, the present invention uses RGB image information as follows: generate a depth map through images and point clouds, and then make a "generated point cloud" to enrich the sparse real point cloud information; use mature methods (such as YOLO) to perform efficient license plate recognition through images to obtain the location information of the license plate on the image and the information contained in the license plate. The license plate location information is used to combine with 3D RoI to determine the 3D area that needs to be cut and retained, and the license plate information is used to give the identified front vehicle "identity information" to ensure that the correct vehicle is tracked.

[0040] The network structure diagram of the entire invention is shown in Figure 1 . The present invention voxelizes the point cloud information so that the point cloud information in a certain space is integrated into spatial grids, laying a foundation for the subsequent point cloud information fusion. The voxelized point cloud is initially obtained through the region generation network (RPN) to obtain some 3D regions of interest (RoI). These 3D RoIs contain the point cloud information of the vehicle and some point cloud information of vehicles that resemble vehicles. Because the point cloud is sparse and has no color, it is difficult to further obtain more detailed information based on the judgment of the point cloud, but this is enough for the method of the present invention. In another line, the method performs license plate recognition on the image and obtains some license plate areas. Based on the judgment logic that a vehicle-like object with a license plate is a vehicle, the method projects the above-obtained 3D RoI onto the image through the coordinate transformation matrix and the camera internal parameters to obtain a set of 2D RoIs, and compares and judges the newly obtained 2D RoI and the license plate area according to the overlap ratio obtained by network training to determine the 3D RoI to be retained, that is, the vehicle-like object area with a license plate will be retained. The above 3D RoI is used to crop the voxelized original point cloud and crop the generated point cloud and voxelize it. Because the generated point cloud will have a lot of invalid information at a distance, cropping first and then voxelizing can save computing costs and improve efficiency. For the processing of generating point clouds, see Figure 2 .

[0041] The projection matrix P used for coordinate transformation is a 3×4 matrix with the following structure:

[0042]

[0043] Among them, f x and f y is the focal length, c x and c y are the coordinates of the principal point of the image.

[0044] At the same time, Tr_velo_to_cam data is also required, which represents the transformation matrix from the lidar to the camera. This data is obtained by real vehicle calibration.

[0045] After that, the generated point cloud region B obtained above and the real point cloud region A obtained above are fused according to their one-to-one corresponding spatial grids in space to enrich the sparse real point cloud information. The fusion process uses the attention module to predict a weight for each pair of grids, and use the weight to weight the pair of grid features to obtain the fused point cloud features. The structure of the point cloud fusion and feature extraction part can be found in Figure 3 .

[0046] Figure 3 In the second half, the dense point cloud information after fusion is reshaped, and the license plate information obtained by license plate recognition is also reshaped to the same size. The two pieces of information are concatenated and then put into the fully connected layer FC to predict the vehicle confidence and (pose) bounding box. At the same time, during the network training process, this part can add auxiliary detection heads to the point cloud before fusion to standardize the network. The auxiliary detection heads are removed during prediction to improve network performance.

[0047] In some actual working conditions of special vehicles, the concealment of special vehicles needs to be considered. At this time, traditional license plates are generally blocked or removed. The present invention also provides a special license plate, namely a "digital camouflage license plate", which is concealed from human vision, but the trained network of the present invention can still recognize the information provided by the digital camouflage license plate. While providing concealment for special vehicles, digital camouflage license plates can assist 3D target detection models in determining vehicle position, (pose) bounding box and vehicle identity information, which is of great strategic significance to autonomous driving special vehicle formations. See the schematic diagram of the "digital camouflage license plate" for details. Figure 4 .

Claims

1. A 3D target detection method for vehicle formation based on license plate information, characterized in that: The following steps are involved: Step (1), hardware preparation: including the laser radar and camera installed at the front of the rear vehicle in the vehicle formation; Step (2), equipment calibration: calibrate the coordinate transformation matrix from the laser radar to the camera and calibrate the camera internal parameters; Step (3), prepare a data set, the data set includes RGB images captured by a camera and point cloud information captured by a lidar; Step (4): the data and "calib" in step (3) are used for model training and verification, and the "calib" obtained in step (2) is used for testing and actual use of the device; Step (5), the original point cloud provided by the laser radar in step (3) is voxelized and used for RPN to perform preliminary region selection to select the region where the object point cloud of the vehicle is located as the preliminary 3D RoI; Step (6), the RGB image captured by the camera in step (3) has two utilization lines. One is: using the RGB image for efficient license plate recognition to determine the license plate information and the position of the license plate on the RGB image. Based on the logic that an object with a license plate that looks like a vehicle must be a vehicle, the approximate area of ​​the vehicle on the RGB image, i.e., the 2D RoI, is preliminarily determined. Using the external parameters obtained in step (7), the 3D RoI is projected onto the RGB image plane to obtain a set of 2D RoIs. The RoI determined by the license plate and the RoI obtained from the real point cloud will jointly determine the target area for point cloud space cropping with a trained weight. Step (8), the RGB image captured by the camera in step (3), the second is: first generate a sparse depth image from the real point cloud information in step (3), then supplement the information from the corresponding RGB image to generate a dense depth image, and then generate a dense point cloud from the dense depth image, hereinafter referred to as "generated point cloud"; Step (9), voxelizing the "generated point cloud"; Step (10), cutting the real point cloud and the generated point cloud respectively according to the target area determined in step (6) to obtain "real point cloud area A" and "generated point cloud area B"; Step (11), "real point cloud area A" and "generated point cloud area B"; the corresponding spatial information is consistent and has the same representation, but the point cloud information density is different. The features of each pair of grids of "real point cloud area A" and "generated point cloud area B" are fused to obtain a dense fused point cloud C; Step (12), reshape the dense fused point cloud C obtained in step (10) and the output of the RGB image license plate detection mentioned in step (6), then connect the two and input them into FC for predicting vehicle confidence and bounding box.

2. The vehicle formation 3D target detection method based on license plate information according to claim 1 is characterized in that: For the real radar point cloud processing flow based on Voxel-RCNN, the real point cloud is voxelized, that is, the space is evenly subdivided into grids, and the points distributed in the grid are aggregated into the information of this cell.

3. The vehicle formation 3D target detection method based on license plate information according to claim 1 is characterized in that: For camera image processing, the process is divided into two parts; The first is to use the YOLO algorithm to first identify the license plate information in the RGB image, including the regional position of the license plate on the original image and the identity information contained in the license plate; The second is to generate a sparse depth map through real point cloud information, and then use the input camera image to assist in depth completion to obtain a dense depth map, and then generate a point cloud from the dense depth map.

4. The vehicle formation 3D target detection method based on license plate information according to claim 1 is characterized in that: The real point cloud and the generated point cloud need to be cropped, and their 3D RoIs are determined by the multiple 3D RoIs obtained by RPN in step (5) and the license plate position obtained in step (6) through the coordinate range: The calib obtained by step (2) is projected onto the plane image to obtain a new 2D RoI and the overlap with the license plate area obtained by license plate detection in step (6) is calculated. The 3D RoI to be retained is determined by the overlap as the cropped point cloud; By projecting the 2D license plate area into the real point cloud 3D space, the probability that the 3D RoI contains voxels appearing in the license plate projection area is calculated. The retained 3D RoI is determined based on this probability and used for cropping; The calib file obtained in step (2) contains the camera intrinsic parameters and the projection matrix P, which is a 3x4 matrix with the following structure: where f x and f y is the focal length, c x and c y are the coordinates of the principal point of the image.

Citation Information

Patent Citations

  • Vehicle type identification method of vehicles of same brand based on space position information

    CN103500327A

  • Vehicle tracking method based on video analysis

    CN105913000A

  • Feature fusion-based vehicle behavior recognition method in urban traffic scene

    CN106781513A

  • Camouflage license plate detection and identification method based on deep learning

    CN118470701A

  • Positive azimuth towing guidance method for road rescue equipment based on license plate corner features

    US20210312653A1