An Autonomous Driving 3D Object Detection Method Based on FPN Swin Transformer and Pointnet++

Through the combination of FPN Swin Transformer and Pointnet++, the feature extraction and point cloud matching from multiple perspectives are used to solve the difficulties of 3D object detection in autonomous driving, improve the comprehensiveness and accuracy of the detection, enhance the robustness of the network, and adapt to complex driving scenarios and inclement weather.

CN116403186BActive Publication Date: 2025-07-25NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310334275.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-07-25
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

In the prior art, in autonomous driving, there are problems such as 3D object detection difficulties, single frame detection single category, detection difficulties caused by the diversity and complexity of driving scenarios, and poor network robustness caused by the impact of light and weather.

Method used

Using a method based on FPN Swin Transformer and Pointnet++, the forward-view image and lidar point cloud data are obtained, inverse perspective transformation and projection fusion are performed, and feature extraction is performed by combining the FPN module and the Swin Transformer module. The detection frame and point cloud region matching are used in multiple perspectives to perform three-dimensional boundary regression and classification.

Benefits of technology

It improves the comprehensiveness and accuracy of object detection, improves the ability to detect small objects, enhances the robustness of the network, reduces the impact of light and weather on detection, and improves the efficiency and accuracy of three-dimensional bounding box regression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403186B_ABST
    Figure CN116403186B_ABST
Patent Text Reader

Abstract

The present invention discloses an autonomous driving three-dimensional object detection method based on FPN Swin Transformer and Pointnet++. This method uses cameras and lidar to obtain the front-view images and point cloud information of the road conditions. Through inverse perspective transformation and projection correspondence, the front-view images and bird's-eye images fused with point cloud information are obtained. Inputting them into the FPN Swin Transformer network for feature extraction can obtain the two-dimensional detection boxes and classification results of the objects from two perspectives. Through the work of frustum point cloud extraction, the candidate point cloud regions of the objects are obtained and the three-dimensional bounding box regression and classification results of the objects can be obtained through the Pointnet++ network. Finally, the final object classification result is obtained by comprehensively considering the object classification results under the two networks. The present invention can effectively solve the problems of incomplete object detection, difficult detection of object three-dimensional information, inaccurate object classification results and poor robustness in the field of autonomous driving by multi-level fusion of image and point cloud information and using the method of three-dimensional bounding box regression based on two-dimensional detection boxes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the three-dimensional object detection task in the field of autonomous driving, specifically to a three-dimensional object detection method for autonomous driving based on FPN SwinTransformer and Pointnet++. Background Art

[0002] In recent years, with the continuous improvement of the market's demand for automotive active safety and intelligence, the huge social and economic value of autonomous driving has become increasingly prominent, and more and more enterprises and research institutions have actively participated in and promoted the development of the field of autonomous driving. Autonomous driving is a complex system combining software and hardware, mainly divided into three major technical modules: perception, decision-making, and control. The perception module mainly provides environmental information for autonomous driving through high-precision sensors such as cameras and lidar; the decision-making module makes decisions such as path planning in the platform based on the vehicle positioning and surrounding environment data provided by the perception system; the control module achieves vehicle control effects in ways such as adaptive control and cooperative control, combined with vehicle hardware devices. Among them, environmental perception involves a variety of different sensors, which is the premise and foundation for the safe, autonomous, and reliable driving of autonomous vehicles, and the object detection task is the most crucial part of the perception task. Object detection refers to the task of giving various information about obstacles such as vehicles in the autonomous driving scenario.

[0003] Patent CN114966603A proposes an image-driven laser point cloud object detection method and system. The frustum point cloud extracted from the two-dimensional detection frame can effectively improve the object detection effect through two-step networks: the detection frame prediction network and the detection frame optimization network. However, it does not make full use of the image feature information extracted in the early stage and the object classification results. Patent CN114387202A proposes a 3D object detection method based on the fusion of vehicle endpoint cloud and image, which reflects the feasibility of obtaining the candidate point cloud region by processing the frustum point cloud from the two-dimensional detection frame of the object. However, there are problems such as incomplete object detection and too large candidate point cloud regions when extracting the point cloud region only through the two-dimensional boundary box of the object in one view, which reduces the extraction speed of subsequent point cloud features. Summary of the Invention

[0004] The purpose of the present invention is to address the problems existing in the above-mentioned prior art, and propose a three-dimensional object detection method for autonomous driving based on FPN SwinTransformer and Pointnet++, which can improve the difficulties in detecting important small objects in 3D object detection in autonomous driving, the problem of detecting a single category with a single framework, the detection difficulties brought by the diversity and complexity of the driving scenario, the influence of light and weather on sensors, and the poor robustness of the object detection network.

[0005] An autonomous driving 3D object detection method based on FPN Swin Transformer and Pointnet++ includes the following steps:

[0006] Step 1), obtain the front view image and lidar point cloud data of the road condition during vehicle driving;

[0007] Step 2), perform inverse perspective transformation on the front view image to obtain a bird's-eye view image of the road condition, project the lidar point cloud onto the front view image and the bird's-eye view image respectively, and obtain the front view image and the bird's-eye view image with fused point cloud features;

[0008] Step 3), respectively perform feature extraction on the front view image and the bird's-eye view image with fused point cloud features through FPN Swim Transformer, and obtain the target 2D detection box, target classification result in the front view perspective, and the target 2D detection box, target classification result in the bird's-eye view perspective;

[0009] The FPN Swin Transformer includes a Swin Transformer module and an FPN module;

[0010] The Swin Transformer module includes a Patch Partition module and first to fourth feature extraction modules;

[0011] The construction steps of the Swin Transformer module are as follows:

[0012] Step 3.1.1), construct a Patch Partition module to block the image with fused point cloud features, and block the input image with fused point cloud features of size H×W×3 into four images of H / 4×W / 4×48;

[0013] Step 3.1.2), construct a first feature extraction module to perform feature extraction on the image of H / 4×W / 4×48 obtained by the Patch Partition module, and obtain a feature map of H / 4×W / 4×C;

[0014] The first feature extraction module consists of a Linear Embeding layer and 2 consecutive SwinTransformer Blocks in sequence;

[0015] The construction method of the Swin Transformer Block is as follows: replace the standard multi-head self-attention module in the Transformer with a module based on a moving window, keep other layers unchanged, and apply a LayerNorm layer before each MSA module and each MLP;

[0016] Step 3.1.3), construct a second feature extraction module for extracting middle-level features, perform feature extraction on the feature map of H / 4×W / 4×C obtained by the first feature extraction module, and obtain a feature map of H / 8×W / 8×2C;

[0017] The second feature extraction module is sequentially composed of a Patch Merging layer and six Swin Transformer Block layers;

[0018] Step 3.1.4), construct a third feature extraction module, perform feature extraction on the feature map of H / 8×W / 8×2C obtained in the second feature extraction stage, and obtain a feature map of H / 16×W / 16×4C;

[0019] The third feature extraction module is sequentially composed of a Patch Merging layer and six Swin Transformer Block layers;

[0020] Step 3.1.5), construct a fourth feature extraction module, perform feature extraction on the feature map of H / 16×W / 16×4C obtained in the third feature extraction stage, and obtain a feature map of H / 32×W / 32×8C;

[0021] The fourth feature extraction module is sequentially composed of a Patch Merging layer and two Swin Transformer Block layers;

[0022] The construction steps of the FPN module are as follows:

[0023] Step 3.2.1), construct four Conv2d(1×1,s1) modules to perform convolution operations on the feature maps obtained by the first to fourth feature extraction modules respectively, transform the feature map of H / 32×W / 32×8C obtained by the fourth feature extraction module into a feature map of H / 32×W / 32×C, transform the feature map of H / 16×W / 16×4C obtained by the third feature extraction module into a feature map of H / 16×W / 16×C, transform the feature map of H / 8×W / 8×2C obtained by the second feature extraction module into a feature map of H / 8×W / 8×C, and transform the feature map of H / 4×W / 4×C obtained by the first feature extraction module into a feature map of H / 4×W / 4×C;

[0024] Step 3.2.2), construct three upsampling working and fusion modules to perform scale change operations on the feature maps obtained by the four Conv2d(1×1, s1) modules respectively and fuse the feature maps of the same scale. Transform the feature map of H / 32×W / 32×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 16×W / 16×C and fuse it with the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 8×W / 8×C and fuse it with the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 4×W / 4×C and fuse it with the feature map of H / 4×W / 4×C obtained by the Conv2d(1×1, s1) module;

[0025] Step 3.2.3), construct four Conv2d(3×3, s1) modules to perform convolution operations on the feature maps obtained by the three upsampling working and fusion modules respectively, and the feature map of H / 32×W / 32×8C obtained by the Conv2d(1×1, s1) module. This convolution operation will not affect the scale of the feature map;

[0026] Step 3.2.4), construct a Maxpool(1×1, s2) module to perform pooling operations on the feature map of H / 32×W / 32×C in the feature maps obtained by the four Conv2d(3×3, s1) modules, and obtain a feature map of H / 64×W / 64×C;

[0027] Step 3.2.5), construct a Contact module to fuse and connect the feature maps of H / 32×W / 32×8C, H / 16×W / 16×C, H / 8×W / 8×C, H / 4×W / 4×C obtained by the four Conv2d(3×3, s1) modules and the feature map of H / 64×W / 64×C obtained by performing pooling operations through the Maxpool(1×1, s2) module, and obtain a fused connection feature map;

[0028] Step 3.2.6), construct a Fully Contected Layer to perform a fully connected operation on the fused connection feature map, and obtain the two-dimensional detection box of the image target and the target classification result;

[0029] Step 4), perform point cloud extraction operations on the two-dimensional detection box of the target in the front view and the two-dimensional detection box of the target in the bird's-eye view respectively, and obtain the frustum point cloud region in the front view and the cylinder point cloud region in the bird's-eye view:

[0030] Step 4.1), based on the camera imaging principle, obtain the frustum region projected from the target 2D detection box in the front view angle to the three-dimensional space according to the target 2D detection box in the front view angle, and obtain the cylinder region projected from the target 2D detection box in the bird's-eye view angle to the three-dimensional space according to the target 2D detection box in the bird's-eye view angle;

[0031] Step 4.2), considering the internal parameters of the camera and the lidar and the rotation matrix and translation vector between the two, realize the coordinate transformation of the point cloud from the lidar coordinate system to the camera coordinate system; if the point cloud is located in the cone region or cylinder region projected from the target 2D detection box to the three-dimensional space, it means that they can be projected into the 2D bounding box of the target, and extract the information of this part of the point cloud for subsequent regression of the target's 3D bounding box; obtain the frustum point cloud space region corresponding to the front view angle and the cylinder point cloud space region corresponding to the bird's-eye view angle through the point cloud coordinate transformation and extraction work respectively;

[0032] Step 5), match the frustum point cloud space region corresponding to the front view angle of each target and the cylinder point cloud space region corresponding to the bird's-eye view angle of each target, and obtain the candidate point cloud region of the target by extracting the overlapping space region:

[0033] Compare the point cloud coordinates of the frustum point cloud space region of each target with the point cloud coordinates of the cylinder point cloud space region. The point cloud coordinates that appear simultaneously in the frustum point cloud space region and the cylinder point cloud space region are the candidate point clouds, and all candidate point clouds form the point cloud candidate region;

[0034] Step 6), perform target point cloud segmentation on the candidate point cloud region and then use Pointnet++ to extract the point cloud features to obtain the target 3D bounding regression box and the target classification result under the spatial point cloud;

[0035] Step 7), obtain the final classification result of the target by comprehensively considering the target classification result in the front view angle, the target classification result in the bird's-eye view angle, and the target classification result under the spatial point cloud.

[0036] As a further optimization scheme of the three-dimensional target detection method for autonomous driving based on FPN Swin Transformer and Pointnet++ of the present invention, in step 1), the lidar point cloud data is collected by the lidar, the front view image of the road condition during the vehicle driving process is collected by the optical camera, and the lidar point cloud and the front view image of the corresponding frame are obtained by intercepting the same timestamp.

[0037] As a further optimization scheme of the three-dimensional target detection method for autonomous driving based on FPN Swin Transformer and Pointnet++ of the present invention, the specific steps of step 2) are as follows:

[0038] Step 2.1), calibrate the camera by the method of checkerboard calibration to obtain the internal and external parameters of the camera, and deduce the conversions between the vehicle body coordinate system, the camera coordinate system, and the pixel coordinate system through coordinate relationships as follows:

[0039]

[0040] In the formula, is the pixel coordinate system, is the camera internal parameter matrix, is the vehicle body coordinate system, is the camera coordinate system, Z c is the distance between this point and the imaging plane in the camera axis direction, f x and f y are the equivalent focal lengths of the camera in the x - direction and y - direction respectively, u0 and v0 are the horizontal and vertical pixel coordinates of the image center respectively, R c is the rotation matrix from the camera coordinate system to the vehicle body coordinate system, T c is the translation matrix from the camera coordinate system to the vehicle body coordinate system;

[0041] Step 2.2), perform inverse perspective transformation on the front - view image by combining the internal and external parameters of the camera, and convert the front - view image from the pixel coordinate system to the top - view angle of the world coordinate system, that is, convert it into a bird's - eye view, eliminate the interference of perspective distortion on road condition information and distance error, and present the true - world top - view characteristics. The mapping relationship between the pixel coordinate system of the perspective image and the top - view plane of the world coordinate system is as follows:

[0042]

[0043]

[0044] In the formula, X and Y are the horizontal and vertical coordinates of the perspective image in the top - view plane of the world coordinate system respectively, u t and v t are the horizontal and vertical pixel coordinates of the perspective image respectively, θ is the angle between the camera optical axis and the horizontal plane on the vehicle mid - plane, h is the distance from the camera to the ground, and d0 is the distance from the camera to the front end of the vehicle;

[0045] The conversion relationship between the pixel coordinate system of the inverse perspective transformation image and the top - view plane of the world coordinate system is as follows:

[0046]

[0047]

[0048] In the formula, u n and v n are the horizontal and vertical pixel coordinates of the inverse perspective transformation image respectively, W IPM and h IPMThey are the pixel width and height of the inverse perspective image respectively. σ1 and σ2 are the actual distances of a unit pixel in the horizontal and vertical directions of the inverse perspective image in the horizontal direction of the world coordinate system respectively. d1 is the distance between the lowest part of the camera's field of view and the front end of the vehicle.

[0049] In step 2.3), after determining the corresponding relationship between the pixels of the front view image and the radar points of the lidar point cloud data, combined with the internal parameters of the camera, solve the linear equations for the rotation matrix and the translation vector to obtain the rotation matrix and the translation vector between the camera and the lidar, and realize the joint calibration of the camera and the lidar:

[0050] In step 2.3.1), according to the perspective imaging model, use the external parameter matrix to multiply with the point cloud coordinates in the Cartesian coordinate system to convert the point cloud to the camera coordinate system; project the point through the internal parameter matrix to the pixel coordinate system to obtain the corresponding pixel point Complete the spatial alignment and registration of the lidar point cloud and the monocular camera image. The conversion relationship is:

[0051]

[0052] In the formula, is the lidar coordinate system coordinate of the point, is the coordinate of the point in the camera coordinate system, is the coordinate of the point in the pixel coordinate system, K is the internal parameter matrix of the camera, is the rotation matrix from the lidar coordinate system to the camera coordinate system, is the translation matrix from the lidar coordinate system to the camera coordinate system.

[0053] As a further optimization scheme of the three-dimensional object detection method for autonomous driving based on FPN Swin Transformer and Pointnet++ of the present invention, when making a comprehensive consideration in step 7), introduce the class confidence formula: P f = 0.4P1 + 0.4P2 + 0.2P3;

[0054] In the formula, P f is the confidence;

[0055] is the judgment of the object category by FPN Swin Transformer under the front view angle, p 1a , p 1b , p 1c are the probability values of the classification results judged by FPN Swin Transformer as category a, category b, and other category c respectively under the front view angle;

[0056] For the judgment of object categories by the FPN Swin Transformer from an aerial view, p 2a 、p 2b 、p 2c are the probability values that the classification results of the FPN Swin Transformer from an aerial view are category a, category b, and other category c respectively;

[0057] For the judgment of object categories by Pointnet++ under the spatial point cloud, p 3a 、p 3b 、p 3c are the probability values that the classification results of the FPN Swin Transformer from an aerial view are category a, category b, and other category c respectively.

[0058] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0059] 1. The object detection method of the present invention uses image and lidar point cloud data to obtain more comprehensive road condition information.

[0060] 2. The present invention projects and fuses the lidar point cloud onto the image, which can enrich the information of the image, thereby solving to a certain extent the problem of incomplete image data caused by poor light and rain or snow weather.

[0061] 3. The FPN Swin Transformer network of the present invention can effectively improve the feature extraction ability of the network by fusing low-level features and high-level features through FPN, so as to improve the accuracy of the target two-dimensional bounding box and target classification;

[0062] 4. By extracting the overlapping part of the two frustum point clouds of the target from different perspectives, the present invention can effectively narrow the range of the point cloud candidate region and improve the accuracy and efficiency of subsequent point cloud segmentation and target three-dimensional bounding box regression.

[0063] 5. By comprehensively judging the classification results of the target in the FPN Swin Transformer network and the Pointnet++ network, the present invention can effectively improve the accuracy of target category detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is the overall framework diagram of the present invention;

[0065] Figure 2 is the schematic diagram of the frustum point cloud optimization process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0066] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings:

[0067] The present invention can be implemented in many different forms and should not be considered limited to the embodiments described herein. On the contrary, these embodiments are provided so that the present disclosure is thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. In the drawings, components are enlarged for clarity.

[0068] As Figure 1 shown, the present invention discloses a 3D object detection method for autonomous driving based on FPN Swin Transformer and Pointnet++, comprising the following steps:

[0069] Step 1), obtaining a front view image and lidar point cloud data of the road condition during vehicle driving;

[0070] Step 2), performing inverse perspective transformation on the front view image to obtain a bird's-eye view image of the road condition, and projecting the lidar point cloud onto the front view image and the bird's-eye view image respectively to obtain a front view image and a bird's-eye view image with fused point cloud features;

[0071] Step 3), respectively performing feature extraction on the front view image and the bird's-eye view image with fused point cloud features through FPN Swim Transformer to obtain a target 2D detection box, a target classification result from the front view perspective, and a target 2D detection box, a target classification result from the bird's-eye view perspective;

[0072] The FPN Swin Transformer includes a Swin Transformer module and an FPN module;

[0073] The Swin Transformer module includes a Patch Partition module and first to fourth feature extraction modules;

[0074] The construction steps of the Swin Transformer module are as follows:

[0075] Step 3.1.1), constructing a Patch Partition module to partition the image with fused point cloud features, and partitioning the input image with fused point cloud features of size H×W×3 into four images of H / 4×W / 4×48;

[0076] Step 3.1.2), constructing a first feature extraction module to perform feature extraction on the image of H / 4×W / 4×48 obtained by the Patch Partition module to obtain a feature map of H / 4×W / 4×C;

[0077] The first feature extraction module consists of a Linear Embeding layer and two consecutive SwinTransformer Blocks in sequence;

[0078] The Swin Transformer Block is constructed as follows: Replace the standard multi-head self-attention module in Transformer with a module based on moving windows, keep other layers unchanged, and apply a LayerNorm layer before each MSA module and each MLP;

[0079] Step 3.1.3), construct a second feature extraction module to extract middle-level features, extract features from the feature map of H / 4×W / 4×C obtained by the first feature extraction module, and obtain a feature map of H / 8×W / 8×2C;

[0080] The second feature extraction module consists of a Patch Merging layer and six Swin Transformer Block layers in sequence;

[0081] Step 3.1.4), construct a third feature extraction module, extract features from the feature map of H / 8×W / 8×2C obtained in the second feature extraction stage, and obtain a feature map of H / 16×W / 16×4C;

[0082] The third feature extraction module consists of a Patch Merging layer and six Swin Transformer Block layers in sequence;

[0083] Step 3.1.5), construct a fourth feature extraction module, extract features from the feature map of H / 16×W / 16×4C obtained in the third feature extraction stage, and obtain a feature map of H / 32×W / 32×8C;

[0084] The fourth feature extraction module consists of a Patch Merging layer and two Swin Transformer Block layers in sequence;

[0085] The construction steps of the FPN module are as follows:

[0086] Step 3.2.1), construct four Conv2d(1×1, s1) modules to perform convolution operations on the feature maps obtained by the first to fourth feature extraction modules respectively, transform the feature map of H / 32×W / 32×8C obtained by the fourth feature extraction module into a feature map of H / 32×W / 32×C, transform the feature map of H / 16×W / 16×4C obtained by the third feature extraction module into a feature map of H / 16×W / 16×C, transform the feature map of H / 8×W / 8×2C obtained by the second feature extraction module into a feature map of H / 8×W / 8×C, and transform the feature map of H / 4×W / 4×C obtained by the first feature extraction module into a feature map of H / 4×W / 4×C;

[0087] Step 3.2.2), construct three upsampling and fusion modules to perform scale change operations on the feature maps obtained by the four Conv2d(1×1, s1) modules respectively and fuse the feature maps of the same scale. Transform the feature map of H / 32×W / 32×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 16×W / 16×C and fuse it with the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 8×W / 8×C and fuse it with the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 4×W / 4×C and fuse it with the feature map of H / 4×W / 4×C obtained by the Conv2d(1×1, s1) module;

[0088] Step 3.2.3), construct four Conv2d(3×3, s1) modules to perform convolution operations on the feature maps obtained by the three upsampling and fusion modules respectively, and the feature map of H / 32×W / 32×8C obtained by the Conv2d(1×1, s1) module. This convolution operation will not affect the scale of the feature map;

[0089] Step 3.2.4), construct a Maxpool(1×1, s2) module to perform pooling operations on the feature map of H / 32×W / 32×C in the feature maps obtained by the four Conv2d(3×3, s1) modules, and obtain a feature map of H / 64×W / 64×C;

[0090] Step 3.2.5), construct the Contact module to fuse and connect the H / 32×W / 32×8C feature map, H / 16×W / 16×C feature map, H / 8×W / 8×C feature map, H / 4×W / 4×C feature map obtained through four Conv2d(3×3,s1) modules and the H / 64×W / 64×C feature map obtained by performing pooling operation through the Maxpool(1×1,s2) module to obtain a fused connection feature map;

[0091] Step 3.2.6), construct the Fully Contected Layer to perform a fully connected operation on the fused connection feature map to obtain the two-dimensional detection box of the image target and the target classification result;

[0092] Step 4), as Figure 2 shown, perform point cloud extraction work on the two-dimensional detection box of the target in the front view and the two-dimensional detection box of the target in the bird's-eye view respectively to obtain the frustum point cloud region in the front view and the cylinder point cloud region in the bird's-eye view:

[0093] Step 4.1), based on the camera imaging principle, obtain the frustum region projected by the two-dimensional detection box of the target in the front view into the three-dimensional space according to the two-dimensional detection box of the target in the front view, and obtain the cylinder region projected by the two-dimensional detection box of the target in the bird's-eye view into the three-dimensional space according to the two-dimensional detection box of the target in the bird's-eye view;

[0094] Step 4.2), considering the internal parameters of the camera and the lidar and the rotation matrix and translation vector between the two, realize the coordinate transformation of the point cloud from the lidar coordinate system to the camera coordinate system; if the point cloud is located in the cone region or cylinder region projected by the two-dimensional detection box of the target into the three-dimensional space, it means that they can be projected into the two-dimensional bounding box of the target, and extract the information of this part of the point cloud for the subsequent regression of the three-dimensional bounding box of the target; through the point cloud coordinate transformation and extraction work, obtain the frustum point cloud space region corresponding to the front view and the cylinder point cloud space region corresponding to the bird's-eye view respectively;

[0095] Step 5), match the frustum point cloud space region corresponding to the front view of each target and the cylinder point cloud space region corresponding to the bird's-eye view of each target, and obtain the candidate point cloud region of the target by extracting the overlapping space region:

[0096] Compare the point cloud coordinates of the frustum point cloud space region of each target with the point cloud coordinates of the cylinder point cloud space region of each target. The point cloud coordinates that appear simultaneously in the frustum point cloud space region and the cylinder point cloud space region are the candidate point clouds, and all the candidate point clouds form the point cloud candidate region;

[0097] Step 6), after performing target point cloud segmentation on the candidate point cloud region, use Pointnet++ to extract point cloud features to obtain the target three-dimensional bounding regression box and target classification result in the spatial point cloud;

[0098] Step 7), by comprehensively considering the target classification results in the front view, the target classification results in the bird's-eye view, and the target classification results in the spatial point cloud, obtain the final classification result of the target.

[0099] In Step 1), lidar point cloud data is collected by a lidar, and a front view image of the road condition during vehicle driving is collected by an optical camera. The lidar point cloud and the front view image of the corresponding frame are obtained by intercepting the same timestamp.

[0100] The specific steps of Step 2 are as follows:

[0101] Step 2.1), calibrate the camera by the method of checkerboard calibration to obtain the internal and external parameters of the camera, and deduce the conversion between the vehicle body coordinate system, the camera coordinate system, and the pixel coordinate system through the coordinate relationship as follows:

[0102]

[0103] In the formula, is the pixel coordinate system, is the internal parameter matrix of the camera, is the vehicle body coordinate system, is the camera coordinate system, Z c is the distance between this point and the imaging plane in the camera axis direction, f x 、f y are the equivalent focal lengths of the camera in the x direction and the y direction respectively, u0 and v0 are the horizontal and vertical pixel coordinates of the image center, R c is the rotation matrix from the camera coordinate system to the vehicle body coordinate system, T c is the translation matrix from the camera coordinate system to the vehicle body coordinate system;

[0104] Step 2.2), combine the internal and external parameters of the camera to perform an inverse perspective transformation on the front view image, convert the front view image from the pixel coordinate system to the bird's-eye view in the world coordinate system, that is, convert it into a bird's-eye view, eliminate the interference of perspective distortion on the road condition information and distance error, and present the true world bird's-eye view characteristics. The mapping relationship between the pixel coordinate system of the perspective image and the bird's-eye view plane of the world coordinate system is as follows:

[0105]

[0106]

[0107] In the formula, X and Y are the horizontal and vertical coordinates of the perspective image in the bird's-eye view plane of the world coordinate system, u t 、vt are the horizontal and vertical coordinate pixels of the perspective image, θ is the angle between the optical axis of the camera on the vertical plane in the car and the horizontal plane, h is the distance from the camera to the ground, and d0 is the distance from the camera to the front end of the car;

[0108] The conversion relationship between the pixel coordinate system of the inverse perspective transformed image and the top-down plane of the world coordinate system is as follows:

[0109]

[0110]

[0111] In the formula, u n 、v n are the horizontal and vertical pixel coordinates of the inverse perspective transformed image, W IPM 、h IPM are the pixel width and height of the inverse perspective image, σ1 and σ2 are the actual distances of the unit pixels in the horizontal and vertical directions of the inverse perspective image in the horizontal direction of the world coordinate system, and d1 is the distance between the bottom of the camera field of view and the front end of the vehicle;

[0112] Step 2.3), after determining the correspondence between the pixels of the front view image and the radar points of the lidar point cloud data, combine the camera's internal parameters to solve the linear equations about the rotation matrix and translation vector, and calculate the rotation matrix and translation vector between the camera and the linear radar to achieve joint calibration of the camera and lidar:

[0113] Step 2.3.1), according to the perspective imaging model, use the external parameter matrix and the point cloud coordinates in the Cartesian coordinate system Multiply them to transform the point cloud into the camera coordinate system; project the point into the pixel coordinate system through the intrinsic parameter matrix to obtain the corresponding pixel point Complete the spatial alignment and registration of the lidar point cloud and the monocular camera image. The conversion relationship is:

[0114]

[0115] In the formula, is the laser radar coordinate system coordinate of the point, is the coordinate of the point in the camera coordinate system, is the pixel coordinate system coordinate of the point, K is the intrinsic parameter matrix of the camera, is the rotation matrix from the laser radar coordinate system to the camera coordinate system, The translation matrix from the LiDAR coordinate system to the camera coordinate system.

[0116] In step 7), when making comprehensive considerations, the category credibility formula is introduced: P f =0.4P1+0.4P2+0.2P3;

[0117] Wherein, P f is the credibility;

[0118] is the judgment of the object category by the FPN Swin Transformer under the front view, and p 1a and p 1b and p 1c are the probability values that the classification results judged by the FPN Swin Transformer under the front view are category a, category b, and other category c respectively;

[0119] is the judgment of the object category by the FPN Swin Transformer under the bird's-eye view, and p 2a and p 2b and p 2c are the probability values that the classification results judged by the FPN Swin Transformer under the bird's-eye view are category a, category b, and other category c respectively;

[0120] is the judgment of the object category by Pointnet++ under the spatial point cloud, and p 3a and p 3b and p 3c are the probability values that the classification results judged by the FPN Swin Transformer under the bird's-eye view are category a, category b, and other category c respectively.

[0121] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0122] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An autonomous driving three-dimensional object detection method based on FPN Swin Transformer and Pointnet++, characterized in that, It includes the following steps: Step 1), obtaining the front view image and lidar point cloud data of the road condition during vehicle driving; Step 2), performing inverse perspective transformation on the front view image to obtain a bird's-eye view image of the road condition, projecting the lidar point cloud onto the front view image and the bird's-eye view image respectively, and obtaining the front view image and the bird's-eye view image with fused point cloud features; Step 3), respectively performing feature extraction on the front view image and the bird's-eye view image with fused point cloud features through an FPN Swin Transformer to obtain the target two-dimensional detection box and target classification result in the front view perspective, and the target two-dimensional detection box and target classification result in the bird's-eye view perspective; The FPN Swin Transformer includes a Swin Transformer module and an FPN module; The Swin Transformer module includes a Patch Partition module and first to fourth feature extraction modules; The construction steps of the Swin Transformer module are as follows: Step 3.1.1), constructing a Patch Partition module to block the image with fused point cloud features, and blocking the input image with fused point cloud features of size H×W×3 into four images of H / 4×W / 4×48; Step 3.1.2), constructing a first feature extraction module to perform feature extraction on the image of H / 4×W / 4×48 obtained by the Patch Partition module, and obtaining a feature map of H / 4×W / 4×C; The first feature extraction module is sequentially composed of a Linear Embeding layer and 2 consecutive Swin Transformer Blocks; The construction method of the Swin Transformer Block is as follows: replacing the standard multi-head self-attention module in the Transformer with a module based on a moving window, keeping other layers unchanged, and applying a LayerNorm layer before each MSA module and each MLP; Step 3.1.3), constructing a second feature extraction module for extracting middle-level features, performing feature extraction on the feature map of H / 4×W / 4×C obtained by the first feature extraction module, and obtaining a feature map of H / 8×W / 8×2C; The second feature extraction module is sequentially composed of a Patch Merging layer and six Swin Transformer Block layers; Step 3.1.4), constructing a third feature extraction module to perform feature extraction on the feature map of H / 8×W / 8×2C obtained in the second feature extraction stage, and obtaining a feature map of H / 16×W / 16×4C; The third feature extraction module is sequentially composed of a Patch Merging layer and six Swin Transformer Block layers; Step 3.1.5), construct a fourth feature extraction module to extract features from the feature map of H / 16×W / 16×4C obtained in the third feature extraction stage, and obtain a feature map of H / 32×W / 32×8C; The fourth feature extraction module consists of a Patch Merging layer and two Swin Transformer Block layers in sequence; The construction steps of the FPN module are as follows: Step 3.2.1), construct four Conv2d(1×1, s1) modules to perform convolution operations on the feature maps obtained by the first to fourth feature extraction modules respectively. Transform the feature map of H / 32×W / 32×8C obtained by the fourth feature extraction module into a feature map of H / 32×W / 32×C, transform the feature map of H / 16×W / 16×4C obtained by the third feature extraction module into a feature map of H / 16×W / 16×C, transform the feature map of H / 8×W / 8×2C obtained by the second feature extraction module into a feature map of H / 8×W / 8×C, and transform the feature map of H / 4×W / 4×C obtained by the first feature extraction module into a feature map of H / 4×W / 4×C; Step 3.2.2), construct three upsampling and fusion modules to perform scale change operations on the feature maps obtained by the four Conv2d(1×1, s1) modules respectively and fuse the feature maps of the same scale. Transform the feature map of H / 32×W / 32×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 16×W / 16×C and fuse it with the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 16×W / 16×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 8×W / 8×C and fuse it with the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module. Transform the feature map of H / 8×W / 8×C obtained by the Conv2d(1×1, s1) module into a feature map of H / 4×W / 4×C and fuse it with the feature map of H / 4×W / 4×C obtained by the Conv2d(1×1, s1) module; Step 3.2.3), construct four Conv2d(3×3, s1) modules to perform convolution operations on the feature maps obtained by the three upsampling and fusion modules and the feature map of H / 32×W / 32×8C obtained by the Conv2d(1×1, s1) module. This convolution operation will not affect the scale of the feature map; Step 3.2.4), construct a Maxpool(1×1, s2) module to perform pooling operations on the feature map of H / 32×W / 32×C in the feature maps obtained by the four Conv2d(3×3, s1) modules, and obtain a feature map of H / 64×W / 64×C; Step 3.2.5), construct the Contact module to fuse and connect the H / 32×W / 32×8C feature map, H / 16×W / 16×C feature map, H / 8×W / 8×C feature map, H / 4×W / 4×C feature map obtained by four Conv2d(3×3,s1) modules and the H / 64×W / 64×C feature map obtained by pooling operation through the Maxpool(1×1,s2) module to obtain the fused connection feature map; Step 3.2.6), construct the Fully Contected Layer to perform a fully connected operation on the fused connection feature map to obtain the two-dimensional detection frame of the image target and the target classification result; Step 4), perform point cloud extraction work on the two-dimensional detection frame of the target in the front view and the two-dimensional detection frame of the target in the bird's-eye view respectively to obtain the frustum point cloud region in the front view and the cylinder point cloud region in the bird's-eye view: Step 4.1), based on the camera imaging principle, obtain the frustum region projected by the two-dimensional detection frame of the target in the front view into the three-dimensional space according to the two-dimensional detection frame of the target in the front view, and obtain the cylinder region projected by the two-dimensional detection frame of the target in the bird's-eye view into the three-dimensional space according to the two-dimensional detection frame of the target in the bird's-eye view; Step 4.2), consider the internal parameters of the camera and the lidar and the rotation matrix and translation vector between the two to realize the coordinate transformation of the point cloud from the lidar coordinate system to the camera coordinate system; if the point cloud is located in the cone region or cylinder region projected by the two-dimensional detection frame of the target into the three-dimensional space, it means that they can be projected into the two-dimensional bounding box of the target, and extract the information of this part of the point cloud for subsequent regression of the three-dimensional bounding box of the target; through the point cloud coordinate transformation and extraction work, the frustum point cloud space region corresponding to the front view and the cylinder point cloud space region corresponding to the bird's-eye view are obtained respectively; Step 5), match the frustum point cloud space region corresponding to the front view of each target and the cylinder point cloud space region corresponding to the bird's-eye view of each target, and obtain the candidate point cloud region of the target by extracting the overlapping space region: Compare the point cloud coordinates of the frustum point cloud space region of each target with the point cloud coordinates of the cylinder point cloud space region. The point cloud coordinates that appear simultaneously in the frustum point cloud space region and the cylinder point cloud space region are the candidate point clouds, and all candidate point clouds form the point cloud candidate region; Step 6), perform target point cloud segmentation on the candidate point cloud region and then use Pointnet++ to extract point cloud features to obtain the three-dimensional boundary regression box of the target under the spatial point cloud and the target classification result; Step 7), obtain the final classification result of the target by comprehensively considering the target classification result in the front view, the target classification result in the bird's-eye view and the target classification result under the spatial point cloud.

2. The 3D object detection method for autonomous driving based on FPN Swin Transformer and Pointnet++, according to claim 1, is characterized in that In Step 1), the lidar point cloud data is collected by the lidar, the front view image of the road condition during the vehicle driving is collected by the optical camera, and the corresponding frame of lidar point cloud and front view image are obtained by intercepting the same timestamp.

3. The 3D object detection method for autonomous driving based on FPN Swin Transformer and Pointnet++, according to claim 2, wherein The specific steps of the said Step 2) are as follows: Step 2.1), calibrate the camera by the method of checkerboard calibration to obtain the internal and external parameters of the camera, and deduce the conversions among the vehicle body coordinate system, the camera coordinate system, and the pixel coordinate system through the coordinate relationship as follows: wherein, is the pixel coordinate system, is the camera intrinsic matrix, is the vehicle body coordinate system, is the camera coordinate system, Z c is the distance between the point and the imaging plane in the camera axis direction, f x and f y are the equivalent focal lengths of the camera in the x - direction and y - direction respectively, u0 and v0 are the horizontal and vertical pixel coordinates of the image center, R c is the rotation matrix from the camera coordinate system to the vehicle body coordinate system, T c is the translation matrix from the camera coordinate system to the vehicle body coordinate system; Step 2.2), perform inverse perspective transformation on the front view image by combining the internal and external parameters of the camera, and convert the front view image from the pixel coordinate system to the top-down view of the world coordinate system, that is, convert it into a bird's-eye view, eliminate the interference of perspective deformation on the road condition information and the distance error, and present the top-down characteristics of the real world. The mapping relationship between the pixel coordinate system of the perspective image and the top-down plane of the world coordinate system is as follows: Wherein, X and Y are respectively the horizontal and vertical coordinates of the perspective image in the top view plane of the world coordinate system, and u t , v t are respectively the horizontal and vertical coordinate pixels of the perspective image, θ is the angle between the optical axis of the camera and the horizontal plane on the vertical plane of the vehicle, h is the distance from the camera to the ground, and d0 is the distance from the camera to the front end of the vehicle; The conversion relationship between the pixel coordinate system of the inverse perspective transformation image and the top-down plane of the world coordinate system is as follows: where u n and v n are the horizontal and vertical pixel coordinates of the inverse perspective transformation image respectively, W IPM and h IPM are the pixel width and height of the inverse perspective image respectively, σ1 and σ2 are the actual distances of a unit pixel in the horizontal and vertical directions of the inverse perspective image in the horizontal direction of the world coordinate system, and d1 is the distance between the lowest part of the camera's field of view and the front end of the vehicle; Step 2.3), after determining the corresponding relationship between the pixels of the front view image and the radar points of the lidar point cloud data, combine the internal parameters of the camera, solve the linear equations about the rotation matrix and the translation vector, and obtain the rotation matrix and the translation vector between the camera and the lidar, so as to realize the joint calibration of the camera and the lidar: Step 2.3.1), according to the perspective imaging model, multiply the external parameter matrix by the point cloud coordinates in the Cartesian coordinate system to transform the point cloud into the camera coordinate system; project the point through the internal parameter matrix onto the pixel coordinate system to obtain the corresponding pixel point Complete the spatial alignment and registration of the lidar point cloud and the monocular camera image, and the conversion relationship is: Wherein, is the coordinate of the point in the lidar coordinate system, is the coordinate of the point in the camera coordinate system, is the coordinate of the point in the pixel coordinate system, K is the internal parameter matrix of the camera, is the rotation matrix from the lidar coordinate system to the camera coordinate system, is the translation matrix from the lidar coordinate system to the camera coordinate system.

4. The 3D object detection method for autonomous driving based on FPN Swin Transformer and Pointnet++, as claimed in claim 1, wherein In step 7), when making a comprehensive consideration, introduce the category credibility formula: P f = 0.4P1 + 0.4P2 + 0.2P3; where P f is the credibility; For the judgment of object categories by the FPN Swin Transformer in the front view, p 1a , p 1b , p 1c are the probability values that the classification results judged by the FPN Swin Transformer in the front view are category a, category b, and other category c respectively; For the judgment of object categories by the FPN Swin Transformer from an aerial view, p 2a , p 2b , p 2c are the probability values that the classification results judged by the FPN Swin Transformer from an aerial view are category a, category b, and other c respectively; For the judgment of object categories by Pointnet++ under spatial point clouds, p 3a , p 3b , p 3c are the probability values that the classification results judged by FPN Swin Transformer from the bird's-eye view are category a, category b, and other c respectively.

Citation Information

Patent Citations

  • 3D target detection method based on vehicle end point cloud and image fusion

    CN114387202A

  • Three-dimensional target detection method and device based on multi-sensor information fusion

    CN110929692A

  • RGB-D multi-modal feature fusion 3D target detection method

    CN113408584A