3D object detection method based on cylindrical projection and polar coordinates

Through cylindrical projection and polar coordinate conversion technology, combined with deep structure and multi-scale feature fusion network, the problem of 3D spatial information extraction in image data in the field of autonomous driving is solved, and more accurate and robust 3D object detection is achieved.

CN120220133APending Publication Date: 2025-06-27安徽中科星驰自动驾驶技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335030.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the field of autonomous driving, image-based 3D object detection technology is difficult to accurately extract 3D spatial information from 2D image data, resulting in poor performance.

Method used

Cylindrical projection is used to project the original image onto the cylinder and convert it into polar coordinates. Combining deep structure and multi-scale feature fusion networks, multi-scale feature maps are generated and feature stitching is performed, and 3D object detection is finally performed through polar coordinate BEV network.

Benefits of technology

Through cylinder projection and polar coordinate conversion, distortion in the long distance part is reduced, the overall structure of the image is maintained, and the accuracy and robustness of object detection are improved, especially when dealing with small objects and rotation or scale changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220133A_ABST
    Figure CN120220133A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D object detection method based on cylindrical projection and polar coordinates, which belongs to the technical field of automatic driving, and is technically characterized by comprising the following steps of: 1, projecting an original image onto a cylindrical surface by adopting cylindrical projection to generate a plane view; 2, preprocessing the image, and carrying out random image cutting on the input image in the transverse direction and the longitudinal direction; 3, extracting high-level semantic features, generating multi-scale feature maps, and fusing the multi-scale feature maps; 4, performing view angle conversion, and outputting features after FPN; 5, a polar coordinate BEV network is established, and a central point regression mode is used; step 6, polar coordinate conversion; step 7, loss calculation: calculating a difference between a model output result and a truth value result; and 8, post-processing: converting the center point, width and height information into a final object prediction frame, and the method has the advantages of reducing the attention to a region far away from the center, and enabling computing resources to be more concentrated on an interested region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and specifically relates to a 3D object detection method based on cylindrical projection and polar coordinates. Background Art

[0002] In the field of autonomous driving, image-based 3D object detection technology mainly relies on visual information obtained from cameras, and infers the three-dimensional spatial position of target objects through computer vision algorithms. Compared with sensors such as LiDAR, image data provides rich visual information, but there are many challenges in extracting accurate 3D spatial information from 2D image data.

[0003] Early attempts mainly solved this problem from the perspective of monocular 3D object detection, where monocular 3D object detection was first performed on each view, and then the predictions of all views were fused through post-processing. Although feasible, this detection scheme ignores the information between views, resulting in unsatisfactory performance.

[0004] Many studies use Bird's-Eye-View (BEV) to integrate cross-view information, eliminating inefficient post-processing fusion, and achieving significant progress in both detection performance and efficiency. In particular, the LSS-based paradigm uses the Lift-SplatShoot (LSS) mechanism to construct an explicit dense BEV representation and has become mainstream. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the purpose of the embodiments of the present invention is to provide a 3D object detection method based on cylindrical projection and polar coordinates to solve the problems in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A 3D object detection method based on cylindrical projection and polar coordinates, comprising the following steps:

[0008] Step 1: Project the original image onto a cylinder using cylindrical projection, and then unfold the cylinder into a plane to generate a planar view;

[0009] Step 2: Preprocess the image. For the input image, perform random image cropping horizontally and vertically, and scale it to the size of the input image 256x704 according to a ratio;

[0010] Step 3: Use the deep structure resnet50 as the backbone network to extract high-level semantic features, transfer the high-level semantic features to a multi-scale feature fusion network for further processing, and generate and fuse multi-scale feature maps;

[0011] Step 4: Perspective transformation. The features after the FPN output are fed into a feed-forward neural network composed of multiple fully connected layers. Through multiple hidden layers and non-linear activation functions, a complex non-linear mapping relationship is learned to predict the depth in polar coordinates, and then feature concatenation is performed.

[0012] Step 5: Polar coordinate BEV network. Use CustomResNet as the backbone network of the bev encoder, adjust the number of layers, activation functions, number of channels, etc. The features are input into FPN_LSS for learning and fusion of multi-scale features. Use the method of center point regression as the detection head for 3D objects to detect the position, offset, direction, and category of the center point.

[0013] Step 6: Polar coordinate transformation. During training, the original ground truth data needs to be transformed from Cartesian coordinates to polar coordinates. Secondly, during inference, the box information predicted by the network needs to be transformed back from polar coordinates to Cartesian coordinates.

[0014] Step 7: Loss calculation. Calculate the difference between the model output result and the ground truth result as an error term, which is then fed back to the model for gradient calculation to update the model's parameters.

[0015] Step 8: Post-processing. Through the decoder, the center point, width, and height information are converted into the final object prediction box.

[0016] As a further solution of the present invention, in Step 1, for the projection, let P(X, Y) be the point coordinates in the original image, and let p(u, v) be the corresponding point on the cylindrical projection plane. The radius of the cylinder is f, that is, the distance from the camera to the cylinder, which is also called the focal length.

[0017] First, calculate the cylindrical projection process of x. Given the horizontal field of view as fov, with the image center as the origin, the θ angles corresponding to the first column and the last column are fov / 2 and -fov / 2 respectively. Arbitrarily take a point x on the image, and the u coordinate corresponding to it on the cylinder is w' / 2 + f×(θ), where w' = f×fov; and the θ angle is expressed as actan((x - w / 2) / f).

[0018] Calculate the cylindrical projection process of y. According to the principle of equilateral triangle, the equation can be obtained: (y - H / 2) / (Y - H / 2) = f / (f(y)), where f(y) = sqrt((x - w / 2)^2 + f^2), and f and f(y) are the distances from the center of the cylinder to v and Y respectively. Thus, the final projection formula is derived as:

[0019]

[0020] Obtain the projection process from P to p.

[0021] As a further solution of the present invention, in step four, the depth in polar coordinates is predicted:

[0022] h = f(Wx + b);

[0023] where W is the weight matrix, b is the bias term, and x is the output or input data of the previous layer, and is the activation function.

[0024] As a further solution of the present invention, in step four, the feature splicing includes combining the extrinsic parameters of the camera and the FOV, calculating the angular range corresponding to each camera, then performing a splicing operation on the features after depth estimation, and filling the features of each camera into the corresponding positions in the clockwise direction. The features are sequentially connected starting from the x-axis direction and spliced to 360 degrees.

[0025] As a further solution of the present invention, in step five, it includes center point detection and center point and offset regression. By regressing the center point of the target, and then expanding or offsetting from the center point to determine the complete position of the object. Each object will have a center point, and the network locates the object by predicting the position of this center point. A regression network is used to predict the center point position of each target and the offset of this point.

[0026] As a further solution of the present invention, in step five, it also includes center point and offset regression. A regression network is used to predict the center point position of each target and the offset of this point. This method makes the positioning of the target more accurate. Especially when dealing with small objects, there may be certain errors in traditional bounding box regression. The center point regression method can usually better capture the geometric center of the object.

[0027] As a further solution of the present invention, the conversion from Cartesian coordinates to polar coordinates in step six includes assuming that the coordinates of a point in the Cartesian coordinate system are (x, y), and to convert it to polar coordinates (r, θ), where r is the distance from the point to the origin, and θ is the angle of the point relative to the x-axis. The conversion formula is as follows:

[0028] ;

[0029] ;

[0030] where r is calculated by the Pythagorean theorem, that is, the straight-line distance from the origin to the point (x, y), and θ is the angle between this point and the x-axis, and then the arctangent function is used to calculate;

[0031] The conversion from polar coordinates to Cartesian coordinates, the formula is as follows:

[0032] ;

[0033] 。

[0034] As a further solution of the present invention, the loss calculation in step seven includes center point regression loss, and by combining Focal Loss and Gaussian Distribution as the loss of center point regression;

[0035] ;

[0036] Where p(t) is the predicted probability that the sample belongs to the correct class, representing the predicted probability of the model for the target class, N(p(t)|μ, σ(2)) is to weight the predicted probability p(t) using Gaussian distribution, similar to adjusting the weight of each sample, σ is the standard deviation of the Gaussian distribution, controlling the size of the attention area. The smaller the standard deviation, the smaller the attention area. α and γ are the same adjustment factors as FocalLoss, which are used to balance class imbalance and adjust the weight of easy-to-classify samples respectively;

[0037] The loss calculation also includes offset regression loss, by calculating the absolute error between the predicted value and the true value of each sample and taking its average value.

[0038] ;

[0039] Where is the predicted value of the i-th sample, y i is the true value of the i-th sample, is the absolute error of the i-th sample.

[0040] To sum up, the embodiments of the present invention have the following beneficial effects compared with the prior art:

[0041] 1. After cylindrical projection, the closer the object in the image is to the center of the image, the smaller its deformation, while the area far from the center will have a larger distortion. By converting to polar coordinates after cylindrical projection, the information at the far end of the image can be compressed into a smaller range, thereby reducing the distortion of the distant part and better maintaining the overall structure of the image. The detection and feature extraction of objects become simpler. After cylindrical projection of the image and then converting to polar coordinates, the edges and symmetry of circular objects may be more obvious. Especially when the image has rotation or scale changes, the model in polar coordinates can perform more robustly. Converting to polar coordinates after cylindrical projection can magnify the central area of the image, which is very useful because many useful information often concentrates in the central area of the image. This way can make the computing resources more concentrated on the area of interest by reducing the attention to the area far from the center.

[0042] 2. Converting multi-view image features to polar coordinate BEV representation instead of Cartesian BEV representation can better adapt to the distribution of image information and is easy to maintain view symmetry;

[0043] In the range of 0 - 30 m, the accuracy of the polar coordinate BEV representation has a significant improvement, indicating that the polar representation is more conducive to the perception of nearby areas. When the vehicle turns at a large angle, the direction of the camera changes significantly. In this case, in order to accurately detect surrounding objects, the autonomous driving system needs to be robust to the change of azimuth angle. The polar coordinate BEV representation has superior generalization ability among different models.

[0044] 3. By splicing image features, the features are transformed to the polar coordinate BEV perspective. Compared with other perspective transformation methods, this method does not use any complex operations or time-consuming operators such as interpolation sampling, which improves the overall running speed of the model.

[0045] To more clearly elaborate the structural features and functions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Description of the Drawings

[0046] Figure 1 It is the algorithm flowchart of the invention embodiment.

[0047] Figure 2 It is the cylindrical projection diagram in the invention embodiment. Detailed Embodiment

[0048] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] The following describes the specific implementation of the present invention in detail with specific embodiments.

[0050] In one embodiment, a 3D object detection method based on cylindrical projection and polar coordinates, see Figures 1 to 2 , includes the following steps:

[0051] Step 1: Use cylindrical projection to project the original image onto a cylinder, and then unfold the cylinder into a plane to generate a planar view;

[0052] Step 2: Preprocess the image. For the input image, perform random image cropping horizontally and vertically, and scale it to the input image size of 256x704 according to a ratio;

[0053] Step 3: By using the deep structure resnet50 as the backbone network, extract high-level semantic features, transfer the high-level semantic features to the multi-scale feature fusion network for further processing, generate multi-scale feature maps and fuse them;

[0054] Step 4: Perspective transformation. The features after the FPN output are fed into a feed-forward neural network composed of multiple fully connected layers. Through multiple hidden layers and non-linear activation functions, learn complex non-linear mapping relationships, predict the depth in polar coordinates, and then perform feature splicing;

[0055] Step 5: Polar coordinate BEV network. Use CustomResNet as the backbone network of the bev encoder, adjust the number of layers, activation functions, number of channels, etc. The features are input into FPN_LSS, perform learning and fusion of multi-scale features, and use the method of center point regression as the detection head for 3D objects to detect the position, offset, direction and category of the center point;

[0056] Step 6: Polar coordinate conversion. During training, the original ground truth data needs to be converted from Cartesian coordinates to polar coordinates. Secondly, during inference, the box information predicted by the network is converted back from polar coordinates to Cartesian coordinates;

[0057] Step 7: Loss calculation. Calculate the gap between the model output result and the ground truth result, and use it as an error term to be sent back to the model for gradient calculation to update the model parameters;

[0058] Step 8: Post-processing. Through the decoder, convert the center point, width and height information into the final object prediction box.

[0059] Further, see Figures 1 to 2 , in Step 1, let P(X, Y) be the point coordinates in the original image, and let p(u, v) be the corresponding point on the cylindrical projection plane. The radius of the cylinder is f, that is, the distance from the camera to the cylinder, also known as the focal length;

[0060] First, calculate the cylindrical projection process of x. Given the horizontal field of view angle fov, with the image center as the origin, the θ angles corresponding to the first column and the last column are fov / 2 and fov / 2 respectively; randomly take a point x on the image, and the u coordinate corresponding to it on the cylinder is w' / 2 + f×(θ), where w' = f×fov; and the θ angle is expressed as actan((x - w / 2) / f);

[0061] Calculate the cylindrical projection process of y. According to the equilateral triangle principle, the equation can be obtained: (y - H / 2) / (Y - H / 2) = f / (f(y)), where f(y) = sqrt((x - w / 2)(2)+f(2)), and f and f(y) are the distances from the cylinder center to v and Y respectively. Finally, the projection formula is derived as:

[0062]

[0063] Obtain the projection process from P to p.

[0064] Furthermore, refer to Figures 1 to 2 , predicting the depth in polar coordinates in step four:

[0065] h = f(Wx + b);

[0066] where W is the weight matrix, b is the bias term, x is the output or input data of the previous layer, and is the activation function.

[0067] Furthermore, refer to Figures 1 to 2 , the feature concatenation in step four includes combining the extrinsic camera parameters and FOV, calculating the angular range corresponding to each camera, then performing a concatenation operation on the features after depth estimation. In the clockwise direction, fill the features of each camera into the corresponding positions, and the features are connected sequentially starting from the x-axis direction and concatenated to 360 degrees.

[0068] Furthermore, refer to Figures 1 to 2 , step five includes center point detection and center point and offset regression. By regressing the center point of the target, and then expanding or offsetting from the center point to determine the complete position of the object. Each object will have a center point, and the network locates the object by predicting the position of this center point, and predicts the center point position and the offset of this point of each target through a regression network.

[0069] Furthermore, refer to Figures 1 to 2 , step five also includes center point and offset regression. Predict the center point position and the offset of this point of each target through a regression network. This method makes the positioning of the target more accurate. Especially when dealing with small objects, there may be certain errors in traditional bounding box regression, and the center point regression method can usually better capture the geometric center of the object.

[0070] Furthermore, refer to Figures 1 to 2 , the conversion from Cartesian coordinates to polar coordinates in step six includes assuming that the coordinates of a point in the Cartesian coordinate system are (x, y), and to convert it to polar coordinates (r, θ), where r is the distance from the point to the origin, and θ is the angle of the point relative to the x-axis. The conversion formula is as follows:

[0071] ;

[0072] ;

[0073] where r is calculated by the Pythagorean theorem, that is, the straight-line distance from the origin to the point (x, y), and θ is the angle between this point and the x-axis, and then the arctangent function is used for calculation;

[0074] The conversion from polar coordinates to Cartesian coordinates is as follows:

[0075] ;

[0076] .

[0077] Furthermore, referring to Figures 1 to 2 , the loss calculation in step 7 includes the center point regression loss, which combines Focal Loss and Gaussian Distribution as the loss for center point regression;

[0078] ;

[0079] where p(t) is the predicted probability that the sample belongs to the correct class, representing the predicted probability of the model for the target class, N(p(t)|μ, σ(2)) is to weight the predicted probability p(t) using the Gaussian distribution, similar to adjusting the weight of each sample, σ is the standard deviation of the Gaussian distribution, controlling the size of the region of interest. The smaller the standard deviation, the smaller the region of interest, and α and γ are the same adjustment factors as FocalLoss, which are used to balance class imbalance and adjust the weight of easy-to-classify samples respectively;

[0080] The loss calculation also includes the offset regression loss, which calculates the absolute error between the predicted value and the true value of each sample and takes its average.

[0081] ;

[0082] where is the predicted value of the i-th sample, y i is the true value of the i-th sample, is the absolute error of the i-th sample.

[0083] In this embodiment, step S1 includes the following steps:

[0084] S1: Cylindrical projection: Cylindrical projection is a method of mapping a three-dimensional scene onto a two-dimensional plane. In the present invention, cylindrical projection is used to project the original image onto a cylinder and then unfold the cylinder into a plane, thereby generating a relatively natural planar view. The following is the mathematical derivation and basic principle of cylindrical projection.

[0085] Let P(X, Y) be the point coordinates in the original image;

[0086] Let p(u, v) be the corresponding point on the cylindrical projection plane;

[0087] The radius of the cylinder is f, which is the distance from the camera to the cylinder and is also called the focal length;

[0088] First, calculate the cylindrical projection process of x. Refer to Figure 2 , given that the horizontal field of view angle is fov, with the image center as the origin, the θ angles corresponding to the first column and the last column are fov / 2 and fov / 2 respectively; randomly take a point x on the image, and the u coordinate on the cylinder corresponding to it is w' / 2 + f×(θ), where w' = f×fov; and the θ angle is expressed as actan((x - w / 2) / f).

[0089] Then, calculate the cylindrical projection process of y. Refer to Figure 2 , according to the equilateral triangle principle, the equation can be obtained: (y - H / 2) / (Y - H / 2) = f / (f(y)), where f(y)=sqrt((x - w / 2)(2)+f(2)), and f and f(y) are the distances from the cylinder center to v and Y; thus, the final projection formula is derived as:

[0090]

[0091] Thus, the projection process from P to p is obtained.

[0092] S2: Image preprocessing: For the input image, first perform randomly cropped images horizontally and vertically, and then scale them to the size of the input image 256x704 according to a ratio.

[0093] S3: Backbone network: By using the deep structure resnet50 as the backbone network to extract rich high-level semantic features, and then transfer these features to the multi-scale feature fusion network (FPN) for further processing, generate multi-scale feature maps and fuse them, so as to improve the detection performance of the model for targets of different sizes.

[0094] S4: Viewpoint conversion:

[0095] S4.1: Depth prediction: The features after the FPN output are further fed into a feedforward neural network (mlp) composed of multiple fully connected layers, learn complex non-linear mapping relationships through multiple hidden layers and non-linear activation functions, and predict the depth in polar coordinates.

[0096] h = f(Wx + b);

[0097] Where: W is the weight matrix, b is the bias term, and x is the output or input data of the previous layer. is the activation function. S4.2: Feature concatenation: First, it is necessary to calculate the angular range corresponding to each camera by combining the extrinsic parameters of the camera and the FOV, and then perform a concatenation operation on the features after depth estimation. Fill the features of each camera into the corresponding positions in the clockwise direction, and connect the features of each camera in sequence starting from the x-axis direction until it is concatenated to 360 degrees.

[0098] S5: Polar coordinate BEV network: Introduce CustomResNet as the backbone network of the bev encoder, and adjust the number of layers, activation function, number of channels, etc. to adapt to the overall network structure and balance the model accuracy and time consumption. Next, the feature is input into FPN_LSS for multi-scale feature learning and fusion. Then, use the method of center point regression as the detection head for 3D objects. The key ideas include:

[0099] S5.1: Center point detection: Determine the complete position of the object by regressing the center point of the target and then expanding or offsetting from the center point. Each object will have a center point, and the network locates the object by predicting the position of this center point.

[0100] S5.2: Center point and offset regression: Predict the center point position and the offset of this point for each target through a regression network. This method makes the object localization more accurate. Especially when dealing with small objects, traditional bounding box regression may have certain errors, while the center point regression method can usually better capture the geometric center of the object.

[0101] S6: Polar coordinate conversion: Polar coordinate conversion involves two parts. First, during training, the original ground truth data needs to be converted from Cartesian coordinates to polar coordinates. Second, during inference, the box information predicted by the network needs to be converted from polar coordinates back to Cartesian coordinates.

[0102] S6.1: Cartesian coordinate to polar coordinate conversion. Assume that in the Cartesian coordinate system, the coordinates of a point are (x, y), and to convert it to polar coordinates (r, θ), where: r is the distance from the point to the origin (polar radius, radius), and θ is the angle of the point relative to the x-axis (polar angle, angle). The conversion formula is as follows:

[0103] ;

[0104] ;

[0105] where: r is calculated by the Pythagorean theorem, that is, the straight-line distance from the origin to the point (x, y). θ is the angle between the point and the x-axis, and the arctangent function is used to calculate it.

[0106] S6.2: Polar coordinate to Cartesian coordinate conversion, the formula is as follows:

[0107] ;

[0108] 。

[0109] S7: Loss calculation:

[0110] S7.1: Center point regression loss. By combining Focal Loss and Gaussian Distribution as the loss for center point regression, Focal Loss reduces the loss weight of easy-to-classify samples by introducing a modulation factor, emphasizing the training of hard-to-classify samples. Then, on this basis, the idea of Gaussian distribution is introduced, making the loss function pay more attention to the area close to the target center, thus providing better performance when detecting small objects.

[0111] ;

[0112] Where: p(t) is the predicted probability that the sample belongs to the correct class, which, like Focal Loss, represents the predicted probability of the model for the target class. N(p(t)|μ, σ(2)) weights the predicted probability p(t) using Gaussian distribution, similar to adjusting the weight of each sample. σ is the standard deviation of the Gaussian distribution, controlling the size of the attention area. The smaller the standard deviation, the smaller the attention area (i.e., closer to the target center). α and γ are the same modulation factors as in Focal Loss, used to balance class imbalance and adjust the weight of easy-to-classify samples respectively.

[0113] S7.2: Offset regression loss: By calculating the absolute error between the predicted value and the true value of each sample and taking its average.

[0114] ;

[0115] Where is the predicted value of the i-th sample, y i is the true value of the i-th sample, is the absolute error of the i-th sample;

[0116] S8: Post-processing: Through the bounding box decoder, the center point, width, height, and some other information are converted into the final bounding box, specifically:

[0117] Regress the coordinates of the center point (i.e., the center position of the target bounding box); regress the width and height of the bounding box.

[0118] During post-processing, the regressed center point and size information are restored to the four coordinates of the bounding box (the coordinates of the upper left corner and the lower right corner).

[0119] The working principle of the present invention is as follows:

[0120] The grid distribution represented by the polar coordinate BEV is consistent with the distribution of the image information, with the image information being dense near and sparse far away. Compared with the Cartesian BEV representation, the polar BEV representation can naturally capture the fine-grained information in the nearby area and also reduce the computational redundancy in the distant area.

[0121] The polar coordinate BEV representation can conveniently maintain the view symmetry of the surround cameras. Assuming that different cameras capture the imaging of the same object, their features are approximately parallel in the polar BEV representation, so only conventional two-dimensional convolution operations are needed to learn azimuth-equivalent object features.

[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A 3D object detection method based on cylindrical projection and polar coordinates, characterized in that: The following steps are involved: Step 1: Use cylindrical projection to project the original image onto a cylinder, and then unfold the cylinder into a plane to generate a plane view; Step 2: Preprocess the image. For the input image, randomly crop the image in the horizontal and vertical directions and scale it to the input image size of 256x704. Step 3: By using the deep structure resnet50 as the backbone network, high-level semantic features are extracted, and the high-level semantic features are passed to the multi-scale feature fusion network for further processing to generate and fuse multi-scale feature maps; Step 4: Perspective conversion, the features after FPN output are sent to a feedforward neural network composed of multiple fully connected layers. Through multiple hidden layers and nonlinear activation functions, complex nonlinear mapping relationships are learned to predict the depth in polar coordinates, and then feature splicing is performed; Step 5: Polar coordinate BEV network, using CustomResNet as the backbone network of bev encoder, adjusting the number of layers, activation function and number of channels, inputting features into FPN_LSS, learning and fusing multi-scale features, using center point regression as the detection head of 3D targets, detecting the position, offset, direction and category of the center point; Step 6: Polar coordinate conversion: during training, the original true value data needs to be converted from Cartesian coordinates to polar coordinates. Secondly, during inference, the box information predicted by the network is converted from polar coordinates back to Cartesian coordinates. Step 7: Loss calculation: calculate the difference between the model output and the true value, and pass it back to the model as the error term for gradient calculation to update the model parameters. Step 8: Post-processing: Through the decoder, the center point, width and height information are converted into the final object prediction box.

2. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 1, characterized in that: In the projection of step 1, let P (X, Y) be the coordinates of the point in the original image, let p (u, v) be the corresponding point on the cylindrical projection plane, and the radius of the cylinder is f, that is, the distance from the camera to the cylinder, also called the focal length; First, calculate the cylindrical projection process of x. The horizontal field of view angle is known to be fov. Taking the center of the image as the origin, the first and last columns correspond to the angles θ of fov / 2 and fov / 2 respectively. Take a random point x on the image and the corresponding u coordinate on the cylinder is , ; where the angle θ is expressed as actan((xw / 2) / f); The cylindrical projection process of y can be calculated according to the equilateral triangle principle: (yH / 2) / (YH / 2)= f / (f(y)), where f(y) = sqrt((xw / 2)(2)+f(2)), f and f(y) are the distances from the center of the cylinder to v and Y, and the projection formula is finally derived as: Get the projection process from P to p.

3. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 2, characterized in that: The depth in polar coordinates is predicted in step 4: h = f (Wx + b); Where W is the weight matrix, b is the bias term, and x is the output or input data of the previous layer. is the activation function.

4. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 3, characterized in that: The feature stitching in step 4 includes combining the camera external parameters and FOV, calculating the angle range corresponding to each camera, and then performing a stitching operation on the features after depth estimation. In a clockwise direction, the features of each camera are filled into the corresponding position. The features are connected in sequence starting from the x-axis direction and stitched to 360 degrees.

5. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 4, characterized in that: The step five includes center point detection and center point and offset regression. The center point of the target is regressed and then expanded or offset from the center point to determine the complete position of the object. Each target object has a center point. The network locates the object by predicting the position of the center point. A regression network is used to predict the center point position of each target and the offset of the point.

6. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 5, characterized in that: The step 5 also includes center point and offset regression, which uses a regression network to predict the center point position of each target and the offset of the point. This method makes the positioning of the target more accurate, especially when dealing with small objects. Traditional bounding box regression may have certain errors. The center point regression method can usually better capture the geometric center of the object.

7. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 6, characterized in that: The conversion of Cartesian coordinates to polar coordinates in step 6 includes assuming that the coordinates of the point in the Cartesian coordinate system are (x, y), and converting them to polar coordinates (r, θ), where r is the distance from the point to the origin, and θ is the angle of the point relative to the x-axis. The conversion formula is as follows: ; ; Where r is calculated using the Pythagorean theorem, i.e. the straight-line distance from the origin to the point (x, y), and θ is the angle between the point and the x-axis, which is calculated using the inverse tangent function; Polar coordinates to Cartesian coordinates conversion, the formula is as follows: ; 。 8. The 3D object detection method based on cylindrical projection and polar coordinates according to claim 7, characterized in that: The loss calculation in step 7 includes the center point regression loss, which is calculated by combining Focal Loss and Gaussian Distribution as the center point regression loss; ; Where p(t) is the predicted probability that the sample belongs to the correct category, which indicates the model's predicted probability for the target category. N(p(t)|μ, σ(2)) is the weighting of the predicted probability p(t) using a Gaussian distribution, which is similar to adjusting the weight of each sample. σ is the standard deviation of the Gaussian distribution, which controls the size of the focus area. The smaller the standard deviation, the smaller the focus area. α and γ are the same adjustment factors as Focal Loss, which are used to balance the category imbalance and adjust the weight of easy-to-classify samples respectively. The loss calculation also includes the offset regression loss, which is calculated by calculating the absolute error between the predicted value and the true value of each sample and taking the average. ; in is the predicted value of the i-th sample, y i is the true value of the i-th sample, is the absolute error of the ith sample.