Target detection method and system based on point cloud data
By constructing and optimizing the three-dimensional and planar target point cloud data detection model and performing information fusion, the shortcomings of point cloud detection models in the existing technology are solved in the process of long-distance targets and single-sided targets, and higher detection accuracy and overall system accuracy are achieved.
Patent Information
- Application Number
- CN202510035087.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
The existing three-dimensional point cloud detection model performs poorly when dealing with forward backward targets, especially when the target distance is far away, and the labeling relies on subjective judgment to cause inconsistency, affecting the model learning effect. In addition, LiDAR sensors have weak performance when identifying distant single-sided targets, which are prone to missed detection.
By acquiring point cloud data, a three-dimensional target point cloud data detection model and a plane target point cloud data detection model are constructed separately, and its loss function is determined. Use the AdamW optimizer to optimize the model, combine the information of three-dimensional and two-dimensional targets to fusion, and output a list of fusion targets to improve detection accuracy.
Excellent performance when dealing with forward backward targets, strong recall ability, able to accurately detect long-distance targets, ensure consistency of labeling, avoid missed detection of LiDAR sensors when identifying distant single-sided targets, and improve the overall accuracy of the system.
Smart Images

Figure CN120014229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a target detection method and system based on point cloud data. Background Art
[0002] With the continuous development of autonomous driving technology, object detection technology, as an important basis for ensuring safe driving, has become one of the key technologies in autonomous driving and intelligent transportation systems. Accurate object detection can effectively identify obstacles, pedestrians and other traffic participants on the road, thereby improving vehicle safety and driving efficiency. At present, point cloud 3D object detection models are widely used in autonomous driving systems.
[0003] However, existing point cloud 3D detection models perform poorly when processing forward and backward targets. When the target is far away, the model often loses its recall ability and cannot accurately detect distant targets.
[0004] Secondly, the existing models rely on the subjective judgment of the annotators when labeling single-sided objects, which easily leads to inconsistency in labeling and affects the learning effect of the model due to the lack of accurate labeling.
[0005] In addition, existing LiDAR sensors have poor performance in identifying single-sided targets at a distance and are prone to missed detections, affecting the overall accuracy of the system. Summary of the invention
[0006] The main purpose of the present invention is to provide a target detection method and system based on point cloud data to solve the problems that the existing point cloud three-dimensional detection model performs poorly when processing forward and backward targets, loses the recall ability when the target is far away, and cannot accurately detect distant targets. The labeling of single-sided targets relies on the subjective judgment of the labeler, which easily leads to inconsistency in labeling. The lack of accurate labeling will affect the learning effect of the model, and the existing LiDAR sensor has weak performance when identifying single-sided targets in the distance, and is prone to missed detection, affecting the overall accuracy of the system.
[0007] In order to achieve the above object, according to one aspect of the present invention, a target detection method based on point cloud data is provided, comprising:
[0008] S1: Acquire point cloud data, wherein the point cloud data includes three-dimensional target point cloud data and plane target point cloud data;
[0009] S2: construct a 3D target point cloud data detection model and a 2D target point cloud data detection model respectively;
[0010] S3: Determine the loss function of the three-dimensional target point cloud data detection model and the planar target point cloud data detection model;
[0011] S4: optimizing the three-dimensional target point cloud data detection model and the plane target point cloud data detection model respectively through an AdamW optimizer with the goal of minimizing the loss function values of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model;
[0012] S5: inputting the three-dimensional target point cloud data and the plane target point cloud data into the optimized three-dimensional target point cloud data detection model and the plane target point cloud data detection model respectively to perform target detection, and outputting the three-dimensional target point cloud data detection result and the plane target point cloud data detection result;
[0013] S6: respectively decoding the three-dimensional target point cloud data detection result and the plane target point cloud data detection result to obtain a three-dimensional target list and a two-dimensional plane list;
[0014] S7: Fusing the three-dimensional target list and the two-dimensional plane list to obtain a fused target list including multiple detection targets.
[0015] Furthermore, the loss function of the three-dimensional target point cloud data detection model includes a first confidence loss function, a first classification loss function and a first positioning loss function;
[0016] The first confidence loss function is specifically:
[0017]
[0018] Among them, H[i,j] represents the first confidence loss function of the grid in the i-th row and j-th column, α represents the positive sample ratio coefficient, T[i,j] represents the heat map value of the 3D target point cloud data in the i-th row and j-th column, P[i,j] represents the confidence prediction value of the 3D target point cloud data in the i-th row and j-th column, and γ represents the hard example adjustment factor;
[0019] The first classification loss function is specifically:
[0020] F(p target )=-α target (1-p target ) γ ln(p target )
[0021]
[0022] Among them, F() represents the first classification loss function, p target represents the predicted probability of the target category, α targetrepresents the first category ratio coefficient of the target category, total represents the total amount of 3D target point cloud data, and quantity(c) represents the total amount of 3D target point cloud data of the cth target category;
[0023] The first positioning loss function is specifically:
[0024]
[0025] Among them, L represents the first positioning loss function, I(BBox p ,BBox t ) represents the intersection volume between the predicted 3D bounding box and the real 3D bounding box, BBox p Represents the predicted three-dimensional bounding box, BBox t Represents the true 3D bounding box, and U() represents the union volume between the predicted 3D bounding box and the true 3D bounding box.
[0026] Furthermore, the loss function of the plane target point cloud data detection model includes a second confidence loss function, a second classification loss function and a second positioning loss function:
[0027] The second confidence loss function is specifically:
[0028]
[0029] Among them, H'[i,j] represents the second confidence loss function of the grid in the i-th row and j-th column, T'[i,j] represents the heat map value of the plane target point cloud data in the i-th row and j-th column, and P'[i,j] represents the confidence prediction value of the plane target point cloud data in the i-th row and j-th column.
[0030] The second classification loss function is specifically:
[0031] F'(p target )=-α′ target (1-p target ) γ ln(p target )
[0032]
[0033] Among them, F() represents the second classification loss function, α′ target represents the second category ratio coefficient of the target category, total' represents the total amount of plane target point cloud data, quantity'(c) represents the total amount of plane target point cloud data of the cth target category;
[0034] The second positioning loss function is specifically:
[0035] L'=|Aa|+|Bb|+|Cc|+|Dd|
[0036] Among them, L' represents the second positioning loss function, Aa represents the distance between the vertex A of the real plane and the vertex a of the predicted plane, Bb represents the distance between the vertex B of the real plane and the vertex b of the predicted plane, Cc represents the distance between the vertex C of the real plane and the vertex c of the predicted plane, and Dd represents the distance between the vertex D of the real plane and the vertex d of the predicted plane.
[0037] Furthermore, the three-dimensional target point cloud data detection result specifically includes: a three-dimensional target heat map, a three-dimensional target bounding box and a three-dimensional target category matrix;
[0038] The plane target point cloud data detection result specifically includes: a plane target heat map, a plane target bounding box and a plane target category matrix.
[0039] Furthermore, the S6 specifically includes:
[0040] S601: Decode the three-dimensional target heat map and the two-dimensional target heat map respectively through the Sigmoid function to obtain a three-dimensional target activation value and a two-dimensional target activation value:
[0041]
[0042] Among them, Sigmoid() represents the Sigmoid function, and x represents the confidence value predicted by the model;
[0043] S602: Decoding the three-dimensional target category matrix and the three-dimensional target bounding box whose three-dimensional target activation value is greater than a first preset value respectively, to obtain a three-dimensional target list having a plurality of three-dimensional targets, wherein the three-dimensional target list includes a three-dimensional category sequence number, a three-dimensional target center point, a three-dimensional target size, and a three-dimensional target heading angle;
[0044] S603: Decode the plane target category matrix and the plane target bounding box whose plane target activation value is greater than the second preset value respectively to obtain a two-dimensional plane list with multiple two-dimensional plane targets, wherein the two-dimensional plane list includes a plane category sequence number, a plane target center point, a plane target size, and a plane target heading angle.
[0045] Furthermore, the decoding of the three-dimensional object bounding box whose three-dimensional object activation value is greater than the first preset value in S602 specifically includes:
[0046] Calculate the center point coordinates of the foreground grid in the three-dimensional target heat map:
[0047] x grid =-x range +(col+0.5)*d
[0048] y grid =-y range +(row+0.5)*d
[0049] z grid =0
[0050] Among them, x grid Represents the x-axis coordinate of the center point of the foreground grid in the 3D target heat map, x range represents the radius range of the foreground grid in the three-dimensional target heat map in the x-axis direction, col represents the column number of the foreground grid in the three-dimensional target heat map, d represents the size of the BEV spatial grid in the three-dimensional target heat map, and y grid Represents the y-axis coordinate of the center point of the foreground grid in the 3D target heat map, y range represents the radius range of the foreground grid in the 3D target heat map in the y-axis direction, row represents the row number of the foreground grid in the 3D target heat map, z grid Represents the z-axis coordinate of the center point of the foreground grid in the three-dimensional target heat map;
[0051] According to the center point coordinates of the foreground grid in the three-dimensional target heat map, the three-dimensional target center point is calculated:
[0052] x=x grid +Sigmoid(bbox tensor [0])*d
[0053] y=y grid +Sigmoid(bbox tensor [1])*d
[0054] z=z grid +Sigmoid(bbox tensor [2])*d
[0055] Among them, x represents the x-axis coordinate of the center point of the three-dimensional target, bbox tensor [0] represents the x-direction offset of the 3D target center point within the grid, y represents the y-axis coordinate of the 3D target center point, and bbox tensor [1] represents the y-direction offset of the 3D target center point within the grid, z represents the z-axis coordinate of the 3D target center point, and bbox tensor [2] represents the z-direction offset of the 3D target center point within the grid;
[0056] Activate the 3D scale channel through an exponential function to generate the 3D target size:
[0057] L = exp(bbox tensor [3])
[0058] W = exp(bboxtensor [4])
[0059] H = exp(bbox tensor [5])
[0060] Among them, L represents the length of the three-dimensional object, exp() represents the exponential function, and bbox tensor [3] represents the predicted value of the length of the three-dimensional object, W represents the width of the three-dimensional object, and bbox tensor [4] represents the predicted value of the width of the three-dimensional object, H represents the height of the three-dimensional object, and bbox tensor [5] represents the predicted height value of the three-dimensional object;
[0061] Calculate the heading angle of the three-dimensional target through the Sigmoid function:
[0062]
[0063] Among them, θ represents the heading angle of the three-dimensional target, bbox tensor [7] represents the predicted angle of the 3D target in the current 180° subinterval, sign() represents the sign function, and bbox tensor [6] represents the predicted value of the three-dimensional target interval.
[0064] Furthermore, decoding the plane target bounding box whose plane target activation value is greater than the second preset value in S603 specifically includes:
[0065] Calculate the center point coordinates of the foreground grid in the planar target heat map:
[0066] x' grid =-x' range +(col'+0.5)*d
[0067] y' grid =-y' range +(row'+0.5)*d
[0068] z' grid =0
[0069] Among them, x' grid Represents the x-axis coordinate of the center point of the foreground grid in the planar target heat map, x' range represents the radius range of the foreground grid in the plane target heat map in the x-axis direction, col' represents the column number of the foreground grid in the plane target heat map, d' represents the size of the BEV spatial grid in the plane target heat map, y' grid Represents the y-axis coordinate of the center point of the foreground grid in the planar target heat map, y' rangeIndicates the radius range of the foreground grid in the plane target heat map in the y-axis direction, row' indicates the row number of the foreground grid in the plane target heat map, z' grid Represents the z-axis coordinate of the center point of the foreground grid in the planar target heat map;
[0070] According to the center point coordinates of the foreground grid in the planar target heat map, the center point of the planar target is calculated:
[0071] x'=x' grid +Sigmoid(bbox′ tensor [0])*d'
[0072] y'=y' grid +Sigmoid(bbox′ tensor [1])*d'
[0073] z'=z' grid +Sigmoid(bbox′ tensor [2])*d'
[0074] Among them, x' represents the x-axis coordinate of the center point of the plane target, bbox′ tensor [0] represents the x-direction offset of the center point of the plane target within the grid, y' represents the y-axis coordinate of the center point of the plane target, and bbox′ tensor [1] represents the y-direction offset of the center point of the plane target within the grid, z' represents the z-axis coordinate of the center point of the plane target, and bbox′ tensor [2] represents the z-direction offset of the center point of the plane target within the grid;
[0075] Activate the plane scale channel through the exponential function to generate the plane target size:
[0076] W' = exp(bbox' tensor [4])
[0077] H' = exp(bbox' tensor [5])
[0078] Among them, W' represents the width of the plane target, bbox′ tensor [4] represents the predicted width of the planar target, H' represents the height of the planar target, and bbox′ tensor [5] represents the predicted height of a planar target;
[0079] The heading angle of the plane target is calculated using the Sigmoid function:
[0080]
[0081] Among them, θ' represents the heading angle of the plane target, bbox′tensor [7] represents the predicted angle of the planar target in the current 180° subinterval, bbox′ tensor [6] represents the predicted value of the plane target interval.
[0082] Furthermore, the S7 specifically includes:
[0083] S701: Calculate the distance between the center point of the three-dimensional target and the center point of the planar target to obtain a distance matrix:
[0084]
[0085] Where D[m,n] represents the distance between the center point of the mth three-dimensional target and the center point of the nth plane target, distance() represents the distance between the center point of the three-dimensional target and the center point of the plane target, || || represents the Euclidean norm, Represents the x-axis coordinate of the center point of the m-th three-dimensional target, Represents the x-axis coordinate of the center point of the nth plane target, Represents the y-axis coordinate of the center point of the m-th three-dimensional target, Represents the y-axis coordinate of the center point of the n-th plane target, Represents the z-axis coordinate of the center point of the m-th three-dimensional target, Represents the z-axis coordinate of the center point of the n-th plane target;
[0086] S702: performing target matching according to the distance matrix by using the Hungarian algorithm to obtain a target matching result;
[0087] S703: According to the target matching result, the three-dimensional target list and the two-dimensional plane list are merged to obtain a fused target list including multiple detection targets.
[0088] Furthermore, the S703 is specifically as follows:
[0089] According to the target matching results, all three-dimensional targets in the three-dimensional target list are retained, the two-dimensional plane targets in the two-dimensional plane list that are successfully matched to the three-dimensional targets are deleted, and the two-dimensional plane targets in the two-dimensional plane list that are not successfully matched to the three-dimensional targets are retained, so as to obtain a fused target list including multiple detection targets.
[0090] According to one aspect of the present invention, there is provided a target detection system based on point cloud data, comprising: a memory and one or more processors;
[0091] One or more applications are stored in the memory, and the one or more applications are suitable for being executed by the one or more processors to implement the target detection method based on point cloud data as described in any one of claims 1 to 9.
[0092] By applying the technical solution of the present invention, the loss functions of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model can be determined, and the loss function values of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model can be minimized. Through the AdamW optimizer, the three-dimensional target point cloud data detection model and the plane target point cloud data detection model are optimized respectively, and excellent performance is achieved when processing forward and backward targets. When the target distance is relatively close, the recall ability is relatively strong, and the long-distance targets can be accurately detected. In the process of labeling single-sided targets, the subjective judgment of the labeler is no longer relied on, and the consistency of the labeling is ensured. The learning effect of the model will not be affected by the lack of accurate labeling. By decoding the three-dimensional target point cloud data detection results and the plane target point cloud data detection results respectively, a three-dimensional target list and a two-dimensional plane list are obtained, and the three-dimensional target list and the two-dimensional plane list are fused to obtain a fused target list including multiple detection targets, thereby avoiding the existing LiDAR sensor having weak performance when identifying single-sided targets in the distance, and not prone to missed detection, thereby improving the overall accuracy of the system.
[0093] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0095] Figure 1 A schematic diagram of a process of a target detection method based on point cloud data provided by an embodiment of the present invention is shown;
[0096] Figure 2 A schematic structural diagram of a target detection system based on point cloud data provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0097] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0098] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0099] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present invention described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0100] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0101] Reference Manual Attached Figure 1 , shows a flow chart of a target detection method based on point cloud data provided by an embodiment of the present invention.
[0102] An embodiment of the present invention provides a target detection method based on point cloud data, comprising:
[0103] S1: Acquire point cloud data, which includes three-dimensional target point cloud data and plane target point cloud data.
[0104] S2: Construct a three-dimensional target point cloud data detection model and a planar target point cloud data detection model respectively.
[0105] In the present invention, the features of three-dimensional targets and two-dimensional plane targets are significantly different. Three-dimensional targets focus on depth, volume and spatial layout, while two-dimensional plane targets mainly focus on plane geometric features and local texture. Constructing models separately allows each model to focus more on its own feature extraction and avoid confusion caused by differences in target characteristics. Each model can independently design its network structure and loss function to better adapt to the corresponding data distribution and feature expression, thereby improving detection performance.
[0106] S3: Determine the loss function of the three-dimensional target point cloud data detection model and the planar target point cloud data detection model.
[0107] Optionally, the loss function of the three-dimensional target point cloud data detection model includes a first confidence loss function, a first classification loss function and a first positioning loss function.
[0108] The first confidence loss function is specifically:
[0109]
[0110] Among them, H[i,j] represents the first confidence loss function of the grid in the i-th row and j-th column, α represents the positive sample ratio coefficient, T[i,j] represents the heat map value of the 3D target point cloud data in the i-th row and j-th column, P[i,j] represents the confidence prediction value of the 3D target point cloud data in the i-th row and j-th column, and γ represents the hard example adjustment factor.
[0111] The first classification loss function is specifically:
[0112] F(p target )=-α target (1-p target ) γ ln(p target )
[0113]
[0114] Among them, F() represents the first classification loss function, p target represents the predicted probability of the target category, α target Represents the first category proportion coefficient of the target category, total represents the total amount of 3D target point cloud data, and quantity(c) represents the total amount of 3D target point cloud data of the cth target category.
[0115] The first positioning loss function is specifically:
[0116]
[0117] Among them, L represents the first positioning loss function, I(BBox p ,BBoxt ) represents the intersection volume between the predicted 3D bounding box and the real 3D bounding box, BBox p Represents the predicted three-dimensional bounding box, BBox t Represents the true 3D bounding box, and U() represents the union volume between the predicted 3D bounding box and the true 3D bounding box.
[0118] Optionally, the loss function of the planar target point cloud data detection model includes a second confidence loss function, a second classification loss function and a second positioning loss function:
[0119] The second confidence loss function is specifically:
[0120]
[0121] Among them, H'[i,j] represents the second confidence loss function of the grid in the i-th row and j-th column, T'[i,j] represents the heat map value of the plane target point cloud data in the i-th row and j-th column, and P'[i,j] represents the confidence prediction value of the plane target point cloud data in the i-th row and j-th column.
[0122] The second classification loss function is specifically:
[0123] F'(p target )=-α′ target (1-p target ) γ ln(p target )
[0124]
[0125] Among them, F() represents the second classification loss function, α′ target represents the second category ratio coefficient of the target category, total' represents the total amount of plane target point cloud data, and quantity'(c) represents the total amount of plane target point cloud data of the cth target category.
[0126] The second positioning loss function is specifically:
[0127] L'=|Aa|+|Bb|+|Cc|+|Dd
[0128] Among them, L' represents the second positioning loss function, Aa represents the distance between the vertex A of the real plane and the vertex a of the predicted plane, Bb represents the distance between the vertex B of the real plane and the vertex b of the predicted plane, Cc represents the distance between the vertex C of the real plane and the vertex c of the predicted plane, and Dd represents the distance between the vertex D of the real plane and the vertex d of the predicted plane.
[0129] In the present invention, the three-dimensional target point cloud data and the plane target point cloud data have significant differences in structure, dimension and feature expression. Designing loss functions separately can ensure that the key features of each data type (such as the volume intersection ratio of the three-dimensional bounding box and the coordinate accuracy of the two-dimensional plane vertices) are fully captured. Three-dimensional targets emphasize the volume relationship in three-dimensional space, while plane targets pay more attention to vertex positions and boundary characteristics. Through independent loss functions, these differentiated characteristics can be optimized separately to improve the accuracy of model detection. Defining confidence, classification and positioning losses as subtasks respectively can achieve joint optimization of multiple objectives, thereby improving the performance of the model in comprehensive performance. Each loss item can adjust the optimization direction by weight balancing, for example, increasing the weights of classification and positioning for complex scenes to meet higher precision requirements.
[0130] S4: With the goal of minimizing the loss function values of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model, the three-dimensional target point cloud data detection model and the plane target point cloud data detection model are optimized respectively through the AdamW optimizer.
[0131] It should be noted that the AdamW optimizer is an improved version of the Adam optimization algorithm, which is specially designed for deep learning model optimization. It adds a weight decay mechanism to the Adam algorithm, and effectively controls the model complexity and reduces the risk of overfitting by directly applying L2 regularization to the weight parameters. Unlike traditional weight decay, AdamW decouples weight decay from gradient updates, avoiding the problem of regularization failure caused by learning rate correction. The AdamW optimizer combines the fast convergence ability of Adam and the regularization effect of weight decay, and is one of the optimization algorithms widely used in current deep learning.
[0132] Specifically, a three-dimensional rectangular block annotation method is used to annotate the double-sided target of the target to be detected (including: the real three-dimensional target center point of the target to be detected, the three-dimensional target heading angle, length, width and height), and only the plane information is annotated for the single-sided target directly behind the target to be detected (including: the real plane target center point of the target to be detected, the plane target heading angle, width and height).
[0133] Furthermore, in the BEV space, the annotated two-sided targets and one-sided targets are encoded respectively to obtain the true three-dimensional target heat map, three-dimensional target bounding box, three-dimensional target category matrix, plane target heat map, plane target bounding box and plane target category matrix of the target to be detected.
[0134] It should be noted that BEV (Bird's Eye View) is a perspective commonly used in autonomous driving and robot vision, which refers to a view from a high altitude perpendicular to the ground. BEV images project objects in three-dimensional space onto a two-dimensional plane so that the scene can be viewed from above. They are usually used to show the spatial distribution of roads, obstacles and other objects. In autonomous driving, BEV views can provide a more intuitive scene understanding and help the system with target detection, path planning and decision making.
[0135] The encoding method of the heat map is as follows: for the marked two-sided targets and one-sided targets, the target matrix is constructed in the BEV space respectively, the grid value where the target center point is located is set to 1, and Gaussian blur processing is used to generate a heat map, with the foreground grid as the center and gradually decaying outward to 0.
[0136] It should be noted that Gaussian blur is an image processing technology that uses a Gaussian function to perform a convolution operation on an image to achieve a blurring effect. It blurs the details and noise in the image by smoothing the pixel values in the image, thereby removing high-frequency noise or achieving a visual smoothing effect. The core idea of Gaussian blur is to assign different weights according to the distance from the center point. The farther the pixel is from the center point, the smaller the weight. It is widely used in image processing, computer vision, edge detection and other fields.
[0137] The encoding method of the bounding box is as follows: the 3D target bounding box stores the offset from the 3D target center point to the grid, the 3D target size, and the 3D target heading angle for each foreground grid. The plane target bounding box only stores the offset from the plane target center point to the grid, the plane target size, and the plane target heading angle.
[0138] The encoding method of the category matrix is as follows: the three-dimensional target category matrix stores the category number in the three-dimensional target category matrix according to the real category of the three-dimensional target. The plane target category matrix only marks the plane information for the single-sided target, and the category number is stored in the plane target category matrix.
[0139] Furthermore, based on the real three-dimensional target heat map, three-dimensional target bounding box, three-dimensional target category matrix, plane target heat map, plane target bounding box and plane target category matrix of the target to be detected, with the goal of minimizing the loss function value of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model, the three-dimensional target point cloud data detection model and the plane target point cloud data detection model are optimized respectively through the AdamW optimizer.
[0140] In the present invention, the three-dimensional target (two-sided target) is annotated as a three-dimensional cuboid, including the center point, heading angle, length, width and height. The planar target (single-sided target) is only annotated with two-dimensional plane information. This annotation method enables the model to more clearly distinguish complex three-dimensional targets from simple two-dimensional targets, avoiding the interference of different target types on model optimization. By encoding the two-sided target and the single-sided target in the BEV (bird's eye view) space respectively, the heat map, bounding box and category matrix of the two-sided target and the single-sided target are obtained, and the geometry, position and category characteristics of the target can be captured, thereby improving the expression ability and detection accuracy of the model. The target center point and its probability distribution are visualized through the BEV heat map, making the detection results more intuitive and convenient for developers to understand and debug.
[0141] S5: input the three-dimensional target point cloud data and the plane target point cloud data into the optimized three-dimensional target point cloud data detection model and the plane target point cloud data detection model respectively for target detection, and output the three-dimensional target point cloud data detection results and the plane target point cloud data detection results.
[0142] Optionally, the three-dimensional target point cloud data detection results specifically include: a three-dimensional target heat map, a three-dimensional target bounding box, and a three-dimensional target category matrix.
[0143] The plane target point cloud data detection results specifically include: plane target heat map, plane target bounding box and plane target category matrix.
[0144] In the present invention, the geometric characteristics and information dimensions of three-dimensional targets (such as vehicles, buildings) and two-dimensional plane targets (such as road markings, walls) are different. By inputting the three-dimensional and two-dimensional targets into the optimization model and detecting them respectively, the characteristic confusion between the targets can be avoided and the accuracy of detection can be improved. The output target detection results (heat map, bounding box, category matrix) provide strong support for the intelligent decision-making of subsequent tasks (such as automatic driving, intelligent security, etc.), while simplifying the result analysis and post-processing process, and enhancing the robustness and scalability of the system. The output three-dimensional and two-dimensional heat maps can clearly display the probability distribution of the target, which is convenient for visualization and debugging of the detection results. The display of three-dimensional and two-dimensional bounding boxes can help analyze the position accuracy and boundary rationality of the target. The output of the category matrix is convenient for developers to quickly confirm whether the model is correctly classified.
[0145] S6: Decoding the three-dimensional target point cloud data detection results and the plane target point cloud data detection results respectively to obtain a three-dimensional target list and a two-dimensional plane list.
[0146] In the present invention, complex data structures such as heat maps, bounding boxes, and category matrices output by the model are converted into easy-to-read target lists (such as the center point, size, and category information of three-dimensional targets) through decoding, which is convenient for subsequent algorithms or modules to call directly. The decoded target list directly contains key information (such as the location, size, and category of the target), and there is no need to parse it again from the model output results, which improves processing efficiency.
[0147] In a possible implementation, S6 specifically includes sub-steps S601 to S603:
[0148] S601: The three-dimensional target heat map and the two-dimensional target heat map are decoded respectively through the Sigmoid function to obtain the three-dimensional target activation value and the two-dimensional target activation value:
[0149]
[0150] Among them, Sigmoid() represents the Sigmoid function, and x represents the confidence value predicted by the model.
[0151] In the present invention, by using the Sigmoid function to decode the three-dimensional target heat map and the plane target heat map, the three-dimensional target activation value and the plane target activation value are obtained, which can effectively convert the confidence value predicted by the model into a probability distribution, thereby accurately judging whether there is a target in each grid.
[0152] It should be noted that the Sigmoid function is a commonly used activation function that maps input values to between 0 and 1 and has the characteristics of an S-shaped curve. The Sigmoid function is often used in classification problems in machine learning, especially in binary classification tasks, to convert predicted values into probability values. In addition, the smoothness of the Sigmoid function enables it to effectively transfer gradients in neural networks and is suitable for dealing with nonlinear problems.
[0153] S602: Decode the three-dimensional target category matrix and the three-dimensional target bounding box whose three-dimensional target activation value is greater than the first preset value respectively to obtain a three-dimensional target list with multiple three-dimensional targets, wherein the three-dimensional target list includes a three-dimensional category number, a three-dimensional target center point, a three-dimensional target size, and a three-dimensional target heading angle.
[0154] In the present invention, the three-dimensional target category matrix and the three-dimensional target bounding box whose three-dimensional target activation value is greater than the first preset value are decoded, and the qualified three-dimensional targets can be effectively extracted from the model output. This approach can accurately identify and locate the target, avoid the interference of redundant or irrelevant targets, and at the same time, by extracting the category, center point, size, heading angle and other information of the three-dimensional target, the accuracy and integrity of the target detection are ensured, thereby improving the overall detection effect.
[0155] Optionally, in S602, decoding the three-dimensional object category matrix whose three-dimensional object activation value is greater than the first preset value is specifically performed as follows:
[0156] The three-dimensional target category matrix is activated through the argmax function to decode the three-dimensional target category matrix and obtain the three-dimensional category sequence number.
[0157] It should be noted that the argmax function is a mathematical operation that returns the index position of the maximum value in a given array or function. Specifically, the argmax function finds the position of the maximum value in the input data and returns the index value of that position. The argmax function is widely used in machine learning and deep learning, especially in classification problems, to determine the category corresponding to the maximum probability in the predicted value.
[0158] In a possible implementation manner, decoding the three-dimensional object bounding box whose three-dimensional object activation value is greater than the first preset value in S602 specifically includes:
[0159] Calculate the center coordinates of the foreground grid in the 3D target heat map:
[0160] x grid =-x range +(col+0.5)*d
[0161] y grid =-y range +(row+0.5)*d
[0162] z grid =0
[0163] Among them, x grid Represents the x-axis coordinate of the center point of the foreground grid in the 3D target heat map, x range represents the radius range of the foreground grid in the three-dimensional target heat map in the x-axis direction, col represents the column number of the foreground grid in the three-dimensional target heat map, d represents the size of the BEV spatial grid in the three-dimensional target heat map, and y grid Represents the y-axis coordinate of the center point of the foreground grid in the 3D target heat map, y range represents the radius range of the foreground grid in the 3D target heat map in the y-axis direction, row represents the row number of the foreground grid in the 3D target heat map, z grid Represents the z-axis coordinate of the center point of the foreground grid in the 3D object heat map.
[0164] According to the center point coordinates of the foreground grid in the 3D target heat map, calculate the 3D target center point:
[0165] x=x grid +Sigmoid(bbox tensor [0])*d
[0166] y=y grid +Sigmoid(bbox tensor [1])*d
[0167] z=z grid +Sigmoid(bbox tensor [2])*d
[0168] Among them, x represents the x-axis coordinate of the center point of the three-dimensional target, bbox tensor [0] represents the x-direction offset of the 3D target center point within the grid, y represents the y-axis coordinate of the 3D target center point, and bbox tensor [1] represents the y-direction offset of the 3D target center point within the grid, z represents the z-axis coordinate of the 3D target center point, and bbox tensor [2] represents the z-direction offset of the 3D target center point within the grid.
[0169] Activate the 3D scale channel through an exponential function to generate the 3D target size:
[0170] L = exp(bbox tensor [3])
[0171] W = exp(bbox tensor [4])
[0172] H = exp(bbox tensor [5])
[0173] Among them, L represents the length of the three-dimensional object, exp() represents the exponential function, and bbox tensor [3] represents the predicted value of the length of the three-dimensional object, W represents the width of the three-dimensional object, and bbox tensor [4] represents the predicted value of the width of the three-dimensional object, H represents the height of the three-dimensional object, and bbox tensor [5] represents the height prediction value of the three-dimensional object.
[0174] It should be noted that an exponential function is a special mathematical function characterized by the fact that the value of the output changes in a rapidly growing or decaying manner as the input changes. Simply put, when the input increases, the output increases rapidly at an exponential level, or if the input decreases, the output decreases at a similar rate. It is widely used in phenomena in nature, such as object decay and population growth, and also plays an important role in finance and computer science, such as describing processes such as compound interest and algorithm optimization.
[0175] Calculate the heading angle of the three-dimensional target through the Sigmoid function:
[0176]
[0177] Among them, θ represents the heading angle of the three-dimensional target, bbox tensor [7] represents the predicted angle of the 3D target in the current 180° subinterval, sign() represents the sign function, and bbox tensor [6] represents the predicted value of the three-dimensional target interval.
[0178] In the present invention, decoding the three-dimensional target category matrix can accurately extract the target category number, determine the target type and refine the attributes, and avoid category recognition errors. The coordinates of the center point of the foreground grid of the three-dimensional target heat map are calculated, and the position of the target center point can be accurately determined by combining the grid coordinates and the bounding box offset value, providing accurate spatial information for target positioning and analysis. The three-dimensional scale channel is activated by an exponential function to generate the target size, ensure the size accuracy, avoid prediction errors, and improve the robustness to the scene of target size changes. The target heading angle is calculated by the Sigmoid function and the sign function, and the spatial orientation of the target is accurately predicted to cope with complex posture changes and ensure accurate restoration of orientation information.
[0179] S603: Decode the plane target category matrix and the plane target bounding box whose plane target activation value is greater than the second preset value respectively to obtain a two-dimensional plane list with multiple two-dimensional plane targets, wherein the two-dimensional plane list includes a plane category serial number, a plane target center point, a plane target size, and a plane target heading angle.
[0180] In the present invention, the plane target category matrix and the plane target bounding box whose plane target activation value is greater than the second preset value are decoded, which can efficiently extract the category number of the plane target and ensure the classification accuracy. At the same time, the center point coordinates and offset information of the foreground grid are combined to accurately locate the center point position of the plane target.
[0181] Optionally, in S603, decoding the plane target category matrix whose plane target activation value is greater than the second preset value is specifically performed as follows:
[0182] The plane target category matrix is activated through the argmax function to decode the plane target category matrix and obtain the plane category sequence number.
[0183] In a possible implementation manner, decoding the plane target bounding box whose plane target activation value is greater than the second preset value in S603 specifically includes:
[0184] Calculate the center point coordinates of the foreground grid in the planar target heat map:
[0185] x' grid =-x' range +(col'+0.5)*d
[0186] y' grid =-y'range +(row'+0.5)*d
[0187] z' grid =0
[0188] Among them, x' grid Represents the x-axis coordinate of the center point of the foreground grid in the planar target heat map, x' range represents the radius range of the foreground grid in the plane target heat map in the x-axis direction, col' represents the column number of the foreground grid in the plane target heat map, d' represents the size of the BEV spatial grid in the plane target heat map, y' grid Represents the y-axis coordinate of the center point of the foreground grid in the planar target heat map, y' range Indicates the radius range of the foreground grid in the plane target heat map in the y-axis direction, row' indicates the row number of the foreground grid in the plane target heat map, z' grid Represents the z-axis coordinate of the center point of the foreground grid in the planar object heat map.
[0189] According to the center point coordinates of the foreground grid in the planar target heat map, calculate the center point of the planar target:
[0190] x'=x' grid +Sigmoid(bbox′ tensor [0])*d'
[0191] y'=y' grid +Sigmoid(bbox′ tensor [1])*d'
[0192] z'=z' grid +Sigmoid(bbox′ tensor [2])*d'
[0193] Among them, x' represents the x-axis coordinate of the center point of the plane target, bbox′ tensor [0] represents the x-direction offset of the center point of the plane target within the grid, y' represents the y-axis coordinate of the center point of the plane target, and bbox′ tensor [1] represents the y-direction offset of the center point of the plane target within the grid, z' represents the z-axis coordinate of the center point of the plane target, and bbox′ tensor [2] represents the z-direction offset of the center point of the planar target within the grid.
[0194] Activate the plane scale channel through the exponential function to generate the plane target size:
[0195] W' = exp(bbox' tensor [4])
[0196] H' = exp(bbox' tensor[5])
[0197] Among them, W' represents the width of the plane target, bbox′ tensor [4] represents the predicted width of the planar target, H' represents the height of the planar target, and bbox′ tensor [5] represents the predicted height of a planar target.
[0198] The heading angle of the plane target is calculated using the Sigmoid function:
[0199]
[0200] Among them, θ' represents the heading angle of the plane target, bbox′ tensor [7] represents the predicted angle of the planar target in the current 180° subinterval, bbox′ tensor [6] represents the predicted value of the plane target interval.
[0201] In the present invention, the planar target category matrix is activated by the argmax function, and the category number can be accurately obtained to ensure accurate calibration and reduce recognition errors. The coordinates of the center point of the foreground grid of the planar target heat map are calculated, and combined with the bounding box offset value, the true position of the target can be efficiently located, which is suitable for scenes that require precise plane positioning. The decoding method based on the grid center point and offset fully considers the slight position offset within the grid to ensure accurate center point decoding. The exponential function activates the scale channel, effectively reduces the nonlinear error of size prediction, and maintains stability and accuracy in scenes with large size changes. Combined with the 180° sub-interval prediction angle and symbol information, the heading angle is accurately calculated to adapt to the angle changes in complex scenes and improve the accuracy of posture analysis.
[0202] It should be noted that those skilled in the art can set the first preset value and the second preset value according to actual needs, and the present invention is not limited thereto.
[0203] S7: Fuse the three-dimensional target list and the two-dimensional plane list to obtain a fused target list.
[0204] In the present invention, the three-dimensional target list and the two-dimensional plane list are fused to fully integrate the three-dimensional spatial information and the two-dimensional plane information, make up for the shortcomings of a single detection model, and improve the comprehensiveness and accuracy of target detection. Through fusion, the spatial position, size, and direction information of the three-dimensional target can be combined with the category and size information of the two-dimensional plane, thereby providing a more complete target representation. This method is particularly suitable for application scenarios that require attention to both the global three-dimensional structure and the local plane features of the target, further enhancing the practicality and reliability of the detection results.
[0205] In a possible implementation, S7 specifically includes sub-steps S701 to S703:
[0206] S701: Calculate the distance between the center point of the three-dimensional target and the center point of the plane target to obtain a distance matrix:
[0207]
[0208] Where D[m,n] represents the distance between the center point of the mth three-dimensional target and the center point of the nth plane target, distance() represents the distance between the center point of the three-dimensional target and the center point of the plane target, || || represents the Euclidean norm, Represents the x-axis coordinate of the center point of the m-th three-dimensional target, Represents the x-axis coordinate of the center point of the nth plane target, Represents the y-axis coordinate of the center point of the m-th three-dimensional target, Represents the y-axis coordinate of the center point of the n-th plane target, Represents the z-axis coordinate of the center point of the m-th three-dimensional target, Indicates the z-axis coordinate of the center point of the n-th plane target.
[0209] S702: Perform target matching based on the distance matrix and the Hungarian algorithm to obtain a target matching result.
[0210] It should be noted that the Hungarian algorithm is an algorithm used to solve optimization problems, especially in bipartite graph matching and minimization / maximization cost problems. It gradually improves the quality of matching and eventually finds the optimal solution. Specifically, the Hungarian algorithm is often used to solve the "optimal matching" problem, such as assigning a set of tasks to a group of workers with the goal of minimizing the total cost. The algorithm constructs a cost matrix and gradually finds the best matching method through a series of adjustment and selection steps. It is widely used in graph theory, operations research and other fields.
[0211] S703: According to the target matching result, the three-dimensional target list and the two-dimensional plane list are merged to obtain a fused target list including multiple detection targets.
[0212] In a possible implementation manner, S703 specifically includes:
[0213] According to the target matching results, all three-dimensional targets in the three-dimensional target list are retained, the two-dimensional plane targets in the two-dimensional plane list that are successfully matched to the three-dimensional target are deleted, and the two-dimensional plane targets in the two-dimensional plane list that are not successfully matched to the three-dimensional target are retained, so as to obtain a fused target list including multiple detected targets.
[0214] In the present invention, by calculating the distance matrix between the three-dimensional target and the two-dimensional plane target and matching them in combination with the Hungarian algorithm, the target correspondence relationship can be accurately established, thereby effectively reducing the matching error. In the fusion process, all three-dimensional target information is retained, while the matched two-dimensional plane targets are eliminated to avoid information redundancy, and the unmatched two-dimensional plane targets are retained to ensure the comprehensiveness of the detection results. In this way, multi-source information can be integrated to improve detection accuracy and robustness, and the structure of the target list can be optimized, providing reliable support for target detection in complex scenes.
[0215] From the above description, it can be seen that the target detection method based on point cloud data provided by the present invention can determine the loss function of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model, and minimize the loss function value of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model. Through the AdamW optimizer, the three-dimensional target point cloud data detection model and the plane target point cloud data detection model are optimized respectively, and the performance is excellent when processing forward and backward targets. When the target distance is close, the recall ability is strong, and the long-distance target can be accurately detected. In the process of labeling single-sided targets, it no longer relies on the subjective judgment of the labeler, ensuring the consistency of the labeling, and will not affect the learning effect of the model due to the lack of accurate labeling. By decoding the three-dimensional target point cloud data detection results and the plane target point cloud data detection results respectively, a three-dimensional target list and a two-dimensional plane list are obtained, and the three-dimensional target list and the two-dimensional plane list are fused to obtain a fused target list including multiple detection targets, thereby avoiding the existing LiDAR sensor. The performance is weak when identifying single-sided targets in the distance, and it is not easy to miss the detection phenomenon, thereby improving the overall accuracy of the system.
[0216] Reference Manual Attached Figure 2 , shows a schematic structural diagram of a target detection system based on point cloud data provided by an embodiment of the present invention.
[0217] An embodiment of the present invention provides a point cloud data-based target detection system 30 , including: a memory 303 and one or more processors 301 .
[0218] One or more application programs are stored in the memory 303 , and the one or more application programs are suitable for being executed by the one or more processors 301 to implement the target detection method based on point cloud data described in the method embodiment.
[0219] The object detection system 30 based on point cloud data includes: a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, for example, via a bus 302.
[0220] The structure of the point cloud data-based object detection system 30 does not constitute a limitation on the embodiment of the present invention.
[0221] Processor 301 may be a CPU, a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0222] The bus 302 may include a path to transmit information between the above components. The bus 302 may be a PCI bus or an EISA bus, etc. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0223] The memory 303 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disk storage (including a compressed optical disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0224] It should be noted that the target detection system 30 based on point cloud data can implement the above-mentioned target detection method based on point cloud data, and can achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.
[0225] From the above description, it can be seen that the target detection system based on point cloud data provided by the present invention can determine the loss function of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model, and minimize the loss function value of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model. Through the AdamW optimizer, the three-dimensional target point cloud data detection model and the plane target point cloud data detection model are optimized respectively, and the performance is excellent when processing forward and backward targets. When the target distance is close, the recall ability is strong, and the long-distance target can be accurately detected. In the process of labeling single-sided targets, it no longer relies on the subjective judgment of the labeler, ensuring the consistency of the labeling, and will not affect the learning effect of the model due to the lack of accurate labeling. By decoding the three-dimensional target point cloud data detection results and the plane target point cloud data detection results respectively, a three-dimensional target list and a two-dimensional plane list are obtained, and the three-dimensional target list and the two-dimensional plane list are fused to obtain a fused target list including multiple detection targets, thereby avoiding the existing LiDAR sensor. The performance is weak when identifying single-sided targets in the distance, and it is not easy to miss detection, thereby improving the overall accuracy of the system.
[0226] Unless otherwise specifically stated, the relative arrangement of the parts and steps described in these embodiments, numerical expressions and numerical values do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to the actual proportional relationship. The technology, method and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but in appropriate cases, the technology, method and equipment should be regarded as a part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once a certain item is defined in an accompanying drawing, it does not need to be further discussed in subsequent drawings.
[0227] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used here to describe the spatial positional relationship between a device or feature and other devices or features as shown in the figure. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figure. For example, if the device in the accompanying drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below other devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.
[0228] In the description of the present invention, it is necessary to understand that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "lateral, vertical, perpendicular, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the devices or elements referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention. The directional words "inside and outside" refer to the inside and outside relative to the contours of each component itself.
[0229] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A target detection method based on point cloud data, characterized in that: include: S1: Acquire point cloud data, wherein the point cloud data includes three-dimensional target point cloud data and plane target point cloud data; S2: construct a 3D target point cloud data detection model and a 2D target point cloud data detection model respectively; S3: Determine the loss function of the three-dimensional target point cloud data detection model and the planar target point cloud data detection model; S4: optimizing the three-dimensional target point cloud data detection model and the plane target point cloud data detection model respectively through an AdamW optimizer with the goal of minimizing the loss function values of the three-dimensional target point cloud data detection model and the plane target point cloud data detection model; S5: inputting the three-dimensional target point cloud data and the plane target point cloud data into the optimized three-dimensional target point cloud data detection model and the plane target point cloud data detection model respectively to perform target detection, and outputting the three-dimensional target point cloud data detection result and the plane target point cloud data detection result; S6: respectively decoding the three-dimensional target point cloud data detection result and the plane target point cloud data detection result to obtain a three-dimensional target list and a two-dimensional plane list; S7: Fusing the three-dimensional target list and the two-dimensional plane list to obtain a fused target list including multiple detection targets.
2. The target detection method based on point cloud data according to claim 1, characterized in that: The loss function of the three-dimensional target point cloud data detection model includes a first confidence loss function, a first classification loss function and a first positioning loss function; The first confidence loss function is specifically: Among them, H[i,j] represents the first confidence loss function of the grid in the i-th row and j-th column, α represents the positive sample ratio coefficient, T[i,j] represents the heat map value of the 3D target point cloud data in the i-th row and j-th column, P[i,j] represents the confidence prediction value of the 3D target point cloud data in the i-th row and j-th column, and γ represents the hard example adjustment factor; The first classification loss function is specifically: F(p target )=-α target (1-p target ) γ ln(p target ) Among them, F() represents the first classification loss function, p target represents the predicted probability of the target category, α target represents the first category ratio coefficient of the target category, total represents the total amount of 3D target point cloud data, and quantity(c) represents the total amount of 3D target point cloud data of the cth target category; The first positioning loss function is specifically: Among them, L represents the first positioning loss function, I(BBox p ,BBox t ) represents the intersection volume between the predicted 3D bounding box and the real 3D bounding box, BBox p Represents the predicted three-dimensional bounding box, BBox t Represents the true 3D bounding box, and U() represents the union volume between the predicted 3D bounding box and the true 3D bounding box.
3. The target detection method based on point cloud data according to claim 1, characterized in that: The loss function of the plane target point cloud data detection model includes a second confidence loss function, a second classification loss function and a second positioning loss function: The second confidence loss function is specifically: Among them, H'[i,j] represents the second confidence loss function of the grid in the i-th row and j-th column, T'[i,j] represents the heat map value of the plane target point cloud data in the i-th row and j-th column, and P'[i,j] represents the confidence prediction value of the plane target point cloud data in the i-th row and j-th column. The second classification loss function is specifically: F'(p target )=-α' target (1-p target ) γ ln(p target ) Among them, F() represents the second classification loss function, α' target represents the second category ratio coefficient of the target category, total' represents the total amount of plane target point cloud data, quantity'(c) represents the total amount of plane target point cloud data of the cth target category; The second positioning loss function is specifically: L'=|Aa|+|Bb|+|Cc|+|Dd| Among them, L' represents the second positioning loss function, Aa represents the distance between the vertex A of the real plane and the vertex a of the predicted plane, Bb represents the distance between the vertex B of the real plane and the vertex b of the predicted plane, Cc represents the distance between the vertex C of the real plane and the vertex c of the predicted plane, and Dd represents the distance between the vertex D of the real plane and the vertex d of the predicted plane.
4. The target detection method based on point cloud data according to claim 1, characterized in that: The three-dimensional target point cloud data detection results specifically include: a three-dimensional target heat map, a three-dimensional target bounding box and a three-dimensional target category matrix; The plane target point cloud data detection result specifically includes: a plane target heat map, a plane target bounding box and a plane target category matrix.
5. The target detection method based on point cloud data according to claim 4, characterized in that: The S6 specifically includes: S601: Decode the three-dimensional target heat map and the two-dimensional target heat map respectively through the Sigmoid function to obtain a three-dimensional target activation value and a two-dimensional target activation value: Among them, Sigmoid() represents the Sigmoid function, and x represents the confidence value predicted by the model; S602: Decoding the three-dimensional target category matrix and the three-dimensional target bounding box whose three-dimensional target activation value is greater than a first preset value respectively, to obtain a three-dimensional target list having a plurality of three-dimensional targets, wherein the three-dimensional target list includes a three-dimensional category sequence number, a three-dimensional target center point, a three-dimensional target size, and a three-dimensional target heading angle; S603: Decode the plane target category matrix and the plane target bounding box whose plane target activation value is greater than the second preset value respectively to obtain a two-dimensional plane list with multiple two-dimensional plane targets, wherein the two-dimensional plane list includes a plane category sequence number, a plane target center point, a plane target size, and a plane target heading angle.
6. The target detection method based on point cloud data according to claim 5, characterized in that: The decoding of the three-dimensional object bounding box whose three-dimensional object activation value is greater than the first preset value in S602 specifically includes: Calculate the center point coordinates of the foreground grid in the three-dimensional target heat map: x grid =-x range +(col+0.5)*d y grid =-y range +(row+0.5)*d z grid =0 Among them, x grid Represents the x-axis coordinate of the center point of the foreground grid in the 3D target heat map, x range represents the radius range of the foreground grid in the three-dimensional target heat map in the x-axis direction, col represents the column number of the foreground grid in the three-dimensional target heat map, d represents the size of the BEV spatial grid in the three-dimensional target heat map, and y grid Represents the y-axis coordinate of the center point of the foreground grid in the 3D target heat map, y range represents the radius range of the foreground grid in the 3D target heat map in the y-axis direction, row represents the row number of the foreground grid in the 3D target heat map, z grid Represents the z-axis coordinate of the center point of the foreground grid in the three-dimensional target heat map; According to the center point coordinates of the foreground grid in the three-dimensional target heat map, the three-dimensional target center point is calculated: x=x grid +Sigmoid(bbox tensor [0])*d y=y grid +Sigmoid(bbox tensor [1])*d z=z grid +Sigmoid(bbox tensor [2])*d Among them, x represents the x-axis coordinate of the center point of the three-dimensional target, bbox tensor [0] represents the x-direction offset of the 3D target center point within the grid, y represents the y-axis coordinate of the 3D target center point, and bbox tensor [1] represents the y-direction offset of the 3D target center point within the grid, z represents the z-axis coordinate of the 3D target center point, and bbox tensor [2] represents the z-direction offset of the 3D target center point within the grid; Activate the 3D scale channel through an exponential function to generate the 3D target size: L=exp(bbox tensor [3]) W=exp(bbox tensor [4]) H=exp(bbox tensor [5]) Among them, L represents the length of the three-dimensional object, exp() represents the exponential function, and bbox tensor [3] represents the predicted value of the length of the three-dimensional object, W represents the width of the three-dimensional object, and bbox tensor [4] represents the predicted value of the width of the three-dimensional object, H represents the height of the three-dimensional object, and bbox tensor [5] represents the predicted height value of the three-dimensional object; Calculate the heading angle of the three-dimensional target through the Sigmoid function: Among them, θ represents the heading angle of the three-dimensional target, bbox tensor [7] represents the predicted angle of the 3D target in the current 180° subinterval, sign() represents the sign function, and bbox tensor [6] represents the predicted value of the three-dimensional target interval.
7. The target detection method based on point cloud data according to claim 5, characterized in that: The decoding of the plane target bounding box whose plane target activation value is greater than the second preset value in S603 specifically includes: Calculate the center point coordinates of the foreground grid in the planar target heat map: x' grid =-x' range +(col'+0.5)*d y' grid =-y' range +(row'+0.5)*d z' grid =0 Among them, x' grid Represents the x-axis coordinate of the center point of the foreground grid in the planar target heat map, x' range represents the radius range of the foreground grid in the plane target heat map in the x-axis direction, col' represents the column number of the foreground grid in the plane target heat map, d' represents the size of the BEV spatial grid in the plane target heat map, y' grid Represents the y-axis coordinate of the center point of the foreground grid in the planar target heat map, y' range Indicates the radius range of the foreground grid in the plane target heat map in the y-axis direction, row' indicates the row number of the foreground grid in the plane target heat map, z' grid Represents the z-axis coordinate of the center point of the foreground grid in the planar target heat map; According to the center point coordinates of the foreground grid in the planar target heat map, the center point of the planar target is calculated: x'=x' grid +Sigmoid(bbox' tensor [0])*d' y'=y' grid +Sigmoid(bbox' tensor [1])*d' z'=z' grid +Sigmoid(bbox' tensor [2])*d' Among them, x' represents the x-axis coordinate of the center point of the plane target, bbox' tensor [0] represents the x-axis offset of the center point of the plane target within the grid, y' represents the y-axis coordinate of the center point of the plane target, and bbox' tensor [1] represents the y-direction offset of the center point of the plane target within the grid, z' represents the z-axis coordinate of the center point of the plane target, and bbox' tensor [2] represents the z-direction offset of the center point of the plane target within the grid; Activate the plane scale channel through the exponential function to generate the plane target size: W'=exp(bbox' tensor [4]) H'=exp(bbox' tensor [5]) Among them, W' represents the width of the plane target, bbox' tensor [4] represents the predicted width of the planar target, H' represents the height of the planar target, and bbox' tensor [5] represents the predicted height of a planar target; The heading angle of the plane target is calculated using the Sigmoid function: Among them, θ' represents the heading angle of the plane target, bbox' tensor [7] represents the predicted angle of the planar target in the current 180° subinterval, bbox' tensor [6] represents the predicted value of the plane target interval.
8. The target detection method based on point cloud data according to claim 5, characterized in that: The S7 specifically includes: S701: Calculate the distance between the center point of the three-dimensional target and the center point of the planar target to obtain a distance matrix: Where D[m,n] represents the distance between the center point of the mth three-dimensional target and the center point of the nth plane target, distance() represents the distance between the center point of the three-dimensional target and the center point of the plane target, || || represents the Euclidean norm, Represents the x-axis coordinate of the center point of the m-th three-dimensional target, Represents the x-axis coordinate of the center point of the nth plane target, Represents the y-axis coordinate of the center point of the m-th three-dimensional target, Represents the y-axis coordinate of the center point of the n-th plane target, represents the z-axis coordinate of the center point of the mth three-dimensional target, z planen Represents the z-axis coordinate of the center point of the n-th plane target; S702: performing target matching according to the distance matrix by using the Hungarian algorithm to obtain a target matching result; S703: According to the target matching result, the three-dimensional target list and the two-dimensional plane list are merged to obtain a fused target list including multiple detection targets.
9. The target detection method based on point cloud data according to claim 8, characterized in that: The S703 is specifically as follows: According to the target matching results, all three-dimensional targets in the three-dimensional target list are retained, the two-dimensional plane targets in the two-dimensional plane list that are successfully matched to the three-dimensional targets are deleted, and the two-dimensional plane targets in the two-dimensional plane list that are not successfully matched to the three-dimensional targets are retained, so as to obtain a fused target list including multiple detection targets.
10. A target detection system based on point cloud data, characterized in that: include: memory and one or more processors; One or more applications are stored in the memory, and the one or more applications are suitable for being executed by the one or more processors to implement the target detection method based on point cloud data as described in any one of claims 1 to 9.