3D object detection method and device with high distant object recall rate, moving tool and storage medium

In the 3D object detection in the field of autonomous driving, coarse detection is performed using voxel features or cylinder features, and combined with fine detection of point features, the problem of low recall of long-distance objects is solved, and efficient and accurate detection of remote objects is achieved.

CN120126092APending Publication Date: 2025-06-10BEIJING ZHIXINGZHE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311673215.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, when performing 3D object detection in the field of autonomous driving, the recall rate of long-distance objects is low, resulting in greater challenges in perception tasks.

Method used

By obtaining point cloud data, using voxel features or cylinder features as input, rough detection is performed to obtain 3D object candidate box information, and then fine detection is performed using point features to perform object classification recognition, improving the recall rate of distant objects.

Benefits of technology

The detection efficiency of candidate boxes and the recognition accuracy of 3D objects have been improved, especially for distant objects, the recall rate and recognition accuracy have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126092A_ABST
    Figure CN120126092A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D object detection method and device with a high distant object recall rate, a mobile tool and a storage medium. The method comprises the following steps: acquiring point cloud data; inputting voxel features or cylinder features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate frame information output by the 3D object detection model; and obtaining point features corresponding to the 3D object candidate frame information, and inputting the point features to a pre-trained candidate frame classifier model, so as to determine an identification result of the point cloud data according to an output result of the candidate frame classifier model. According to the invention, the voxel feature or the cylinder feature corresponding to the point cloud data is used as the input part of the 3D object detection model to obtain the 3D object candidate frame information, and the point feature of the point cloud data corresponding to the 3D object candidate frame information is used as the input part of the candidate frame classifier model to carry out object classification and recognition. The object recall rate of 3D object detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of obstacle perception for autonomous driving, and particularly to a 3D object detection method, device, mobile tool and storage medium with a high recall rate for distant objects. Background Art

[0002] In the prior art, when performing 3D object detection in the field of autonomous driving, the input data is the point cloud data sensed by a mechanical lidar. During the process of 3D object detection, the sparsity of the point cloud and the accuracy of the detection algorithm will both cause the point cloud features of the 3D object to be unstable, resulting in missed detections. Commonly, for an object at a relatively far distance from the vehicle, when the distance between the object and the vehicle is greater than 60m, due to the long distance, the point cloud data sensed for the object is scarce, and it is easy to have unclear point cloud features of the object, thus causing missed detections. Currently, the recall rates of most 3D object detection algorithms for objects within a range greater than 60m are relatively low, thus posing a great challenge to the perception task. It can be seen that there is an urgent need in the industry for a solution that can improve the recall rate of detected distant objects. Summary of the Invention

[0003] Embodiments of the present invention provide a 3D object detection method, device, mobile tool and storage medium with a high recall rate for distant objects to solve the problem of the relatively low recall rate of distant objects in the prior art.

[0004] In a first aspect, embodiments of the present invention provide a 3D object detection method with a high recall rate for distant objects, including:

[0005] Obtain point cloud data;

[0006] Input the voxel features or column features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model;

[0007] Obtain point features corresponding to the 3D object candidate box information and input them into a pre-trained candidate box classifier model to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

[0008] In a second aspect, embodiments of the present invention provide a 3D object detection device with a high recall rate for distant objects, including:

[0009] A point cloud data acquisition module, configured to obtain point cloud data;

[0010] A candidate box acquisition module, configured to input the voxel features or column features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model;

[0011] An object recognition module, configured to obtain point features corresponding to the 3D object candidate box information and input the point features into a pre-trained candidate box classifier model, so as to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

[0012] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0013] A memory, configured to store executable instructions; and

[0014] A processor, configured to execute the executable instructions stored in the memory, and the executable instructions, when executed by the processor, implement the steps of the 3D object detection method with a high recall rate for distant objects in the first aspect described above.

[0015] In a fourth aspect, an embodiment of the present invention provides a mobile tool, where the mobile tool includes: the 3D object detection device with a high recall rate for distant objects in the second aspect described above or the electronic device in the third aspect described above.

[0016] In a fifth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and the program, when executed by a processor, implements the steps of the 3D object detection method with a high recall rate for distant objects in the first aspect described above.

[0017] In a sixth aspect, an embodiment of the present invention provides a computer program product, where the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the 3D object detection method with a high recall rate for distant objects in the first aspect described above.

[0018] The beneficial effects of the embodiments of the present invention are as follows: The 3D object detection method with a high recall rate for distant objects provided by the embodiments of the present invention first uses the voxel features or cylinder features corresponding to the point cloud data as the input part of the 3D object detection model to obtain 3D object candidate box information, and then uses the point features corresponding to the 3D object candidate box information as the input part of the candidate box classifier model to perform object classification and recognition. With such a design, it is possible to first perform rough detection using relatively efficient but less accurate voxel features or cylinder features to quickly and efficiently obtain 3D object candidate boxes within the entire detection range, and then use the point features corresponding to the relatively slower but more accurate 3D object candidate box information for fine detection to accurately classify the roughly detected 3D object candidate boxes. The combination of the two can not only effectively improve the detection efficiency of the candidate boxes, but also effectively improve the recognition accuracy of 3D objects; in addition, in the extraction of features, the method of the embodiments of the present invention fully considers the characteristics of the global features of distant objects being not obvious due to the sparsity of the point cloud and being easily occluded, etc. By extracting more localized features, it is also possible to effectively improve the object recall rate of distant 3D object detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of a 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention;

[0021] Figure 2 It is a flowchart of a training method of a 3D object detection model adopted in step S12 of the 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention;

[0022] Figure 3 It is a flowchart of the method of step S22 in the 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention;

[0023] Figure 4 It is a flowchart of a training method of a candidate box classifier model adopted in step S13 of the 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention;

[0024] Figure 5 It is a flowchart of a method for obtaining a second training data set in step S41 of the 3D fifth detection method with a high recall rate for distant objects according to an embodiment of the present invention;

[0025] Figure 6 The principle block diagram of a 3D object detection device with a high recall rate for distant objects according to an embodiment of the present invention;

[0026] Figure 7 The principle block diagram of a mobile tool according to an embodiment of the present invention;

[0027] Figure 8 The structural schematic diagram of an embodiment of an electronic device according to the present invention. Detailed implementation manners

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0029] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0030] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0031] In the present invention, "module", "device", "system", etc. refer to relevant entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable component, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server may both be components. One or more components may be in a process and / or thread in execution, and the components may be localized on one computer and / or distributed between two or more computers, and may be run by various computer-readable media. The components may also communicate through local and / or remote processes according to a signal having one or more data packets, for example, a signal from a data interacting with another component in a local system, a distributed system, and / or a signal interacting with other systems through a network on the Internet.

[0032] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or also include elements inherent to such a process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device comprising the said elements.

[0033] The 3D object detection method with a high recall rate for distant objects in the embodiments of the present invention can be applied in an object detection device, so that the user can use the method of the present invention for efficient, high-precision and high-recall object detection, especially for the detection of distant objects, and the recall rate and recognition accuracy can also be significantly improved. These object detection devices include but are not limited to the perception module on an autonomous driving device, the controller on an autonomous driving device, a smart phone, a smart tablet, a personal PC, a computer, a cloud server, etc., and the present invention does not make any limitation thereto. In particular, the 3D object detection method with a high recall rate for distant objects in the embodiments of the present invention can also be applied to autonomous driving devices such as autonomous driving vehicles (such as passenger cars, buses, minibuses, buses, trucks, sanitation vehicles, sweeping vehicles, floor washing vehicles, etc.), floor cleaning robots, motorcycles, electric vehicles, etc., and the present invention does not make any limitation thereto.

[0034] The present invention will be further described in detail below with reference to the accompanying drawings.

[0035] Regarding the problem of low recall rate of distant object detection existing in the algorithm detection model in the prior art, the inventor has found through research that the existing algorithm models often habitually adopt global features when performing object detection. For distant objects, due to the sparse and occluded characteristics of the point cloud, their global features are not obvious. Therefore, when using these existing algorithm models to detect distant objects, it is very easy to cause problems such as missed detection and low recall rate. Based on this, the inventor thought of finding more localized and accurate features to use these features to train and generate an algorithm model with a high recall rate. Under the guidance of this research concept, the inventor proposed a technical solution that combines rough detection and fine detection to improve the recall rate and detection efficiency of distant objects. Taking the algorithm model trained to include a 3D object detection model in the rough detection stage and a candidate box classifier model in the fine detection stage, and the fine detection feature used is the point feature corresponding to the 3D object candidate box information detected in the rough detection stage as an example, Figure 1Schematically shows a 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0036] Step S11: Obtain point cloud data;

[0037] Step S12: Input the voxel features or column features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model;

[0038] Step S13: Obtain point features corresponding to the 3D object candidate box information and input them into a pre-trained candidate box classifier model to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

[0039] In step S11, the obtained point cloud data is the point cloud data collected by the vehicle itself and to be used for object detection and recognition. As some possible implementation manners, the acquisition of the point cloud data can be real-time point cloud data or point cloud data within a specified time collected based on devices such as a laser point cloud radar or a camera provided on the vehicle itself. Since this part of the content is a conventional technical means in the prior art, no further elaboration will be made here.

[0040] Step S12 is a step of analyzing and processing the point cloud data based on the point cloud data obtained in step S11 to obtain the information of the recognized 3D object candidate boxes. Among them, in step S12, the 3D object detection model adopted is a pre-trained algorithm model that takes the point cloud data as the input and the 3D object candidate box information as the output. Among them, the basic algorithm model selected during training can be, for example, the pointpillars algorithm model, etc. The selected pointpillars algorithm model can be trained accordingly to obtain a 3D object detection model that meets the expectations. As a preferred implementation manner, during model training, the training objective of the 3D object detection model of the embodiment of the present invention is to generate an object detection model with high efficiency and high recall, especially a 3D object detection model suitable for high-efficiency and high-recall detection of distant objects. Based on this, in the embodiment of the present invention, when training the selected basic algorithm model such as the pointpillars algorithm model, the input data used is preferably the voxel feature (voxel feature) or pillar feature (pillar-based feature) extracted from the point cloud data. Thus, in step S12, the point cloud data input into the 3D object detection model is the voxel feature or pillar feature corresponding to the obtained point cloud data. Compared with the point cloud data of general point features, the point cloud data of voxel features or pillar features can better reflect the contour of the detected 3D object. And because the accuracy of voxel / pillar-based features is relatively low, the 3D object detection model trained by using it can quickly and roughly detect the 3D object candidate boxes in the entire detection range, greatly improving the efficiency and recall rate of 3D object candidate box detection. Especially for distant objects with sparse point clouds and easy to be occluded, the detection efficiency and recall rate of 3D object candidate boxes are also very high. Therefore, compared with the 3D object detection method that uses the point features corresponding to the point cloud data as the input, the 3D object detection method that uses the voxel feature or pillar feature corresponding to the point cloud data as the input can effectively improve the detection rate of candidate boxes. Among them, the acquisition method of the voxel feature or pillar feature corresponding to the point cloud data can be obtained by transforming according to the x, y, z coordinates of the point cloud in the voxel or pillar. Exemplarily, there are N points [(x 1 ,y 1 ,z 1 ),......,(x n ,y n ,z n)]For example, these three-dimensional coordinate data can be transformed into high-dimensional voxel features or cylinder features. The transformation method can be, for example, through a CNN neural network. In the embodiments of the present invention, specifically, these high-dimensional voxel features or cylinder features are selected for model training. Among them, the specific implementation of transforming three-dimensional coordinate data into high-dimensional voxel features or cylinder features through a CNN neural network can be processed with reference to relevant existing technologies, and this part of the content will not be elaborated here. Preferably, in the embodiments of the present invention, ten dimensions can be selected from the transformed high-dimensional voxel features or cylinder features to form the high-dimensional voxel features or cylinder features as the input part of the model for training. The input part of the resulting model is then [f 1 , f 2 ,......, f n , where f i is a ten-dimensional feature vector. For the 3D object candidate box information obtained through the 3D object detection model, exemplarily, it can include information such as the position of the 3D object candidate box, the size of the 3D object candidate box, and the yaw angle, that is, the coordinate position (x, y, z) of the 3D object candidate box, the length, width, and height (l, w, h), and the yaw angle theta data, so as to be able to represent and form the corresponding 3D object candidate box. As a possible implementation manner, the confidence threshold of the 3D object detection model can be set to a relatively low value, such as 0.2. With such a design, it can further effectively improve the object detection rate in the rough detection stage, thereby being able to recall objects as much as possible and improve the recall rate of objects. It should be noted that in the embodiments of the present invention, the recall rate is an index used to reflect the number of detected objects. The more objects are detected, the higher the recall rate, and the fewer objects are detected, the lower the recall rate. In the object detection methods of the prior art, due to the sparse point clouds of distant objects and being easily blocked, etc., it is very easy to miss detections, so there is a problem of relatively low recall rate for distant objects.

[0041] Step S13 is a step of analyzing and processing the point cloud data corresponding to the 3D object candidate box information based on the 3D object candidate box information obtained in step S12 to classify and identify the object type corresponding to the 3D object candidate box information. Among them, in step S13, the candidate box classifier model used is a classification model that takes the point cloud data as the input and the object classification information obtained by classification and recognition as the output. The corresponding candidate box classifier model can be obtained by pre-training the selected algorithm model. The selected algorithm model can exemplarily be a high-precision classification model such as pointnet / pointnet++, or a pointpillars algorithm model. Different from the 3D object detection model, the candidate box classifier model takes the point features corresponding to the point cloud data as the input. With such a design, in cooperation with the 3D object detection model that takes the voxel features or cylinder features of the point cloud data as the input, it is possible to first quickly and efficiently detect the candidate box of the 3D object by means of rough detection, and then use the fine detection method to accurately classify the detected 3D object candidate box. As a result, not only can the detection rate of the candidate box be effectively improved, but also the accuracy of object recognition can be effectively improved. The combination of the two can effectively improve the object recall rate of 3D object detection. As a preferred implementation manner, when training the candidate box classifier model in the embodiment of the present invention, the training objective of the model is to train a high-precision 3D candidate box classification model to accurately classify the output result of the foregoing 3D object detection model. Based on this, the input of the candidate box classifier model trained in the embodiment of the present invention is the point cloud feature data corresponding to the 3D object candidate box output by the 3D object detection model, and the output of the trained candidate box classifier model is the object classification information including the object type, object position, object size, and yaw angle. Since the embodiment of the present invention performs classification and recognition by extracting localized features (i.e., the point cloud features corresponding to the detected 3D object candidate box) in the fine detection stage, it can effectively avoid the problems of high omission rate and low recall rate caused by the unclear global features of distant objects due to sparse point clouds and easy occlusion. Therefore, the detection scheme combining rough detection and fine detection provided by the embodiment of the present invention has a high recall rate for distant objects, can effectively avoid the deficiencies of the detection schemes in the prior art, and provides an effective solution for the high-recall rate detection of distant objects.

[0042] In order to further improve the recall rate of distant objects, when training the algorithm model, the inventor also thought of conducting research and proposing solutions from the perspective of the training dataset. Specifically, the inventor found that the weight of distant objects in the adopted dataset has a direct impact on the recall rate of the algorithm model trained for distant objects. Based on this, the inventor thought of forming a training dataset by dynamically adjusting the weights of the sample set to train the algorithm model, so as to obtain an algorithm model with a high recall rate for distant objects, in order to solve the problem of low recall rate of distant objects in the prior art. Specifically, Figure 2 Schematically shows the training method flow of the 3D object detection model adopted in step S12 in the 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention. Refer to Figure 2 As shown, the method can be specifically implemented as including the following steps:

[0043] Step S21: Obtain a first sample dataset, where the first sample dataset includes a sample acquisition center, at least one sample object, and first sample data corresponding to the sample object. Among them, the first sample data includes point cloud data corresponding to the sample object and 3D object candidate box information corresponding to the sample object;

[0044] Step S22: Weight the first sample data according to the distance between each sample object in the first sample dataset and the sample acquisition center to form a first training dataset, where the first training dataset includes the weighted point cloud data corresponding to the sample object and 3D object candidate box information corresponding to the sample object;

[0045] Step S23: Use the voxel feature or column feature corresponding to the point cloud data in the first training dataset as the input, and use the 3D object candidate box information in the first training dataset as the output to train the selected algorithm model, and obtain a 3D object detection model with the voxel feature or column feature as the input and the candidate box information of the 3D object as the output.

[0046] In step S21, the first sample data set obtained is a sample data set for training a 3D object detection model. Therefore, in the first sample data set, there should be at least the first sample data corresponding to the sample object. Among them, the first sample data includes the point cloud data corresponding to the sample object, which is the input part of the 3D object detection model, and the 3D object candidate box information corresponding to the point cloud data, which is the output part of the 3D object detection model. In order to dynamically adjust the weight of distant objects, in the first sample data set of the embodiment of the present invention, there is also a sample acquisition center point and at least one sample object. The sample object is the basis of the first sample data. The point cloud data and the 3D object candidate box information in the first sample data are both data information corresponding to the sample object. The sample acquisition center point is the data acquisition center position information (exemplarily, it can be the position of the vehicle itself) when collecting the first sample data corresponding to each sample object in the first sample data set. According to the sample acquisition center point, the distance between each sample object in the first sample data set and the sample acquisition center point can be obtained (for example, specifically, the distance between the two can be determined by the coordinates of the sample acquisition center point and the coordinates of each sample object). In the embodiment of the present invention, preferably, the weight of distant objects is dynamically adjusted according to the distance between the sample object and the sample acquisition center point, so that the recall rate of the trained algorithm model for distant objects is higher. Exemplarily, a weight expression can be added to the first sample data corresponding to the sample object according to the distance between the sample object and the sample acquisition center point to optimize the training target, so that the trained model has a high recall rate for distant objects.

[0047] In a possible implementation manner of the present invention, a weight expression is added to the first sample data corresponding to the sample object according to the distance between the sample object and the sample acquisition center point, which can be achieved by means of weighted processing. Specifically, step S22 is a step of performing weighted processing on the first sample data corresponding to each sample object in the first sample data set. Among them, when weighting each first sample data, the first sample data is weighted according to the distance between each sample object in the first sample data set and the sample acquisition center point. Among them, the distance between the sample object and the sample acquisition center point can be understood as the distance between the center point of the sample object and the sample acquisition center point, or the distance between the coordinate position of the sample object and the coordinate position of the sample acquisition center point. Preferably, by way of example, the coordinate position of the sample acquisition center point can be taken from the coordinate origin of the ego-vehicle coordinate system when the vehicle performs acquisition. Thus, the distance between the sample object and the sample acquisition center point refers to the distance between the coordinate position where the sample object is located and the coordinate origin of the ego-vehicle coordinate system when the ego-vehicle performs data acquisition. After weighting the first sample data in the first sample data set, a first training data set for training the 3D object detection model is formed. In this first training data set, it includes the weighted point cloud data corresponding to the sample object and the 3D object candidate box information corresponding to the sample object. Since the point cloud data in the first training data set is designed with corresponding weighting weights according to the distance between its corresponding sample object and the sample acquisition center point, when training the 3D object detection model, the 3D object detection model can be trained according to the weighting weights of each sample data, so as to improve the detection rate of the 3D object candidate box of distant objects in the trained 3D object detection model according to the designed weighting weights.

[0048] Figure 3 Schematically shows the method flow of step S22 in the 3D object detection method with a high recall rate for distant objects according to an embodiment of the present invention. Refer to Figure 3 As shown, the method can be specifically implemented as including the following steps:

[0049] Step S31: Determine the weighting weight of the point cloud data corresponding to the sample object according to the distance between each sample object in the first sample data set and the sample acquisition center point;

[0050] Step S32: Weight each point cloud data corresponding to the sample object according to the weighting weight of each point cloud data corresponding to the sample object to form a first training data set.

[0051] Step S31 is a step for determining the weighted weights of the point cloud data corresponding to each sample object in the first sample dataset. Specifically, the weighted weights of the point cloud data corresponding to each sample object are determined based on the distance between the sample object corresponding to each point cloud data and the sample acquisition center point.

[0052] As a possible implementation, when determining the weighted weights of the point cloud data corresponding to each sample object in the first sample data, the weighted weights of the sample data corresponding to each sample object can be determined according to the distance between the sample object corresponding to each point cloud data and the sample acquisition center point, and a preset quantization distance value. Among them, the specific value of the preset quantization distance value can be designed according to the requirements of the actual application scenario and the design experience of the designer. Exemplarily, in the scenario where a distant object refers to an object whose distance from the vehicle itself is greater than 60 meters and is within the laser point cloud range [-80, -80, 80, 80], when it is necessary to make the value of the designed weighted weight within the range of 1 + 60 / 80 to 1 + 80 / 80, the quantization distance value can be designed as 8; and when it is necessary to make the value of the designed weighted weight within the range of 1 + 60 / 10 to 1 + 80 / 10, the quantization distance value can be designed as 1. In this implementation, the weighted weights of the point cloud data corresponding to each sample object can be determined based on the following formula:

[0053] weight 1 = 1 + dist / (10 * Edge)

[0054] where, weight 1 is the weighted weight of the point cloud data corresponding to the sample object, dist is the distance between the sample object and the sample acquisition center point, and Edge is the preset quantization distance value. Exemplarily, taking a sample dataset with a laser point cloud range of [-80m, -80m, 80m, 80m] as an example, if the value of Edge is designed as 8, and if the distance between a certain sample object and the sample acquisition center point is 65m, then the weighted weight weight 1 of the point cloud data corresponding to this sample object is 1 + 65 / (10 * 8) = 1.8125. Designed in this way, for sample objects that are farther away from the sample acquisition center point, their weighted weights will be larger.

[0055] As another possible implementation, when determining the weighted weights of the point cloud data corresponding to each sample object in the first sample data, the sample objects can be classified according to the distances between the sample objects in the first sample data set and the sample acquisition center point. After classification, the weighted weights of the point cloud data corresponding to each sample object are determined according to the classification categories of each sample object. Among them, when classifying each sample object, the classification categories can be divided based on the distance between the sample object and the sample acquisition center point. Exemplarily, it can be divided into near-distance samples and far-distance samples. For the classification criteria between near-distance samples and far-distance samples, it can be designed that when the distance between the sample object and the sample acquisition center point is not greater than a preset distance threshold, the sample object is classified as a near-distance sample; when the distance between the sample object and the sample acquisition center point is greater than the preset distance threshold, the sample object is classified as a far-distance sample. Among them, the preset distance threshold can be designed according to the actual situation. In this implementation, the distance threshold can be designed to be 60m. At this time, in this implementation, the sample objects with the distance between the sample object and the sample acquisition center point not greater than 60m will be classified as near-distance samples, and the sample objects with the distance between the sample object and the sample acquisition center point greater than 60m will be classified as far-distance samples. At the same time, in this implementation, the weighted weights of the point cloud data corresponding to each sample object can be determined based on the following formula:

[0056]

[0057] Among them, weight 2is the weighted weight of the sample data corresponding to the sample object, N1 is the number of near-distance samples, N2 is the number of far-distance samples, classone represents the classification category of near-distance samples, and classtwo represents the classification category of far-distance samples. Exemplarily, taking the first sample dataset including 10 sample objects, with 6 near-distance samples and 4 far-distance samples as an example, the weighted weight of the point cloud data corresponding to the sample object of the near-distance sample classification category is 1.4; the weighted weight of the point cloud data corresponding to the sample object of the far-distance sample classification category is 1.6. Since usually, the data volume of far-distance samples is very small, that is, the number N2 of far-distance samples (i.e., greater than 60 meters) is much smaller than the number N1 of near-distance samples. Correspondingly, in the traditional model training method, the training frequency of far-distance samples will be very low, so that the algorithm model cannot well adapt to far-distance samples. Through this weighting method, the weighted weight of far-distance samples can be increased, so that the samples of the two categories are balanced during training, that is, the embodiment of the present invention balances the training of far-distance samples and near-distance samples as much as possible by increasing the weight of far-distance samples and decreasing the weight of near-distance samples, so that the trained algorithm model can detect distant objects as much as possible.

[0058] After determining the weighted weights corresponding to the point cloud data of each sample object in the first sample dataset, according to step S32, the point cloud data corresponding to each sample object can be weighted according to the weighted weights of the point cloud data corresponding to each sample object to form a training dataset for training the 3D object detection model. As an implementation manner, when weighting each sample data, the method of directly multiplying by the weight can be used, and then the point cloud data corresponding to each sample object in the sample dataset is weighted to obtain the first training dataset for training the 3D object detection model after weighting.

[0059] After obtaining the first training dataset, return to step S23. In step S23, the selected algorithm model is further trained based on the obtained first training dataset to obtain a 3D object detection model. When training the selected algorithm model, the voxel features or cylinder features corresponding to the point cloud data in the first training dataset are used as the input part, and the 3D object candidate box information in the first training dataset is used as the output part to train the selected algorithm model. Since each point cloud data in the first training dataset has been weighted, and through the weighting of the point cloud data of the corresponding sample objects in the first training dataset, the training weight of the distant sample objects is enhanced, making the trained algorithm model pay more attention to the distant sample objects of the sparse point cloud. Therefore, the 3D object detection model trained based on the first training dataset can effectively improve the detection rate of the 3D object candidate boxes of the distant objects, enabling the trained 3D object detection model to be relatively more adaptable to the distant objects when identifying the 3D object candidate box information based on the point cloud data, improving the recognition rate of the 3D object candidate boxes of the distant objects, and further improving the recognition rate of the distant objects of the overall algorithm.

[0060] Since the goal of the 3D object detection model in the embodiment of the present invention is to quickly and efficiently detect distant objects with a high recall rate, and the features used in object detection are relatively rough voxel / pillar-based features, there will inevitably be some false detections when using the above-trained 3D object detection model for object detection. To accurately distinguish distant objects, sparse negative sample objects in the vicinity, or false detection candidate boxes, the embodiment of the present invention will also use a more accurate feature extraction method to train a model that can more accurately classify the detection results of the 3D object detection model. Among them, the more accurate feature can be the point feature corresponding to the 3D object candidate box information (that is, a more localized and accurate feature). Taking this as an example, Figure 4 Schematically shows the training method flow of the candidate box classifier model adopted in step S13 in the 3D object detection method with a high distant object recall rate according to an embodiment of the present invention. Refer to Figure 4 As shown, the method can be specifically implemented as including the following steps:

[0061] Step S41: Obtain a second training dataset, where the second training dataset includes object information corresponding to sample objects and point features corresponding to 3D object candidate box information;

[0062] Step S42: Using the point features corresponding to the 3D object candidate box information in the second training dataset as the input, and the object information of the sample object corresponding to the 3D object candidate box information in the second training dataset as the output, train the selected classification model to obtain a candidate box classifier model that takes the point features corresponding to the 3D object candidate box information as the input and the object information of the object corresponding to the 3D object candidate box information as the output. Among them, the object information includes the object classification type, object position, object size, and object yaw angle.

[0063] In step S41, the obtained second training dataset is the dataset for training the candidate box classifier model. Therefore, in the second training dataset, it should at least include the point features corresponding to the 3D object candidate box information corresponding to the sample object, which is the input part of the candidate box classifier model, and the object information corresponding to the sample object, which is the output part of the candidate box classifier. Among them, the object information includes the object classification type, object position, object size, and object yaw angle. The object classification type corresponding to the sample object can be confirmed and obtained according to the attributes of the sample object itself, while the object position, object size, and object yaw angle corresponding to the sample object can be determined according to the 3D object candidate box information corresponding to the sample object or the point features corresponding to the sample object.

[0064] In step S42, using the point features corresponding to the 3D object candidate box information in the second training dataset as the input, and the object information of the sample object corresponding to the 3D object candidate box information in the second training dataset as the output, train the selected classification model to obtain a candidate box classifier model with high recognition accuracy.

[0065] In some embodiments, the 3D object candidate box information in the second training dataset can be the 3D object candidate box information obtained by the 3D object detection model trained by the above steps S21 to S23 based on the point cloud data corresponding to the sample object. Figure 5 Schematically shows the method flow of obtaining the second training dataset in step S41 of the 3D object detection method with high far - away object recall rate according to an embodiment of the present invention. Refer to Figure 5 As shown, the method can be specifically implemented as including the following steps:

[0066] Step S51: Obtain a second sample dataset, where the second sample dataset includes at least one sample object and second sample data corresponding to the sample object. Among them, the second sample data includes point cloud data corresponding to the sample object and object information;

[0067] Step S52: Obtain 3D object candidate box information corresponding to each sample object in the second sample data set according to the second sample data set and the pre-trained 3D object detection model, and determine the point features corresponding to the 3D object candidate box information to form a second training data set.

[0068] In step S51, the obtained second sample data set is a sample data set used to form a second training data set based on the trained 3D object detection model. Therefore, in the second sample data set, it should at least include second sample data corresponding to the sample object, which is the input part of the pre-trained 3D object detection model. Among them, the second sample data includes point cloud data corresponding to the sample object. At the same time, since the formed second training data set also contains the output part for training the selected classification model, the second sample data also includes object information corresponding to the sample object.

[0069] In step S52, it is necessary to first obtain 3D object candidate box information based on the second sample data set and the 3D object detection model trained through the above steps S21 to S23 in step S12. Then, based on the obtained 3D object candidate box information, the corresponding 3D object candidate box is determined, and the point cloud data corresponding to the 3D object candidate box is determined, thereby forming a second training data set including the point features corresponding to the 3D object candidate box information and the object classification type of the sample object corresponding to the 3D object candidate box information. With such a design, the 3D object candidate box information obtained by the 3D object detection model trained through the above steps S21 to S23 with a higher recognition rate for distant objects can be used as the input part of the second training data set. On the one hand, it can make the trained candidate box classifier model have a better effect on the recall of distant objects when used in conjunction with the 3D object detection model. On the other hand, since the output result of the 3D object detection model trained through the above steps S21 to S23 with a higher recognition rate for distant objects can make the extracted localized features more accurate, using this data as the input part of the training data of the candidate box classifier model can make the recognition accuracy of the trained candidate box classifier model higher. It should be noted that the point cloud data corresponding to the 3D object candidate box needs to correspond to the sample object in the second sample data set. If there is a situation where the point cloud data corresponding to the 3D object candidate box cannot correspond to the sample object in the second sample data set, the corresponding data is not used. Different from the general candidate box classifier training method, by first using the trained 3D object detection model to process the point cloud data to obtain 3D object candidate box information, the candidate box classifier model obtained by using these 3D object candidate box information and the point cloud data corresponding to these 3D object candidate box information as the input part for training will be more matched with the 3D object detection model obtained in the training of steps S21 to S23 during recognition. Thus, in actual use, the distant object recall rate of the overall method can be effectively improved.

[0070] The 3D object detection method with a high distant object recall rate provided by the embodiments of the present invention uses the voxel feature or cylinder feature corresponding to the point cloud data as the input part of the 3D object detection model to obtain 3D object candidate box information, and uses the point features of the point cloud corresponding to the 3D object candidate box information as the input part of the candidate box classifier model for object classification and recognition. With such a design, the detection rate of the candidate box can be effectively improved by using a rough detection method, and at the same time, the accuracy of object recognition can be effectively improved by using a fine detection method. The combination of the two can effectively improve the object recall rate of 3D object detection.

[0071] Moreover, when training the 3D object detection model, the sample data corresponding to each sample object in the sample dataset can also be weighted based on the distance between each sample object and the sample acquisition center point. Training the 3D object detection model with the training dataset obtained after such weighting processing can effectively improve the detection rate of candidate boxes for distant objects when the 3D object detection model is actually used. At the same time, when training the candidate box classifier model, the sample dataset is processed in combination with the previously trained 3D object detection model, and the candidate box classifier model is trained based on the candidate boxes output by the 3D object detection model and the data in the sample dataset, which can make the obtained candidate box classifier model more adaptable to the 3D object detection model and can further improve the recall rate and recognition accuracy of distant object detection.

[0072] Figure 6 Schematically shows a 3D object detection device with a high distant object recall rate according to an embodiment of the present invention, as Figure 6 shown, the device includes:

[0073] A point cloud data acquisition module 1 for acquiring point cloud data;

[0074] A candidate box acquisition module 2 for inputting the voxel features or column features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model;

[0075] An object recognition module 3 for inputting point features corresponding to the 3D object candidate box information into a pre-trained candidate box classifier model to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

[0076] Among them, the 3D object detection model adopted by the candidate box acquisition module 2 is trained based on the following method:

[0077] Obtain a first sample dataset, which includes a sample acquisition center, at least one sample object, and first sample data corresponding to the sample object. Among them, the first sample data includes point cloud data corresponding to the sample object and 3D object candidate box information corresponding to the sample object;

[0078] Weight the first sample data according to the distance between each sample object and the sample acquisition center in the first sample dataset to form a first training dataset, where the first training dataset includes weighted point cloud data corresponding to the sample object and 3D object candidate box information corresponding to the sample object;

[0079] Using the voxel features or cylinder features corresponding to the point cloud data in the first training dataset as input and the 3D object candidate box information in the first training dataset as output, train the selected algorithm model to obtain a 3D object detection model that takes the point cloud data with voxel features or cylinder features as input and the candidate box information of the 3D object as output.

[0080] The candidate box classifier model adopted by the object recognition module 3 is trained based on the following method:

[0081] Obtain a second training dataset, where the second training dataset includes object information corresponding to the sample object, 3D object candidate box information, and point features corresponding to the 3D object candidate box information.

[0082] Using the 3D object candidate box information and the point features corresponding to the 3D object candidate box information in the second training dataset as input and the object information of the sample object corresponding to the 3D object candidate box information in the second training dataset as output, train the selected classification model to obtain a candidate box classifier model that takes the point features corresponding to the 3D object candidate box information as input and the object information of the object corresponding to the 3D object candidate box information as output, where the object information includes object classification type, object position, object size, and object yaw angle.

[0083] The second training dataset is obtained based on the following method:

[0084] Obtain a second sample dataset, where the second sample dataset includes at least one sample object and second sample data corresponding to the sample object, and the second sample data includes point cloud data corresponding to the sample object.

[0085] According to the second sample dataset and the pre-trained 3D object detection model, obtain the 3D object candidate box information corresponding to each sample object in the second sample dataset to form the second training dataset.

[0086] It should be noted that the implementation process and implementation principle of the 3D object detection device with a high recall rate for distant objects in the embodiments of the present invention can be specifically referred to the corresponding descriptions in the above embodiments of the 3D object detection method with a high recall rate for distant objects. For example, the corresponding processes such as obtaining candidate boxes and performing object recognition and classification on the candidate boxes in the method embodiment part are not described in detail here. Exemplarily, the 3D object detection device with a high recall rate for distant objects in the embodiments of the present invention can be any intelligent device with a processor, including but not limited to computers, smartphones, personal computers, robots, cloud servers, etc.

[0087] Figure 7 Schematically shows a mobile tool according to an embodiment of the present invention, such as Figure 7As shown, the mobile tool includes:

[0088] A fuselage 41;

[0089] A 3D object detection device 42 with a high recall rate for distant objects, which is arranged on the fuselage 41 and is used to obtain the perception result of the current point cloud scene at the current moment.

[0090] Among them, the specific implementation process and principle of the 3D object detection device 42 with a high recall rate for distant objects can be specifically referred to the corresponding description of the above embodiments, and will not be elaborated here. It should be noted that the mobile tool in the embodiments of the present invention can be an unmanned cleaning vehicle, an unmanned sweeper, a cleaning robot, etc. with an automatic cleaning function.

[0091] The "mobile tool" referred to in the embodiments of the present invention can be a vehicle of the L0-L5 autonomous driving technology level formulated by the Society of Automotive Engineers International (SAE International) or the Chinese national standard "Automotive Driving Automation Classification".

[0092] Exemplarily, the mobile tool can be a vehicle device or a robot device with the following various functions:

[0093] (1) A manned function, such as a family car, a bus, etc.;

[0094] (2) A cargo-carrying function, such as a general freight truck, a van, a trailer, an enclosed truck, a tank truck, a flatbed truck, a container truck, a dump truck, a special structure truck, etc.;

[0095] (3) A tool function, such as a logistics distribution vehicle, an automatic guided vehicle (AGV), a patrol vehicle, a crane, a hoist, an excavator, a bulldozer, a forklift, a roller, a loader, an off-road engineering vehicle, an armored engineering vehicle, a sewage treatment vehicle, a sanitation vehicle, a vacuum cleaner, a floor washer, a sprinkler, a floor sweeping robot, a food delivery robot, a shopping guide robot, a lawn mower, a golf cart, etc.;

[0096] (4) An entertainment function, such as an entertainment vehicle, an automatic driving device in a playground, a segway, etc.;

[0097] (5) A special rescue function, such as a fire truck, an ambulance, an electric power repair vehicle, an engineering rescue vehicle, etc.

[0098] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the 3D object detection method with a high recall rate of distant objects in any of the above embodiments of the present invention.

[0099] In some embodiments, the embodiments of the present invention further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the 3D object detection method with a high recall rate of distant objects in any of the above embodiments.

[0100] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the 3D object detection method with a high recall rate of distant objects in any of the above embodiments.

[0101] In some embodiments, the embodiments of the present invention further provide a storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, the 3D object detection method with a high recall rate of distant objects in any of the above embodiments is implemented.

[0102] Figure 8 is a schematic diagram of the hardware structure of an electronic device for executing the 3D object detection method with a high recall rate of distant objects provided by another embodiment of the present invention, as Figure 8 shown, the device includes:

[0103] One or more processors 710 and a memory 720, Figure 8 Taking one processor 710 as an example.

[0104] The device for executing the 3D object detection method for recalling distant objects may further include: an input device 730 and an output device 740.

[0105] The processor 710, the memory 720, the input device 730, and the output device 740 may be connected through a bus or other means, Figure 8 Taking connection through a bus as an example.

[0106] The memory 720, being a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the 3D object detection method with a high recall rate for distant objects in the embodiments of the present invention. The processor 710 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 720, that is, to implement the 3D object detection method with a high recall rate for distant objects in the above method embodiments.

[0107] The memory 720 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the 3D object detection method with a high recall rate for distant objects, etc. In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 720 may optionally include a memory remotely disposed relative to the processor 710, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0108] The input device 730 can receive input digital or character information, and generate signals related to user settings and function control of the image processing device. The output device 740 may include display devices such as a display screen.

[0109] The one or more modules are stored in the memory 720, and when executed by the one or more processors 710, execute the 3D object detection method with a high recall rate for distant objects in any of the above method embodiments.

[0110] The above product can execute the method provided by the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present invention.

[0111] The electronic device in the embodiments of the present invention exists in various forms, including but not limited to:

[0112] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.

[0113] (2) Ultra-mobile personal computer devices: These devices fall within the category of personal computers, have computing and processing capabilities, and generally also possess the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc., such as the iPad.

[0114] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players (such as the iPod), handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.

[0115] (4) Servers: Devices that provide computing services. The composition of a server includes a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but due to the need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0116] (5) Other electronic devices with data interaction functions.

[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0119] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A 3D object detection method with a high recall rate for distant objects, characterized in that, it includes: Obtain point cloud data; Input the voxel features or cylinder features corresponding to the point cloud data into a pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model; Obtain point features corresponding to the 3D object candidate box information and input them into a pre-trained candidate box classifier model to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

2. The method according to claim 1, characterized in that, the 3D object detection model is pre-trained in the following manner: Obtain a first sample data set, which includes a sample acquisition center, at least one sample object, and first sample data corresponding to the sample object. Among them, the first sample data includes point cloud data corresponding to the sample object and 3D object candidate box information corresponding to the sample object; Weight the first sample data according to the distance between each sample object in the first sample data set and the sample acquisition center to form a first training data set, where the first training data set includes the weighted point cloud data corresponding to the sample object and the 3D object candidate box information corresponding to the sample object; Use the voxel features or cylinder features corresponding to the point cloud data in the first training data set as input, and the 3D object candidate box information in the first training data set as output, and train the selected algorithm model to obtain a 3D object detection model with voxel features or cylinder feature point cloud data as input and 3D object candidate box information as output.

3. The method according to claim 2, characterized in that, the step of weighting the first sample data according to the distance between each sample object in the first sample data set and the sample acquisition center point to form a first training data set includes: Determine the weighting weight of the point cloud data corresponding to the sample object according to the distance between each sample object in the first sample data set and the sample acquisition center point; Weight each point cloud data corresponding to the sample object according to the weighting weight of each point cloud data corresponding to the sample object to form a first training data set.

4. The method according to claim 3, characterized in that, the step of determining the weighting weight of the point cloud data corresponding to the sample object according to the distance between each sample object in the first sample data set and the sample acquisition center point includes: Classify the sample objects according to the distance between each sample object in the first sample data set and the sample acquisition center point; Determine the weighting weight of the point cloud data corresponding to the sample object according to the category of each sample object's belonging classification.

5. The method according to claim 3, characterized in that, the step of determining the weighting weight of the point cloud data corresponding to the sample object according to the distance between each sample object in the first sample data set and the sample acquisition center point includes: Determine the weighting weight of the point cloud data corresponding to the sample object according to the distance between each sample object in the first sample data set and the sample acquisition center point and a preset quantization distance value.

6. The method according to claim 2, characterized in that, The candidate box classifier model is pre-trained in the following manner: Obtain a second training dataset, where the second training dataset includes object information corresponding to sample objects and point features corresponding to 3D object candidate box information; Using the point features corresponding to the 3D object candidate box information in the second training dataset as input and the object information of the sample objects corresponding to the 3D object candidate box information in the second training dataset as output, train the selected classification model to obtain a candidate box classifier model with the point features corresponding to the 3D object candidate box information as input and the object information of the objects corresponding to the 3D object candidate box information as output, where the object information includes object classification type, object position, object size, and object yaw angle.

7. The method according to claim 6, wherein, the obtaining of the second training dataset includes: Obtain a second sample dataset, where the second sample dataset includes at least one sample object and second sample data corresponding to the sample object, and the second sample data includes point cloud data corresponding to the sample object and object information; According to the second sample dataset and the pre-trained 3D object detection model, obtain 3D object candidate box information corresponding to each sample object in the second sample dataset, and determine the point features corresponding to the 3D object candidate box information to form a second training dataset.

8. The method according to claim 1, wherein, the confidence threshold of the 3D object detection model is set to 0.

2.

9. A 3D object detection device with a high recall rate for distant objects, wherein, it includes: A point cloud data acquisition module for acquiring point cloud data; A candidate box acquisition module for inputting the voxel features or cylinder features corresponding to the point cloud data into the pre-trained 3D object detection model to obtain 3D object candidate box information output by the 3D object detection model; An object recognition module for inputting the point features corresponding to the 3D object candidate box information into the pre-trained candidate box classifier model to determine the recognition result of the point cloud data according to the output result of the candidate box classifier model.

10. An electronic device, wherein, it includes: A memory for storing executable instructions; and A processor for executing the executable instructions stored in the memory, and the executable instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 8.

11. A mobile tool, wherein, the mobile tool includes: the device according to claim 9 or the electronic device according to claim 10.

12. A storage medium having a computer program stored thereon, wherein, the program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.