Three-dimensional point cloud detection model training method and device, electronic equipment and storage medium
Through the method of automated pre-labeling and parameter adjustment, the problems of low efficiency and low accuracy of manual labeling in the training of 3D point cloud target detection models are solved, and high-quality automated training and the ability of the model to adapt to specific scenarios are achieved.
Patent Information
- Application Number
- CN202410345938.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing 3D point cloud target detection model training, dataset labeling and data processing rely entirely on manual labor, which is inefficient and easily affected by human factors, resulting in low model accuracy.
By acquiring continuous multi-frame point cloud data and trajectory pose data, pre-labeling and filtering are performed to obtain target detection results, and the error back propagation algorithm is used to adjust the model parameters to achieve automated training.
It improves the quality of automatically labeled data for model training, enhances the model's detection performance and obstacle marking accuracy in specific scenarios, and possesses strong scene generalization capabilities.
Smart Images

Figure CN120708182A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of model training technology, and in particular to a three-dimensional point cloud detection model training method, device, electronic device and storage medium. Background Art
[0002] When a vehicle is driving on the road, there may be pedestrians or other vehicles around it. The movement trajectories of these obstacles are uncertain and may change at any time, affecting the driving of the vehicle. These obstacles can be labeled and identified through 3D point cloud data. Currently, the training of 3D point cloud target detection models first requires the construction of a 3D point cloud target detection dataset, and then the point cloud data to be labeled is manually labeled before training can be carried out. In this way, the labeling and data processing of the dataset are all done manually, which is not only inefficient but also easily affected by human factors, resulting in low accuracy of the trained model. Summary of the Invention
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present application provide a three-dimensional point cloud detection model training method, device, electronic device and storage medium to solve the problem that the labeling and data processing of data sets in the prior art are all completed manually, which is not only inefficient but also easily affected by human factors, thereby resulting in low accuracy of the trained model.
[0004] In order to achieve the above objectives, the technical solutions provided in the embodiments of the present application are as follows:
[0005] In a first aspect, an embodiment of the present application provides a three-dimensional point cloud detection model training method, the method comprising: obtaining at least one original data packet, each original data packet comprising: point cloud data and trajectory pose data corresponding to a plurality of consecutive frames;
[0006] Pre-labeling the at least one original data packet to obtain a pre-labeling result corresponding to each original data packet, wherein the pre-labeling result is used to indicate obstacle information included in each frame of point cloud data;
[0007] Filtering the pre-brush marking results according to a preset classification method and specified obstacle information to obtain a target detection result corresponding to the original data packet;
[0008] The original data packet and the target detection result are used to train an initial three-dimensional point cloud detection model, and the specified parameters in the initial three-dimensional point cloud detection model are adjusted through an error back propagation algorithm to obtain a target point cloud detection model.
[0009] As an optional implementation manner, in the first aspect of the embodiment of the present application, pre-marking the at least one original data packet to obtain a pre-refresh marking result corresponding to the original data packet includes:
[0010] For each of the original data packets, determining a plurality of detection frames from each frame of point cloud data, each detection frame being used to indicate an obstacle;
[0011] Determine a 3D detection result corresponding to each detection frame, wherein the 3D detection result includes the size, position, angle, type, and confidence level of the detection frame;
[0012] The pre-brush annotation result corresponding to the original data packet is determined through a multi-target tracking algorithm according to the three-dimensional detection result corresponding to each frame of point cloud data and the trajectory posture data.
[0013] As an optional implementation, in the first aspect of the embodiment of the present application, determining the pre-brush annotation result corresponding to the original data packet based on the three-dimensional detection result corresponding to each frame of point cloud data and the trajectory pose data using a multi-target tracking algorithm includes:
[0014] Calculate the similarity between each detection box in the current frame and the trajectory in the previous frame;
[0015] Matching each detection frame with the trajectory according to the similarity to obtain a matching result corresponding to each detection frame;
[0016] According to the matching result, multiple detection frames are processed to obtain the pre-refresh marking result corresponding to the original data packet.
[0017] As an optional implementation, in the first aspect of the embodiment of the present application, filtering the pre-brush annotation results according to the preset classification method and the specified obstacle information to obtain the target detection result corresponding to the original data packet includes:
[0018] Classify the pre-brush marking results according to obstacle categories to obtain pre-brush marking results corresponding to each specified category;
[0019] The pre-brush marking results are filtered according to the spacing and confidence included in the pre-brush marking results to obtain the target detection result corresponding to the original data packet.
[0020] As an optional implementation manner, in the first aspect of the embodiment of the present application, filtering the pre-brush annotation results according to the spacing and confidence included in the pre-brush annotation results to obtain the target detection result corresponding to the original data packet includes:
[0021] Determine the spacing and confidence level corresponding to each pre-brush annotation result, where the spacing is used to indicate the distance between the obstacle and the vehicle;
[0022] The pre-brush marking result whose spacing is smaller than a preset distance threshold and whose confidence is greater than a preset confidence threshold is determined as the target detection result.
[0023] As an optional implementation, in the first aspect of the embodiments of the present application, adjusting the specified parameters of the initial three-dimensional point cloud detection model by using an error back propagation algorithm to obtain a target point cloud detection model includes:
[0024] Freezing first parameters in the initial three-dimensional point cloud detection model and adjusting second parameters using the error back propagation algorithm, wherein the first parameters include early layer parameters of the initial three-dimensional point cloud detection model, the second parameters include late layer parameters of the initial three-dimensional point cloud detection model, and a learning rate of the first parameters is smaller than a learning rate of the second parameters;
[0025] The initial three-dimensional point cloud detection model after parameter adjustment is trained to obtain the target point cloud detection model.
[0026] As an optional implementation manner, in the first aspect of the embodiments of the present application, the method further includes:
[0027] Get the evaluation dataset;
[0028] Inputting the evaluation data set into the target point cloud detection model to obtain an evaluation detection result;
[0029] The evaluation detection result is matched and calculated with the evaluation true value to obtain the evaluation result of the target point cloud detection model, and the evaluation result includes: the accuracy and recall rate corresponding to each category.
[0030] In a second aspect, an embodiment of the present application provides a three-dimensional point cloud detection model training device, the three-dimensional point cloud detection model training device comprising: an acquisition module for acquiring at least one original data packet, each original data packet comprising: point cloud data and trajectory pose data corresponding to a plurality of consecutive frames;
[0031] a 3D point cloud object detection pre-annotation module, configured to pre-annotate the at least one original data packet to obtain a pre-annotation result corresponding to each original data packet, wherein the pre-annotation result is used to indicate obstacle information included in each frame of point cloud data;
[0032] A training data cleaning and packaging module is used to filter and screen the pre-brush annotation results according to a preset classification method and specified obstacle information to obtain a target detection result corresponding to the original data packet;
[0033] An automated training module is used to train an initial three-dimensional point cloud detection model using the original data packet and the target detection result, and to adjust specified parameters in the initial three-dimensional point cloud detection model through an error back propagation algorithm to obtain a target point cloud detection model.
[0034] As an optional implementation, in the second aspect of the embodiment of the present application, the three-dimensional point cloud object detection pre-annotation module is specifically configured to determine, for each original data packet, a plurality of detection boxes from each frame of point cloud data, each detection box being configured to indicate an obstacle;
[0035] The 3D point cloud object detection pre-annotation module is specifically used to determine the 3D detection result corresponding to each detection frame, wherein the 3D detection result includes the size, position, angle, type and confidence of the detection frame;
[0036] The three-dimensional point cloud target detection pre-brush labeling module is specifically used to determine the pre-brush labeling result corresponding to the original data packet based on the three-dimensional detection result corresponding to each frame of point cloud data and the trajectory posture data through a multi-target tracking algorithm.
[0037] As an optional implementation, in the second aspect of the embodiment of the present application, the three-dimensional point cloud object detection pre-annotation module is specifically used to calculate the similarity between each detection box in the current frame and the trajectory in the previous frame;
[0038] The three-dimensional point cloud object detection pre-annotation module is specifically used to match each detection frame with the trajectory according to the similarity to obtain a matching result corresponding to each detection frame;
[0039] The three-dimensional point cloud target detection pre-brush annotation module is specifically used to process multiple detection frames according to the matching results to obtain the pre-brush annotation results corresponding to the original data packet.
[0040] As an optional implementation, in the second aspect of the embodiment of the present application, the training data cleaning and packaging module is specifically used to classify the pre-brush annotation results according to obstacle categories to obtain pre-brush annotation results corresponding to each specified category;
[0041] The training data cleaning and packaging module is specifically used to filter the pre-brush annotation results according to the spacing and confidence included in the pre-brush annotation results to obtain the target detection results corresponding to the original data packets.
[0042] As an optional implementation, in the second aspect of the embodiment of the present application, the training data cleaning and packaging module is specifically used to determine the spacing and confidence level corresponding to each pre-brush annotation result, where the spacing is used to indicate the distance between the obstacle and the vehicle;
[0043] The training data cleaning and packaging module is specifically configured to determine the pre-brush labeling results whose spacing is less than a preset distance threshold and whose confidence is greater than a preset confidence threshold as the target detection results.
[0044] As an optional implementation, in a second aspect of the embodiments of the present application, the fine-tuning training module is specifically used to freeze first parameters in the initial three-dimensional point cloud detection model and adjust second parameters through the error back propagation algorithm, wherein the first parameters include early layer parameters of the initial three-dimensional point cloud detection model, the second parameters include late layer parameters of the initial three-dimensional point cloud detection model, and the learning rate of the first parameters is smaller than the learning rate of the second parameters;
[0045] The fine-tuning training module is specifically used to train the initial three-dimensional point cloud detection model after parameter adjustment to obtain the target point cloud detection model.
[0046] As an optional implementation, in the second aspect of the embodiment of the present application, the three-dimensional point cloud detection model training device further includes: a model evaluation module;
[0047] The model evaluation module is used to obtain an evaluation data set;
[0048] The model evaluation module is further configured to input the evaluation data set into the target point cloud detection model to obtain an evaluation detection result;
[0049] The model evaluation module is further used to match and calculate the evaluation detection results with the evaluation true values to obtain the evaluation results of the target point cloud detection model. The evaluation results include: the accuracy and recall rate corresponding to each category.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0051] a memory storing executable program code;
[0052] a processor connected to the memory;
[0053] The processor calls the executable program code stored in the memory to execute the three-dimensional point cloud detection model training method in the first aspect of the embodiment of the present application.
[0054] In a fourth aspect, embodiments of the present application provide a computer-readable storage medium storing a computer program that causes a computer to execute the three-dimensional point cloud detection model training method of the first aspect of the embodiments of the present application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, or an optical disk.
[0055] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute part or all of the steps of any one of the methods of the first aspect.
[0056] In a sixth aspect, an embodiment of the present application provides an application publishing platform, which is used to publish a computer program product, wherein when the computer program product runs on a computer, the computer executes part or all of the steps of any one of the methods of the first aspect.
[0057] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0058] The embodiments of the present application provide a three-dimensional point cloud detection model training method, device, electronic device and storage medium, which obtain at least one original data packet, each of which includes: point cloud data and trajectory pose data corresponding to multiple consecutive frames; pre-label the at least one original data packet to obtain a pre-brush labeling result corresponding to the original data packet, and the pre-brush labeling result is used to indicate the obstacle information included in each frame of point cloud data; according to a preset classification method and specified obstacle information, the pre-brush labeling result is filtered and screened to obtain a target detection result corresponding to the original data packet; the original data packet and the target detection result are used to train an initial three-dimensional point cloud detection model, and the specified parameters in the initial three-dimensional point cloud detection model are adjusted through an error back propagation algorithm to obtain a target point cloud detection model. In this solution, no manual operation is required. The obstacle annotation information can be obtained directly through a large amount of point cloud data and trajectory pose data, and the annotation information will also be filtered. This can ensure that the annotation data for automated model training has higher quality. After the model training, the model parameters will be adjusted and further trained. This will allow the model to better adapt to the data distribution and characteristics of specific tasks, thereby improving the detection performance of the model in specific scenarios, making the model have stronger scene generalization capabilities, and also improving the model's accuracy in obstacle marking. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0061] Figure 1 This is a flow chart of a 3D point cloud detection model training method provided in an embodiment of the present application. Figure 1 ;
[0062] Figure 2 This is a flow chart of a 3D point cloud detection model training method provided in an embodiment of the present application. Figure 2 ;
[0063] Figure 3 Schematic diagram of a three-dimensional point cloud detection model training device provided in an embodiment of the present application;
[0064] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be noted that, in the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. Obviously, the described embodiments are part of the embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0066] The terms "first" and "second" and the like in the description and claims of the present invention are used to distinguish different objects rather than to describe a specific order of the objects.
[0067] The terms "including" and "having" and any variations thereof in the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatus.
[0068] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0069] Training a 3D point cloud object detection model first requires constructing a 3D point cloud object detection dataset. This involves extracting frames from the collected data based on the scene and content of the acquisition to extract the point cloud data to form a labeled dataset. This labeled dataset is then manually labeled and, after completion, undergoes quality inspection to obtain a training dataset. Once the labeled training dataset is obtained, manual data cleaning and dataset packaging are required. Finally, the packaged training dataset is loaded onto the training platform to train the 3D point cloud object detection model. This results in a completely manually labeled dataset, which is costly, inefficient, and time-consuming, making it difficult to respond to urgent object detection needs. Furthermore, manual processing of the labeled dataset is susceptible to human factors, and the quality of the processed dataset can also affect the results of model training.
[0070] In order to solve some or all of the above-mentioned technical problems, the embodiments of the present application provide a three-dimensional point cloud detection model training method, device, electronic device and storage medium, obtain at least one original data packet, each original data packet includes: point cloud data and trajectory pose data corresponding to multiple consecutive frames; pre-label the at least one original data packet to obtain a pre-brush labeling result corresponding to each original data packet, and the pre-brush labeling result is used to indicate the obstacle information included in each frame of point cloud data; according to a preset classification method and specified obstacle information, the pre-brush labeling result is filtered and screened to obtain a target detection result corresponding to the original data packet; the original data packet and the target detection result are used to train the initial three-dimensional point cloud detection model, and the specified parameters in the initial three-dimensional point cloud detection model are adjusted through the error back propagation algorithm to obtain a target point cloud detection model. In this solution, no manual operation is required. The obstacle annotation information can be obtained directly through a large amount of point cloud data and trajectory pose data, and the annotation information will also be filtered. This can ensure that the annotation data for automated model training has higher quality. After the model training, the model parameters will be adjusted and trained in one step. This allows the model to better adapt to the data distribution and characteristics of specific tasks, thereby improving the detection performance of the model in specific scenarios, making the model have stronger scene generalization capabilities, and also improving the accuracy of the model for obstacle marking.
[0071] like Figure 1 As shown, Figure 1A flowchart of a three-dimensional point cloud detection model training method provided in an embodiment of the present application, the method may include the following steps:
[0072] 101. Obtain at least one original data packet.
[0073] In an embodiment of the present application, each original data packet may include: point cloud data and trajectory pose data corresponding to multiple consecutive frames.
[0074] It should be noted that the original data packet is obtained through some sensors on the collection vehicle. The data collected by the collection vehicle mainly includes various types of data continuously recorded by sensors such as lidar, camera, odometer, inertial measurement unit (IMU) over a period of time. The data packet collected in this way is called a clip, which is the basic data unit for automated training.
[0075] It's important to note that point cloud data is 3D data of urban road scenes collected by a collection vehicle using LiDAR. This data is represented as a point cloud. A point cloud is a collection of three-dimensional points, generated by lasers reflected from obstacle surfaces scanned by the LiDAR. The coordinates of these points are (x, y, z). A large number of these points form a point cloud. As the collection vehicle drives along urban roads, the LiDAR continuously collects point cloud data, generating multiple frames of continuous point cloud data over a period of time. The trajectory pose data represents the vehicle's trajectory during this period.
[0076] 102. Pre-label at least one original data packet to obtain a pre-labeling result corresponding to each original data packet.
[0077] In an embodiment of the present application, after the original data packet is obtained, the original data packet can be pre-labeled, that is, all obstacles are determined from the point cloud data included in the original data packet and labeled.
[0078] The pre-refresh annotation result can be used to indicate the obstacle information included in each frame of point cloud data.
[0079] It should be noted that the obstacles determined by point cloud data may include many types, that is, any object on the street or road can be counted as an obstacle, such as cars, trucks, pedestrians, trash cans, bus stops, etc., which will be marked. The specific marking form is a detection frame, that is, the obstacle is circled by the detection frame. The pre-brush marking result can be represented in the form of a detection frame, which may include: the three-dimensional coordinates of the center point of the detection frame, the length, width, and height values of the detection frame, the orientation angle of the detection frame, the detection target type (type) information and the confidence score (confidence score) information of the detection frame, where the detection target type is the type of obstacle, and the confidence score is the probability that the inferred obstacle detection target belongs to this predicted category when the 3D detection model infers each detection frame, because the model also uses the principle of probability statistics. This probability is the confidence score, which is generally a value between 0 and 1.
[0080] 103. According to the preset classification method and the specified obstacle information, the pre-brush marking results are filtered and screened to obtain the target detection results corresponding to the original data packet.
[0081] In an embodiment of the present application, after all obstacles are indicated by pre-labeling results, these obstacles can be filtered and screened to select obstacles that meet the requirements, that is, the target detection results corresponding to the original data packets. The target detection results are the detection results for subsequent model training. Specifically, since obstacles include many types, filtering and screening can be performed based on the classification of the obstacle types, and filtering and screening can also be performed based on specified obstacle information.
[0082] In some embodiments, the pre-annotation results of each frame are cleaned, and the pre-annotation results are filtered according to the classification of the automated training model to obtain detection results of the specified category; secondly, the pre-annotation results are filtered according to the distance and confidence score between the detection target and the vehicle to obtain detection results with higher confidence within different distance ranges, thereby ensuring that the annotation data of the automated training has higher quality.
[0083] 104. Use the original data packet and the target detection results to train the initial three-dimensional point cloud detection model, and adjust the specified parameters in the initial three-dimensional point cloud detection model through the error back propagation algorithm to obtain the target point cloud detection model.
[0084] In an embodiment of the present application, after determining the target detection result that meets the requirements, training can be carried out. Specifically, the initial three-dimensional point cloud detection model is trained with the original data packet as input and the target detection result as output; and further, the specified parameters can be adjusted based on the initial three-dimensional point cloud detection model. The initial three-dimensional point cloud detection model (pretrain model) is trained in a large number of different scenarios and has strong scene generalization capabilities, which can improve the detection capabilities of the detection model in various scenarios. Parameter fine-tuning (fine-tune) training is based on the initial three-dimensional point cloud detection model, freezing part of the model parameters and fine-tuning the remaining parameters of the model. Fine-tune allows the model to better adapt to the data distribution and characteristics of specific tasks, so as to improve the detection performance of the model in specific scenarios, thereby obtaining a target point cloud detection model.
[0085] In some embodiments, before training the original data packet and target detection results, the target detection results can also be batch packaged. The target detection results of all frames in the original data packet are packaged into a binary file (pkl) in units of original data packet clip. During the automated training process, this binary file needs to be input to obtain the annotation information of each frame to complete the model training. Loading this detection result binary file can further improve the efficiency of model training.
[0086] In some embodiments, when performing model training, a specified number of original data packet clips and packaged detection result pkl files are combined into an automated training set (training dataset). This automated training set can be trained to obtain an initial three-dimensional point cloud detection model. In addition, multiple different automated training sets can be input simultaneously for training according to task requirements to obtain an initial three-dimensional point cloud detection model (pretrain model).
[0087] Among them, the organizational structure of each data packet is the same. One directory stores the original several-frame point cloud data files, and the other directory stores the pre-flash detection result files corresponding to these point cloud data files. The difference between the data packets is the number of frames. For example, some data packets have 150 frames of point cloud data files and corresponding pre-flash detection result files, and some data packets only have 140 frames of point cloud data files and corresponding pre-flash detection result files.
[0088] Among them, the task requirements are generally that when the collection vehicle collects certain scenes, such as snowy and rainy scenes, the data of such scenes are needed to participate in automated training, so that the 3D target detection model can adapt to these rare scenes. At this time, it is necessary to find a data set of such scenes to add to the automated training process, so that the trained model can better adapt to the obstacle detection tasks in these scenes.
[0089] A three-dimensional point cloud detection model training method provided in an embodiment of the present application can obtain obstacle annotation information directly through a large amount of point cloud data and trajectory pose data, and can also filter the annotation information, so as to ensure that the annotation data for automated model training has higher quality. After model training, the model parameters will be adjusted and further trained, so that the model can better adapt to the data distribution and characteristics of specific tasks, thereby improving the detection performance of the model in specific scenarios, making the model have strong scene generalization capabilities, and also improving the accuracy of the model for obstacle marking.
[0090] like Figure 2 As shown, Figure 2 A flowchart of another three-dimensional point cloud detection model training method provided in an embodiment of the present application, the method may include the following steps:
[0091] 201. Obtain at least one original data packet.
[0092] In the embodiment of the present application, for the description of step 201, please refer to the detailed description of step 101 in the above embodiment, and the embodiment of the present invention will not be repeated.
[0093] 202. For each original data packet, determine multiple detection boxes from each frame of point cloud data.
[0094] In an embodiment of the present application, after obtaining the original data packet, the point cloud data in each original data packet can be processed, that is, the detection frame indicating the obstacle can be determined from each frame of point cloud data. Each frame of point cloud data can include a large number of detection frames, and each detection frame is used to indicate an obstacle.
[0095] 203. Determine the three-dimensional detection result corresponding to each detection frame.
[0096] In an embodiment of the present application, a three-dimensional detection result can be formed for each detection frame. The three-dimensional detection result can be defined as a 7-dimensional vector (x, y, z, l, w, h, heading), where x, y, z represent the three-dimensional coordinates of the center point of the detection frame, l, w, h represent the length, width, and height values of the detection frame, heading represents the orientation angle of the detection frame, and there is also information indicating the type of the detection target and the confidence score information of the detection frame.
[0097] In some embodiments, each frame of point cloud data can be voxelized to adapt to the operation of the model network, and then after operation and processing by the 3D target detection model, the detection frame of the obstacle corresponding to each frame of point cloud data and the obstacle category represented by the detection frame are output.
[0098] In some embodiments, 3D point cloud voxelization is the process of converting a 3D point cloud into a voxel grid. A voxel grid is a grid composed of a series of 3D cubes (also called voxels), each of which contains the value of one or more points. Voxelization can be used to simplify 3D point clouds, making them easier to process. In essence, it converts large amounts of point cloud data into a 3D grid, which is intended to facilitate the calculation and processing of 3D object detection models.
[0099] 204. Through the multi-target tracking algorithm, according to the three-dimensional detection results and trajectory pose data corresponding to each frame of point cloud data, determine the pre-brush annotation results corresponding to the original data packet.
[0100] It should be noted that a large number of detection frames are circled in each frame of point cloud data. The obstacles circled in each frame may be newly appeared targets, or targets that existed in the previous frame, or some targets may disappear. For example, if two objects a and b are detected in the previous frame and two objects c and d are detected in the next frame, how to match a in the previous frame with c or d in the next frame, and how to match b in the previous frame with c or d in the next frame; for example, when a cyclist meets a pedestrian, the pedestrian is blocked. For the computer, it thinks that the tracking of the pedestrian's ID has ended. After a while, the pedestrian reappears in the field of view, but the computer thinks that this is a new object and assigns a new ID. In this case, an ID exchange occurs. Similarly, the same is true when the object is blocked by other objects such as telephone poles. Therefore, a multi-target tracking algorithm is needed to associate obstacles between each frame.
[0101] In the embodiment of the present application, when using the multi-target tracking algorithm, it is also necessary to combine the trajectory posture data, and the trajectory of the collection vehicle will also affect the collection of surrounding obstacles.
[0102] In some embodiments, a multi-target tracking algorithm is used to determine the pre-refreshed annotation results corresponding to the original data packet based on the three-dimensional detection results and trajectory pose data corresponding to each frame of point cloud data. Specifically, this may include: calculating the similarity between each detection frame in the current frame and the trajectory in the previous frame; matching each detection frame with the trajectory based on the similarity to obtain the matching result corresponding to each detection frame; and processing multiple detection frames based on the matching results to obtain the pre-refreshed annotation results corresponding to the original data packet.
[0103] It should be noted that the matching calculation in 3D multi-target tracking refers to the process of matching the detection box in the current frame with the trajectory in the previous frame. The purpose of the matching calculation is to determine which trajectory the detection box in the current frame comes from, or whether it is a new target.
[0104] First, we need to calculate the similarity measure between the detection box in the current frame and the trajectory in the previous frame. The similarity metric used here is 3D Intersection over Union (IoU), which is the most commonly used measure of 3D bounding box similarity. The IoU is the ratio of the intersection and union of the detection box and the trajectory.
[0105] Then, the Hungarian algorithm can be used for matching: the detection box in the current frame is matched with the tracklet in the previous frame using the Hungarian algorithm. The Hungarian algorithm is an optimal matching algorithm that can find the best matching pair.
[0106] Finally, the matching results are processed. Specifically, the unmatched detection frames can be deleted, which may be new targets; the matched tracks can be updated by adding the detection frames in the current frame to the tracks; and the unmatched tracks can be processed, which may have disappeared.
[0107] In the matching calculation, who matches whom depends on the value of the similarity metric. If the value of the similarity metric is large enough, the detection box in the current frame is matched with the trajectory in the previous frame. If the value of the similarity metric is too small, the detection box in the current frame is considered a new target.
[0108] Therefore, in addition to the obstacle detection box output by the 3D target detection model inference mentioned above, the pre-brush annotation results obtained in this way will also have the tracking results output by the multi-target tracking algorithm. The final representation is still the detection result of the obstacles in the scene corresponding to the collected point cloud data. Unlike the obstacle detection box output by the 3D target detection model inference alone, the detection box in this process will supplement the targets missed by some frames and carry tracking information, making the detection results more complete.
[0109] 205. Classify the pre-brush marking results according to the obstacle category to obtain the pre-brush marking results corresponding to each specified category.
[0110] In an embodiment of the present application, since each detection frame indicates an obstacle, and the types of obstacles are diverse, and not every obstacle needs to be finally marked, the pre-brush marking results can be classified according to categories, and some specified categories can be determined in advance. The specified categories can be determined based on actual needs or scene information, and all pre-brush marking results in all frames are classified according to the specified categories.
[0111] For example, the pre-brush labeling results include cars, pedestrians, cyclists, trucks, cones, bus stops, trash cans, bicycles and other obstacle targets, but the classification of the automated training model only requires the four categories of cars, pedestrians, cyclists and trucks. In this case, the pre-brush labeling results need to be filtered according to the category to obtain pre-brush labeling results containing only the four categories of cars, pedestrians, cyclists and trucks.
[0112] 206. Filter the pre-brush annotation results according to the spacing and confidence level included in the pre-brush annotation results to obtain the target detection results corresponding to the original data packet.
[0113] In an embodiment of the present application, after the pre-brush annotation results are classified by category, they can be further filtered, that is, filtered according to the distance and confidence of the obstacle target from the vehicle. When the collection vehicle is driving, the range of data collected by the sensor is limited, so it is necessary to filter the data from the two perspectives of distance and confidence to obtain the final detection results that can be used for model training.
[0114] In some embodiments, the pre-brush annotation results are filtered according to the spacing and confidence included in the pre-brush annotation results to obtain the target detection results corresponding to the original data packet, which may specifically include: determining the spacing and confidence corresponding to each pre-brush annotation result, where the spacing is used to indicate the distance between the obstacle and the vehicle; and determining the pre-brush annotation results with a spacing less than a preset distance threshold and a confidence greater than a preset confidence threshold as target detection results.
[0115] It should be noted that a preset distance threshold and a preset confidence threshold can be pre-set as filtering criteria. From all pre-brush annotation results, only pre-brush annotation results with a distance less than the preset distance threshold and a confidence greater than the preset confidence threshold are selected as target detection results. Pre-brush annotation results with a distance greater than or equal to the preset distance threshold, or a confidence less than or equal to the preset confidence threshold, can be directly deleted without subsequent training. The purpose of cleaning is to ensure that the data set input to the automated training model is of higher quality and continuously enhance the detection capability of the 3D target detection model.
[0116] For example, the preset distance threshold is 60m and the preset confidence threshold is 0.5. If a detection result of a car is 30m away from the ego vehicle and the confidence level is 0.7, then based on the distance and confidence level of this detection result, this detection result needs to be retained. If a detection result of a car is 35m away from the ego vehicle and the confidence level is 0.3, then this detection result needs to be filtered out.
[0117] 207. Use the original data packet and target detection results to train the initial 3D point cloud detection model.
[0118] In the embodiment of the present application, for the description of step 207, please refer to the detailed description of step 104 in the above embodiment, and the embodiment of the present invention will not be repeated.
[0119] 208. Freeze the first parameter in the initial three-dimensional point cloud detection model, and adjust the second parameter through an error back propagation algorithm.
[0120] The first parameter includes early layer parameters of the initial three-dimensional point cloud detection model, the second parameter includes late layer parameters of the initial three-dimensional point cloud detection model, and the learning rate of the first parameter is less than the learning rate of the second parameter.
[0121] In the embodiment of the present application, during the fine-tuning process of a 3D object detection model, the frozen parameters are typically the parameters of the early layers of the model, such as the convolutional layers and the fully connected layers. The remaining parameters are typically the parameters of the later layers of the model, such as the classification layer and the regression layer.
[0122] The method of distinguishing frozen parameters from remaining parameters is usually to control them using the learning rate. For frozen parameters, that is, the first parameters, the learning rate is usually set to 0, indicating that these parameters are not updated. For remaining parameters, that is, the second parameters, the learning rate is usually set to a smaller value, indicating that the amplitude of the update is smaller.
[0123] Fine-tuning the remaining parameters is typically accomplished using backpropagation. Backpropagation is a fundamental algorithm in deep learning that calculates the gradient of the loss function with respect to the model parameters. Using gradient descent, the model parameters can be updated to minimize the loss function.
[0124] Specifically, during fine-tuning, the parameters of the early layers of the model are usually frozen, and then only the parameters of the later layers of the model are updated. This prevents the parameters of the early layers of the model from overfitting the training data, thereby improving the generalization ability of the model.
[0125] In some embodiments, the method for fine-tuning the remaining parameters includes: first initializing the remaining parameters, then setting the learning rate, then iteratively calculating the loss function and gradients, and then updating the remaining parameters using a gradient descent algorithm, and continuously repeating the iteration and update steps until the parameters achieve the desired effect. In actual applications, the fine-tuning parameter settings can be adjusted according to specific circumstances. For example, the number of frozen parameters and the fine-tuning results can be determined based on the model structure and the quality of the training data.
[0126] The learning rate is a hyperparameter that must be set during automated training. It is one of the key parameters that influences model training performance. It controls the amplitude of model parameter updates. If the learning rate is too high, model parameters will become unstable during training, prone to oscillation or divergence. If the learning rate is too low, model training will be slow or even fail to converge. Setting the learning rate is crucial for training a better 3D object detection model.
[0127] 209. The initial 3D point cloud detection model after parameter adjustment is trained to obtain a target point cloud detection model.
[0128] In an embodiment of the present application, after the parameters of the initial three-dimensional point cloud detection model are adjusted, further training can be performed again.
[0129] In some embodiments, the parameter-adjusted initial 3D point cloud detection model can be trained using refined annotated data. This refined annotated data is data manually annotated according to 3D annotation requirements. To make the model more accurate, some data can be selected from the original data package and manually annotated again to obtain refined annotated data. Specifically, the parameter-adjusted initial 3D point cloud detection model is trained using the original data package as input and the refined annotated data as output to obtain a target point cloud detection model.
[0130] 210. Obtain the evaluation dataset.
[0131] In an embodiment of the present application, after obtaining the target point cloud detection model, it is also necessary to evaluate the target point cloud detection model, that is, to detect the quality of the target point cloud detection model. Therefore, an evaluation data set can be obtained. The evaluation data set is also point cloud data and has an evaluation true value, so that it can be compared with the results obtained by the model.
[0132] 211. Input the evaluation data set into the target point cloud detection model to obtain the evaluation detection results.
[0133] In an embodiment of the present application, the evaluation data set is input into the target point cloud detection model to obtain an evaluation detection result, which is the obstacle information of the required category determined based on the point cloud data of the evaluation data set.
[0134] 212. Match the evaluation test results with the evaluation true values to obtain the evaluation results of the target point cloud detection model.
[0135] In an embodiment of the present application, the evaluation detection results obtained by the target point cloud detection model need to be matched and calculated with the evaluation true value to detect the accuracy of the target point cloud detection model. The evaluation results may specifically include: the accuracy and recall rate corresponding to each category.
[0136] In some embodiments, when performing matching calculations on the evaluation detection results and the evaluation true values, the 3DIOU method may also be used for calculations.
[0137] Recall refers to the ratio of the number of objects detected by the model to the number of objects actually present. Recall is an important metric for measuring model detection performance. The formula for calculating recall is: Recall = Number of Detected Objects / Number of Actual Objects. A higher recall indicates that the model detects more objects and performs better.
[0138] A three-dimensional point cloud detection model training method provided in an embodiment of the present application can obtain obstacle annotation information directly through a large amount of point cloud data and trajectory pose data, and can also filter the annotation information, so as to ensure that the annotation data for automated model training has higher quality. After the model training, the model parameters will be adjusted and further trained in combination with the refined annotation data, so that the model can further better adapt to the data distribution and characteristics of specific tasks, thereby improving the detection performance of the model in specific scenarios, making the model have stronger scene generalization capabilities, and also improving the accuracy of the model for obstacle marking.
[0139] like Figure 3 As shown, an embodiment of the present application provides a three-dimensional point cloud detection model training device, which may include:
[0140] An acquisition module 301 is configured to acquire at least one original data packet, each of which includes: point cloud data and trajectory pose data corresponding to a plurality of consecutive frames;
[0141] The 3D point cloud object detection pre-annotation module 302 is configured to pre-annotate at least one original data packet to obtain a pre-annotation result corresponding to each original data packet. The pre-annotation result is used to indicate obstacle information included in each frame of point cloud data.
[0142] The training data cleaning and packaging module 303 is used to filter the pre-brush annotation results according to the preset classification method and the specified obstacle information to obtain the target detection results corresponding to the original data packet;
[0143] The automated training module 304 is used to train the initial three-dimensional point cloud detection model using the original data packet and the target detection results, and to adjust the specified parameters in the initial three-dimensional point cloud detection model through the error back propagation algorithm to obtain the target point cloud detection model.
[0144] In some embodiments, the 3D point cloud object detection pre-annotation module 302 is specifically configured to determine, for each original data packet, a plurality of detection boxes from each frame of point cloud data, each detection box being used to indicate an obstacle;
[0145] The 3D point cloud object detection pre-annotation module 302 is specifically used to determine the 3D detection result corresponding to each detection frame. The 3D detection result includes the size, position, angle, type and confidence of the detection frame;
[0146] The 3D point cloud target detection pre-annotation module 302 is specifically used to determine the pre-annotation result corresponding to the original data packet based on the 3D detection result and trajectory pose data corresponding to each frame of point cloud data through a multi-target tracking algorithm.
[0147] In some embodiments, the 3D point cloud object detection pre-annotation module 302 is specifically configured to calculate the similarity between each detection box in the current frame and the trajectory in the previous frame;
[0148] The 3D point cloud object detection pre-annotation module 302 is specifically used to match each detection frame with the trajectory based on similarity to obtain a matching result corresponding to each detection frame;
[0149] The 3D point cloud target detection pre-brush annotation module 302 is specifically used to process multiple detection frames according to the matching results to obtain pre-brush annotation results corresponding to the original data packet.
[0150] In some embodiments, the training data cleaning and packaging module 303 is specifically used to classify the pre-brush annotation results according to the obstacle category, and obtain the pre-brush annotation results corresponding to each specified category;
[0151] The training data cleaning and packaging module 303 is specifically used to filter the pre-brush annotation results according to the spacing and confidence included in the pre-brush annotation results to obtain the target detection results corresponding to the original data packets.
[0152] In some embodiments, the training data cleaning and packaging module 303 is specifically used to determine the spacing and confidence level corresponding to each pre-brush annotation result, where the spacing is used to indicate the distance between the obstacle and the vehicle;
[0153] The training data cleaning and packaging module 303 is specifically configured to determine the pre-brush labeling results whose spacing is less than a preset distance threshold and whose confidence is greater than a preset confidence threshold as target detection results.
[0154] In some embodiments, the 3D point cloud detection model training apparatus further includes: a fine-tuning training module 305;
[0155] A fine-tuning training module 305 is specifically configured to freeze first parameters in the initial 3D point cloud detection model and adjust second parameters using an error back propagation algorithm, wherein the first parameters include parameters of an early layer of the initial 3D point cloud detection model, and the second parameters include parameters of a later layer of the initial 3D point cloud detection model, and the learning rate of the first parameters is smaller than the learning rate of the second parameters;
[0156] The fine-tuning training module 305 is specifically used to train the initial three-dimensional point cloud detection model after parameter adjustment by refining the labeled data to obtain a target point cloud detection model.
[0157] In some embodiments, the 3D point cloud detection model training apparatus further includes: a model evaluation module 306;
[0158] Model evaluation module 306, used to obtain evaluation data sets;
[0159] The model evaluation module 306 is further used to input the evaluation data set into the target point cloud detection model to obtain the evaluation detection results;
[0160] The model evaluation module 306 is further used to match and calculate the evaluation detection results with the evaluation true values to obtain the evaluation results of the target point cloud detection model. The evaluation results include: the accuracy and recall rate corresponding to each category.
[0161] In the embodiment of the present application, each module can implement the three-dimensional point cloud detection model training method provided by the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0162] like Figure 4 As shown, an embodiment of the present application further provides an electronic device, which may include:
[0163] A memory 401 storing executable program code;
[0164] a processor 402 connected to the memory 401;
[0165] The processor 402 calls the executable program code stored in the memory 401 to execute the three-dimensional point cloud detection model training method executed by the electronic device in the above-mentioned method embodiments.
[0166] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the three-dimensional point cloud detection model training method in the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0167] An embodiment of the present application also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the three-dimensional point cloud detection model training method in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0168] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0169] In the several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0170] In the present disclosure, a processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0171] In this disclosure, memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0172] In this disclosure, those skilled in the art will appreciate that all or part of the steps in the various methods of the above-described embodiments can be performed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, including permanent and non-permanent, removable and non-removable storage media. The storage medium can use any method or technology to store information, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, Parallel Random Access Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Programmable Read Only Memory (PROM), Erasable Programmable Read Only Memory (EPROM), other types of Random Access Memory (RAM), Read Only Memory (ROM), One Time Programmable Read Only Memory (OTPROM), Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory or other memory technology, Compact Disc Read Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0173] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0174] It should be understood that “one embodiment” or “an embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, “in one embodiment” or “in an embodiment” appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present invention. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments. Dividing them into multiple embodiments is only used to highlight the different technical features in different embodiments. Those skilled in the art should be aware that the above-mentioned multiple embodiments can also be combined in any combination.
[0175] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the above-mentioned processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0176] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of this embodiment.
[0177] In addition, the functional units in the embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0178] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute some or all of the steps of the above-mentioned methods of various embodiments of the present invention.
[0179] The above are merely specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to these embodiments, but is to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A three-dimensional point cloud detection model training method, characterized in that: The method comprises: Acquire at least one original data packet, each original data packet including: point cloud data and trajectory pose data corresponding to a plurality of consecutive frames; Pre-labeling the at least one original data packet to obtain a pre-labeling result corresponding to each original data packet, wherein the pre-labeling result is used to indicate obstacle information included in each frame of point cloud data; Filtering the pre-brush marking results according to a preset classification method and specified obstacle information to obtain a target detection result corresponding to the original data packet; The original data packet and the target detection result are used to train an initial three-dimensional point cloud detection model, and the specified parameters in the initial three-dimensional point cloud detection model are adjusted through an error back propagation algorithm to obtain a target point cloud detection model.
2. The method according to claim 1, characterized in that The pre-marking of the at least one original data packet to obtain a pre-brush marking result corresponding to the original data packet includes: For each of the original data packets, determining a plurality of detection frames from each frame of point cloud data, each detection frame being used to indicate an obstacle; Determine a 3D detection result corresponding to each detection frame, wherein the 3D detection result includes the size, position, angle, type, and confidence level of the detection frame; The pre-brush annotation result corresponding to the original data packet is determined through a multi-target tracking algorithm according to the three-dimensional detection result corresponding to each frame of point cloud data and the trajectory posture data.
3. The method according to claim 2, characterized in that Determining the pre-brush annotation result corresponding to the original data packet according to the three-dimensional detection result corresponding to each frame of point cloud data and the trajectory pose data through the multi-target tracking algorithm includes: Calculate the similarity between each detection box in the current frame and the trajectory in the previous frame; Matching each detection frame with the trajectory according to the similarity to obtain a matching result corresponding to each detection frame; According to the matching result, multiple detection frames are processed to obtain the pre-refresh marking result corresponding to the original data packet.
4. The method according to claim 1, wherein The filtering and screening of the pre-brush marking results according to the preset classification method and the specified obstacle information to obtain the target detection result corresponding to the original data packet includes: Classify the pre-brush marking results according to obstacle categories to obtain pre-brush marking results corresponding to each specified category; The pre-brush marking results are filtered according to the spacing and confidence included in the pre-brush marking results to obtain the target detection result corresponding to the original data packet.
5. The method according to claim 4, characterized in that The filtering of the pre-brush annotation results according to the spacing and confidence included in the pre-brush annotation results to obtain the target detection result corresponding to the original data packet includes: Determine the spacing and confidence level corresponding to each pre-brush annotation result, where the spacing is used to indicate the distance between the obstacle and the vehicle; The pre-brush marking result whose spacing is smaller than a preset distance threshold and whose confidence is greater than a preset confidence threshold is determined as the target detection result.
6. The method according to claim 1, wherein The step of adjusting the specified parameters of the initial three-dimensional point cloud detection model by using an error back propagation algorithm to obtain a target point cloud detection model includes: Freezing first parameters in the initial three-dimensional point cloud detection model and adjusting second parameters using the error back propagation algorithm, wherein the first parameters include early layer parameters of the initial three-dimensional point cloud detection model, the second parameters include late layer parameters of the initial three-dimensional point cloud detection model, and a learning rate of the first parameters is smaller than a learning rate of the second parameters; The initial three-dimensional point cloud detection model after parameter adjustment is trained to obtain the target point cloud detection model.
7. The method according to claim 1, characterized in that The method further comprises: Get the evaluation dataset; Inputting the evaluation data set into the target point cloud detection model to obtain an evaluation detection result; The evaluation detection result is matched and calculated with the evaluation true value to obtain the evaluation result of the target point cloud detection model, and the evaluation result includes: the accuracy and recall rate corresponding to each category.
8. A 3D point cloud detection model training device, characterized in that: include: An acquisition module is used to acquire at least one original data packet, each of which includes: point cloud data and trajectory pose data corresponding to multiple consecutive frames; a 3D point cloud object detection pre-annotation module, configured to pre-annotate the at least one original data packet to obtain a pre-annotation result corresponding to each original data packet, wherein the pre-annotation result is used to indicate obstacle information included in each frame of point cloud data; A training data cleaning and packaging module is used to filter and screen the pre-brush annotation results according to a preset classification method and specified obstacle information to obtain a target detection result corresponding to the original data packet; An automated training module is used to train an initial three-dimensional point cloud detection model using the original data packet and the target detection result, and to adjust specified parameters in the initial three-dimensional point cloud detection model through an error back propagation algorithm to obtain a target point cloud detection model.
9. An electronic device, characterized in that: include: a memory storing executable program code; and a processor connected to said memory; The processor calls the executable program code stored in the memory to execute the three-dimensional point cloud detection model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that include: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the processor, the three-dimensional point cloud detection model training method according to any one of claims 1 to 7 is implemented.