Training method for point cloud 3D detection model and point cloud detection method

By associating training data input with iteration cycles and reducing irrelevant data, the method addresses domain gaps in point cloud 3D detection models, improving their accuracy and robustness for autonomous driving.

CN116386026BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310340658.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-07-15
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

In the prior art, the point cloud 3D detection model has domain gap problems during training, resulting in low detection accuracy and difficulty in adapting to different test scenarios.

Method used

By correlating the iteration rounds of the point cloud 3D detection model with the number of input training data, the amount of manually labeled 3D detection box data is gradually reduced, the amount of radar collected data related to the current test scenario is increased, and the training process is optimized.

Benefits of technology

It effectively avoids domain gap problems, improves the accuracy and robustness of point cloud 3D detection models, and improves the safety of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386026B_ABST
    Figure CN116386026B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method for a point cloud 3D detection model and a point cloud detection method, which relate to the field of artificial intelligence technology, and particularly to the fields of deep learning, point cloud detection, and autonomous driving technology. The specific implementation solution is as follows: obtaining first training data and second training data, where the first training data is point cloud data obtained from a labeled 3D detection box, and the second training data is point cloud data collected by a radar; obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target quantity is associated with the iteration round of the point cloud 3D detection model; for any iteration round in the training process of the point cloud 3D detection model, inputting the second training data and the first training data with the target quantity corresponding to this iteration round into the point cloud 3D detection model to train the point cloud 3D detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, point cloud detection, and autonomous driving technology, and specifically relates to a method for training a point cloud 3D detection model, a point cloud detection method, a device for training a point cloud 3D detection model, a point cloud detection device, an electronic device, a storage medium, a computer program product, and an autonomous driving vehicle. Background Art

[0002] LiDAR plays an important role in an autonomous driving system. With LiDAR, the autonomous driving system can accurately perform real-time three-dimensional (3D) modeling of the vehicle's environment, which can improve the safety of the autonomous driving system and accurately perceive the position, size, and attitude of a certain 3D target in the LiDAR point cloud coordinate system. Currently, the detection task of 3D point cloud targets is usually achieved through a neural network model. Summary of the Invention

[0003] The present disclosure provides a method for training a point cloud 3D detection model, a point cloud detection method, a device for training a point cloud 3D detection model, a point cloud detection device, an electronic device, a storage medium, a computer program product, and an autonomous driving vehicle.

[0004] According to a first aspect of the present disclosure, there is provided a method for training a point cloud 3D detection model, including:

[0005] Obtaining first training data and second training data, where the first training data is point cloud data obtained from a labeled 3D detection box, and the second training data is point cloud data collected by a radar;

[0006] Obtaining the target number of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target number is associated with the iteration round of the point cloud 3D detection model;

[0007] For any iteration round in the training process of the point cloud 3D detection model, inputting the second training data and the first training data with the target number corresponding to this iteration round into the point cloud 3D detection model to train the point cloud 3D detection model;

[0008] Wherein, the input of the trained point cloud 3D detection model is the point cloud data collected by the radar, and the output is a 3D detection box.

[0009] According to a second aspect of the present disclosure, there is provided a point cloud detection method, including:

[0010] Obtaining the point cloud data collected by the radar;

[0011] Input the point cloud data into a point cloud 3D detection model, and obtain the 3D detection bounding boxes output by the point cloud 3D detection model;

[0012] Wherein, the point cloud 3D detection model is a model obtained after being trained based on the training method of the point cloud 3D detection model as described in the first aspect.

[0013] According to the third aspect of the present disclosure, there is provided a training device for a point cloud 3D detection model, including:

[0014] A first acquisition module, configured to acquire first training data and second training data, wherein the first training data is point cloud data obtained from labeled 3D detection bounding boxes, and the second training data is point cloud data collected by a radar;

[0015] A second acquisition module, configured to acquire the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, and the target quantity is associated with the iteration round of the point cloud 3D detection model;

[0016] A training module, configured to input the second training data and the first training data of the target quantity corresponding to the iteration round into the point cloud 3D detection model for any iteration round during the training process of the point cloud 3D detection model, so as to train the point cloud 3D detection model;

[0017] Wherein, the input of the trained point cloud 3D detection model is the point cloud data collected by the radar, and the output is the 3D detection bounding box.

[0018] According to the fourth aspect of the present disclosure, there is provided a point cloud detection device, including:

[0019] A third acquisition module, configured to acquire the point cloud data collected by the radar;

[0020] A fourth acquisition module, configured to input the point cloud data into the point cloud 3D detection model, and obtain the 3D detection bounding boxes output by the point cloud 3D detection model;

[0021] Wherein, the point cloud 3D detection model is a model obtained after being trained based on the training device of the point cloud 3D detection model as described in the third aspect.

[0022] According to the fifth aspect of the present disclosure, there is provided an electronic device, including:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect or the second aspect.

[0026] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect or the second aspect.

[0027] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method described in the first aspect or the second aspect.

[0028] According to an eighth aspect of the present disclosure, there is provided an autonomous driving vehicle including a point cloud detection device as described in the fourth aspect.

[0029] In the embodiments of the present disclosure, the number of the first training data input to the point cloud 3D detection model is associated with the number of iterations of the model, and also makes the training data input to the point cloud 3D detection model variable, so as to be able to control the input data of the point cloud 3D detection model, making the training of the point cloud 3D detection model more flexible, and also being able to improve the accuracy of the trained point cloud 3D detection model by controlling the number of input data.

[0030] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0032] Figure 1 is a flowchart of a method for training a point cloud 3D detection model provided by an embodiment of the present disclosure;

[0033] Figure 2 is a relationship diagram between the input quantity of the first training data and the number of iterations in a method for training a point cloud 3D detection model provided by an embodiment of the present disclosure;

[0034] Figure 3 is a flowchart of a point cloud detection method provided by an embodiment of the present disclosure;

[0035] Figure 4 is a structural diagram of a device for training a point cloud 3D detection model provided by an embodiment of the present disclosure;

[0036] Figure 5It is the second structural diagram of a training device for a point cloud 3D detection model provided by an embodiment of the present disclosure;

[0037] Figure 6 It is the structural diagram of a point cloud detection device provided by an embodiment of the present disclosure;

[0038] Figure 7 It is a block diagram of an electronic device for implementing the training method or the point cloud detection method of the point cloud 3D detection model according to the embodiment of the present disclosure. Detailed implementation manners

[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0040] For better understanding, the following explains relevant concepts and principles that may be involved in the embodiments of the present disclosure.

[0041] The point cloud 3D detection task refers to performing real-time 3D modeling of the current environment through a lidar, and simultaneously perceiving the position, size, and pose of a certain 3D target in the lidar point cloud coordinate system. Usually, the data collected by the lidar is displayed and processed in the form of a point cloud. Simply put, a point cloud is N points in space, and each point contains three floating-point values XYZ representing the spatial position and an R value representing the echo intensity. A point cloud is a type of data containing the geometric shape information of the object surface. In the related art, the point cloud 3D detection is usually realized through a point cloud 3D detection model.

[0042] Please refer to Figure 1 , Figure 1 It is a flowchart of a training method for a point cloud 3D detection model provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps:

[0043] Step S101, obtain first training data and second training data, where the first training data is point cloud data obtained from a labeled 3D detection box, and the second training data is point cloud data collected by the radar.

[0044] It should be noted that the method provided by the embodiments of the present disclosure can be applied to electronic devices such as computers, tablets, mobile phones, etc. In the following embodiments, the electronic device will be used as the execution subject to explain the specific implementation process of the method provided by the embodiments of the present disclosure.

[0045] In an embodiment of the present disclosure, the first training data may be point cloud data corresponding to a point cloud object in a 3D detection box (which may also be referred to as a 3D bounding box, a 3D envelope box, etc.) manually labeled. For example, after obtaining the point cloud data collected by a radar for a certain vehicle, a user manually labels a 3D detection box for the vehicle based on the point cloud data, and then the first training data is the point cloud data corresponding to the vehicle. It should be noted that the user may obtain point cloud data through a large number of manually labeled 3D detection boxes, and a point cloud database is obtained based on these point cloud data. The electronic device may obtain the first training data from this point cloud database. For example, a part of the point cloud data may be randomly selected as the first training data.

[0046] Optionally, the second training data may be point cloud data collected by a lidar on a vehicle, and the vehicle may be an autonomous vehicle.

[0047] Step S102: Obtain the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target quantity is associated with the iteration round of the point cloud 3D detection model.

[0048] It can be understood that a neural network model usually needs to be obtained through multiple rounds of iterative training. The point cloud 3D detection model in the embodiment of the present disclosure also needs to be obtained through multiple rounds of iterative training. One iteration round represents one iterative training of the point cloud 3D detection model. In each iteration round, the quantity of the training data input into the point cloud 3D detection model may be different.

[0049] In the embodiment of the present disclosure, the iteration round of the point cloud 3D detection model is associated with the quantity of the first training data to be input. Optionally, the association relationship between the iteration round and the quantity of the first training data to be input may be preset. For example, the association relationship may be: as the iteration round increases, the quantity of the first training data to be input decreases accordingly. Or, it may also be set that in the first N iteration rounds, the quantity of the first training data to be input in each iteration round is x, and after the Nth iteration round, the quantity of the first training data to be input in each iteration round is y, where x is different from y. Of course, the association relationship between the iteration round and the quantity of the first training data to be input may also be other possible situations, which will not be listed in detail here.

[0050] Step S103: For any iteration round in the training process of the point cloud 3D detection model, input the second training data and the first training data with the target quantity corresponding to this iteration round into the point cloud 3D detection model to train the point cloud 3D detection model.

[0051] For example, taking the first target iteration round as an example, the first target iteration round is any iteration round during the training process of the point cloud 3D detection model. Suppose the target quantity of the first training data to be input into the point cloud 3D detection model corresponding to the first target iteration round is x. Then, when the point cloud 3D detection model is being trained in this first target iteration round, the second training data and the first training data with a target quantity of x are input into the point cloud 3D detection model to train the point cloud 3D detection model in this first target iteration round. It can be understood that for any iteration round during the training process of the point cloud 3D detection model, the model training can be carried out in the above manner, that is, inputting the second training data and the first training data with the target quantity corresponding to this iteration round into the point cloud 3D detection model to achieve the training of the point cloud 3D detection model in this iteration round. Thus, based on such a method, the training of the point cloud 3D detection model in all iteration rounds can be completed.

[0052] In the embodiments of the present disclosure, after obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, for a certain iteration round among them, the first training data with the target quantity corresponding to this iteration round and the second training data are input into the point cloud 3D detection model to train the point cloud 3D detection model for this iteration round, and based on such a method, the training of the point cloud 3D detection model in all iteration rounds is completed, so as to obtain the trained point cloud 3D detection model.

[0053] It should be noted that during the training process of each iteration round of the point cloud 3D detection model, the quantity of the input second training data can be the same, while the quantity of the input first training data is related to the current iteration round, that is, in different iteration rounds, the quantity of the first training data input into the point cloud 3D detection model can be different, that is, the input quantity of the point cloud data obtained through manually labeled 3D detection frames can be different. In this way, it also makes the training data input into the point cloud 3D detection model variable, so as to realize the control of the input data of the point cloud 3D detection model, make the training of the point cloud 3D detection model more flexible, and also be able to improve the accuracy of the trained point cloud 3D detection model by controlling the quantity of the input data.

[0054] In the embodiments of the present disclosure, the trained point cloud 3D detection model can be applied to an autonomous driving vehicle. The input of the trained point cloud 3D detection model is the point cloud data collected by the lidar on the autonomous driving vehicle, and the output is a 3D detection frame, so as to better assist the autonomous driving of the autonomous driving vehicle and improve the safety of the autonomous driving vehicle.

[0055] Optionally, in step S102, obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round may specifically further include:

[0056] When the iteration round is the first N iteration rounds, determining that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a first preset quantity;

[0057] When the iteration round is an iteration round after the Nth iteration round, determining that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a second preset quantity;

[0058] Wherein, the first preset quantity is greater than the second preset quantity, the value of N is less than the value of the last iteration round of the point cloud 3D detection model, and the value of N can be a preset value.

[0059] In the embodiments of the present disclosure, in the first N iteration rounds of the point cloud 3D detection model, the target quantity of the first training data to be input in each iteration round is set as the first preset quantity, and in the iteration rounds after the Nth iteration round, the target quantity of the first training data to be input in each iteration round is set as the second preset quantity, and the second preset quantity is less than the first preset quantity. That is, after the point cloud 3D detection model completes a certain number of iterative trainings, for the subsequent iterative round trainings, the quantity of the first training data input into the point cloud 3D detection model will decrease. And the first training data is point cloud data obtained based on manually labeled 3D detection frames, and these point cloud data may not be relevant to the current test scenario, while the second training data is point cloud data collected by a lidar, that is, the second training data is collected on-site based on the current test scenario, and the second training data is relevant to the current test scenario.

[0060] It should be noted that for a point cloud 3D detection model, if the training data and test data of the model target different scenarios respectively, there may be a problem of domain gap. The domain gap refers to an obvious distribution difference between two parts of data. If one part of the data is used for training and the other for testing, the performance of the model will be worse than that obtained from training and testing with data from a unified distribution. For example, the most intuitive case of domain gap is: Suppose a batch of training data is collected on a certain ring road in City A, and a batch of test data is collected on a certain ring road in City B. The model trained with the training data from City A will not perform well when running on the test data from City B because there are obvious distribution differences between the two data sets, such as: the dimensions of vehicle length, width and height, road traffic participants, the distribution of point cloud background objects, etc. These distribution differences are not seen by the model during training, so the model will perform poorly during testing and the accuracy will be relatively low.

[0061] In the embodiments of the present disclosure, by reducing the quantity of the first training data, it is possible to reduce the influence of the first training data that is not relevant to the current test scenario on the point cloud 3D detection model during the training process of the point cloud 3D detection model, so that the point cloud 3D detection model focuses more on completing the training through the second training data related to the current test scenario in the later stage of training, thereby enabling the trained point cloud 3D detection model to be more applicable to the current test scenario, effectively avoiding the generation of the domain gap problem, and thus improving the accuracy of the trained point cloud 3D detection model.

[0062] Optionally, in the case of an iteration round after the Nth iteration round, determining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round as a second preset quantity includes:

[0063] In the case of an iteration round after the Nth iteration round, determine the (N + 1)th iteration round to the (N + n)th iteration round, where n takes values of 1, 2, 3... n, and the value of N + n is less than or equal to the value of the last iteration round of the point cloud 3D detection model;

[0064] In the (N + 1)th iteration round to the (N + n)th iteration round, determine the second preset quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the second preset quantity gradually decreases as the value of the iteration round increases.

[0065] Exemplarily, it is assumed that the iterative training of the point cloud 3D detection model includes a total of 100 iterative rounds, and the value of N is 60. Then, in the first 60 iterative rounds, the number of the first training data to be input into the point cloud 3D detection model in each iterative round is the first preset quantity. It is assumed that the value of n is 20, that is, in the 61st to 80th iterative rounds, the number of the first training data to be input into the point cloud 3D detection model in each iterative round gradually decreases. For example, the number of the first training data decreases linearly. In this way, in the later stage of the training of the point cloud 3D detection model, the influence of the first training data that is not relevant to the current test scenario can be gradually reduced.

[0066] Optionally, the values of N and n can be preset values. Furthermore, when training the point cloud 3D detection model, by the preset values of N and n, it can be determined in which iterative rounds the number of the first training data to be input into the point cloud 3D detection model gradually decreases. Thus, during the training process of these iterative rounds, the random sampling amount of the first training data can be gradually reduced, so as to gradually reduce the number of the first training data input into the point cloud 3D detection model. In this way, in the later stage of the training of the point cloud 3D detection model, by gradually reducing the number of the input first training data, the influence of the first training data that is not relevant to the current test scenario on the point cloud 3D detection model can be gradually reduced, so that the point cloud 3D detection model focuses more on completing the training through the second training data related to the current test scenario in the later stage of training, effectively avoiding the generation of the domain gap problem and improving the accuracy of the point cloud 3D detection model.

[0067] It should be noted that after the (N + n)th iterative round, the number of the first training data input into the point cloud 3D detection model can remain unchanged. For example, it remains the same as the number of the first training data input in the (N + n)th iterative round, or the number of the input first training data can also continue to decrease, so as to effectively reduce the influence of the first training data that is not relevant to the current test scenario on the point cloud 3D detection model, so that the trained point cloud 3D detection model has higher accuracy.

[0068] Optionally, the second preset quantity corresponding to the N + n-th iteration round is 0. That is, in the N + n-th iteration round, the quantity of the first training data input into the point cloud 3D detection model is 0. That is to say, in the iteration rounds from the N + 1-th iteration round to the N + n-th iteration round, as the iteration round increases, the quantity of the first training data input into the point cloud 3D detection model gradually decreases and finally decreases to 0. In this way, in the later stage of the training of the point cloud 3D detection model, its training data only includes the second training data related to the current test scenario, and no longer includes the first training data unrelated to the current test scenario, which makes the training of the point cloud 3D detection model more suitable for the current test scenario, effectively avoiding the generation of domain gap problems, and thus ensuring that the trained point cloud 3D detection model has higher accuracy when applied to the current test scenario and improving the robustness of the output data of the trained point cloud 3D detection model.

[0069] It should be noted that after the N + n-th iteration round, the quantity of the first training data input into the point cloud 3D detection model may still remain 0, which makes the point cloud 3D detection model only use the second training data related to the current test scenario for model training in the later stage of training, and enables the trained point cloud 3D detection model to obtain better and more robust test results on the test data.

[0070] Optionally, in the iteration rounds from the N + 1-th iteration round to the N + n-th iteration round, determining the second preset quantity of the first training data that needs to be input into the point cloud 3D detection model in each iteration round includes:

[0071] Obtaining the target quantity of the first training data that needs to be input into the point cloud 3D detection model in the N-th iteration round;

[0072] According to the target quantity, the value of N + 1, and the value of N + n, determining the quantity of the first training data that needs to be input into the point cloud 3D detection model in the target iteration round;

[0073] Wherein, the target iteration round is any iteration round from the N + 1-th iteration round to the N + n-th iteration round.

[0074] Exemplarily, assume that the target quantity of the first training data that needs to be input into the point cloud 3D detection model in the N-th iteration round is N, the N + 1-th iteration round is e0, that is, the value of N + 1 is represented by e0, and the N + n-th iteration round is e1, that is, the value of N + n is represented by e1. Then, the quantity of the first training data that needs to be input into the point cloud 3D detection model in any iteration round from the N + 1-th iteration round to the N + n-th iteration round can be determined according to N, e0, and e1.

[0075] In the embodiments of the present disclosure, it is also possible to determine the quantity of the first training data to be input to the point cloud 3D detection model in any iteration round from the (N + 1)-th iteration round to the (N + n)-th iteration round according to the value of the iteration round and the target quantity of the first training data to be input to the point cloud 3D detection model in the N-th iteration round. Thus, it is possible to flexibly adjust the first training data input to the point cloud 3D detection model, so as to better control the training process of the point cloud 3D detection model and ensure that the trained point cloud 3D detection model has better accuracy.

[0076] Optionally, in the (N + 1)-th iteration round to the (N + n)-th iteration round, the second preset quantity decreases linearly or curvilinearly with respect to the iteration round. For example, a preset algebraic relationship may be satisfied among N, e0, e1, and the quantity of the first training data.

[0077] Exemplarily, in an alternative embodiment, the algebraic relationship satisfied among N, e0, e1, and the quantity of the first training data to be input is as follows:

[0078]

[0079] where f(e) represents the quantity of the first training data to be input to the point cloud 3D detection model in the e-th iteration round (i.e., the second target iteration round), e0 represents the (N + 1)-th iteration round, e1 represents the (N + n)-th iteration round, and N represents the target quantity of the first training data to be input to the point cloud 3D detection model in the N-th iteration round. In this case, as Figure 2 shown, in the (N + 1)-th iteration round to the (N + n)-th iteration round, the quantity of the first training data to be input to the point cloud 3D detection model in each iteration round decreases linearly with respect to the iteration round.

[0080] Alternatively, in another alternative embodiment, the algebraic relationship satisfied among N, e0, e1, and the quantity of the first training data to be input is as follows:

[0081]

[0082] where f(e) represents the quantity of the first training data to be input to the point cloud 3D detection model in the e-th iteration round (i.e., the second target iteration round), e0 represents the (N + 1)-th iteration round, e1 represents the (N + n)-th iteration round, and N represents the target quantity of the first training data to be input to the point cloud 3D detection model in the N-th iteration round. In this case, in the (N + 1)-th iteration round to the (N + n)-th iteration round, the quantity of the first training data to be input to the point cloud 3D detection model in each iteration round decreases curvilinearly with respect to the iteration round.

[0083] In the embodiments of the present disclosure, through different algebraic relations, it is possible to flexibly adjust the quantity of the first training data to be input into the point cloud 3D detection model in each iteration round from the (N + 1)-th iteration round to the (N + n)-th iteration round, and it can be ensured that the quantity of the first training data input in each iteration round gradually decreases. Thus, in the later stage of the training of the point cloud 3D detection model, its training data gradually shifts towards the second training data related to the current test scenario, which makes the training of the point cloud 3D detection model more suitable for the current test scenario, effectively avoiding the generation of domain gap problems, and ensuring that the trained point cloud 3D detection model has higher accuracy when applied to the current test scenario.

[0084] Please refer to Figure 3 , Figure 3 which is a flowchart of a point cloud detection method provided by the embodiments of the present disclosure. As Figure 3 shown, the method includes the following steps:

[0085] Step S301: Obtain the point cloud data collected by the radar.

[0086] Optionally, the radar may be a lidar installed on an autonomous vehicle. In the embodiments of the present disclosure, the point cloud detection method may be applied to an autonomous vehicle.

[0087] Step S302: Input the point cloud data into the point cloud 3D detection model, and obtain the 3D detection box output by the point cloud 3D detection model.

[0088] Wherein, the point cloud 3D detection model is a model obtained after being trained by the training method of the point cloud 3D detection model as described in the above embodiments.

[0089] In the embodiments of the present disclosure, the point cloud 3D detection model is trained by the training method as described above, and the point cloud 3D detection model has better and more robust detection effects, which makes the 3D detection box output by the point cloud 3D detection model have higher accuracy, and thus can better assist the driving of the autonomous vehicle and effectively improve the safety of the autonomous vehicle.

[0090] Please refer to Figure 4 , Figure 4 which is one of the structural diagrams of a training device for a point cloud 3D detection model provided by the embodiments of the present disclosure. As Figure 4 shown, the training device 400 for the point cloud 3D detection model includes:

[0091] The first acquisition module 401 is configured to acquire first training data and second training data, where the first training data is point cloud data obtained from labeled 3D detection boxes, and the second training data is point cloud data collected by a radar;

[0092] The second acquisition module 402 is configured to acquire the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target quantity is associated with the iteration round of the point cloud 3D detection model;

[0093] The training module 403 is configured to, for any iteration round during the training process of the point cloud 3D detection model, input the second training data and the first training data with the target quantity corresponding to this iteration round into the point cloud 3D detection model to train the point cloud 3D detection model;

[0094] Wherein, the input of the trained point cloud 3D detection model is the point cloud data collected by the radar, and the output is a 3D detection box.

[0095] Optionally, please further refer to Figure 5 The second acquisition module 402 includes:

[0096] The first determination unit 4021 is configured to, when the iteration round is the first N iteration rounds, determine that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a first preset quantity;

[0097] The second determination unit 4022 is configured to, when the iteration round is an iteration round after the Nth iteration round, determine that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a second preset quantity;

[0098] Wherein, the first preset quantity is greater than the second preset quantity, and the value of N is less than the value of the last iteration round of the point cloud 3D detection model.

[0099] Optionally, the second determination unit 4022 is further configured to:

[0100] When the iteration round is an iteration round after the Nth iteration round, determine the (N + 1)th iteration round to the (N + n)th iteration round, where the value of n is 1, 2, 3... n, and the value of N + n is less than or equal to the value of the last iteration round of the point cloud 3D detection model;

[0101] In the (N + 1)-th to (N + n)-th iteration rounds, determine the second preset quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the second preset quantity gradually decreases as the value of the iteration round increases.

[0102] Optionally, the second determination unit 4022 is further configured to:

[0103] Obtain the target quantity of the first training data to be input into the point cloud 3D detection model in the N-th iteration round;

[0104] Determine the quantity of the first training data to be input into the point cloud 3D detection model in the target iteration round according to the target quantity, the value of N + 1, and the value of N + n;

[0105] Wherein, the target iteration round is any iteration round from the (N + 1)-th to the (N + n)-th iteration rounds.

[0106] Optionally, in the (N + 1)-th to (N + n)-th iteration rounds, the second preset quantity decreases linearly or curvilinearly with respect to the iteration round.

[0107] Optionally, the second preset quantity corresponding to the (N + n)-th iteration round is 0.

[0108] It should be noted that the device provided in the embodiments of the present disclosure can implement all the technical processes of the above Figure 1 training method of the point cloud 3D detection model and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0109] Please refer to Figure 6 , Figure 6 which is a structural diagram of a point cloud detection device provided in the embodiments of the present disclosure. As Figure 6 shown, the point cloud detection device 600 includes:

[0110] A third acquisition module 601, configured to acquire point cloud data collected by a radar;

[0111] A fourth acquisition module 602, configured to input the point cloud data into a point cloud 3D detection model and acquire a 3D detection box output by the point cloud 3D detection model;

[0112] Wherein, the point cloud 3D detection model is a model obtained by training based on the training device of the point cloud 3D detection model as described above.

[0113] It should be noted that the device provided in the embodiments of the present disclosure can implement the above Figure 3The entire technical process of the point cloud detection method and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0114] An embodiment of the present disclosure also provides an autonomous vehicle, including the point cloud detection device described above. By adopting the point cloud detection device described above, the autonomous vehicle provided by the embodiment of the present disclosure can obtain a 3D detection frame with higher accuracy, thereby better assisting the driving of the autonomous vehicle and effectively improving the safety of the autonomous vehicle.

[0115] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0116] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Figure 7 A schematic block diagram of an exemplary electronic device 700 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0118] As Figure 7 shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0119] Multiple components in device 700 are connected to I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0120] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the training method or the detection method of the point cloud 3D detection model. For example, in some embodiments, the training method or the detection method of the point cloud 3D detection model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method or the detection method of the point cloud 3D detection model described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the above-described training method or the detection method of the point cloud 3D detection model by any other suitable means (e.g., by means of firmware).

[0121] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0123] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0126] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0127] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A training method for a point cloud 3D detection model, comprising: Obtaining first training data and second training data, wherein the first training data is point cloud data obtained from labeled 3D detection boxes, and the second training data is point cloud data collected by a radar; Obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target quantity is associated with the iteration round of the point cloud 3D detection model; For any iteration round during the training process of the point cloud 3D detection model, inputting the second training data and the first training data of the target quantity corresponding to this iteration round into the point cloud 3D detection model to train the point cloud 3D detection model; Wherein, the input of the trained point cloud 3D detection model is the point cloud data collected by the radar, and the output is a 3D detection box; The obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round includes: When the iteration round is the first N iteration rounds, determining that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a first preset quantity; When the iteration round is an iteration round after the Nth iteration round, determining that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a second preset quantity; Wherein, the first preset quantity is greater than the second preset quantity, and the value of N is less than the value of the last iteration round of the point cloud 3D detection model.

2. The method according to claim 1, wherein The determining that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a second preset quantity when the iteration round is an iteration round after the Nth iteration round includes: When the iteration round is an iteration round after the Nth iteration round, determining the (N + 1)th iteration round to the (N + n)th iteration round, where the value of n is 1, 2, 3... n, and the value of N + n is less than or equal to the value of the last iteration round of the point cloud 3D detection model; In the (N + 1)th iteration round to the (N + n)th iteration round, determining the second preset quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, wherein the second preset quantity gradually decreases as the value of the iteration round increases.

3. The method according to claim 2, wherein, The determining the second preset quantity of the first training data to be input into the point cloud 3D detection model in each iteration round in the (N + 1)th iteration round to the (N + n)th iteration round includes: Obtaining the target quantity of the first training data to be input into the point cloud 3D detection model in the Nth iteration round; According to the target quantity, the value of N + 1, and the value of N + n, determining the quantity of the first training data to be input into the point cloud 3D detection model in the target iteration round; Wherein, the target iteration round is any iteration round in the (N + 1)th iteration round to the (N + n)th iteration round.

4. The method according to claim 3, wherein In the (N + 1)-th to (N + n)-th iteration rounds, the second preset quantity decreases linearly or curvilinearly with respect to the iteration rounds.

5. The method according to claim 2, wherein, The second preset quantity corresponding to the (N + n)-th iteration round is 0.

6. A point cloud detection method, comprising: Obtaining point cloud data collected by a radar; Inputting the point cloud data into a point cloud 3D detection model, and obtaining a 3D detection box output by the point cloud 3D detection model; Wherein, the point cloud 3D detection model is a model obtained by training based on the training method of the point cloud 3D detection model according to any one of claims 1 - 5.

7. A training device for a point cloud 3D detection model, comprising: A first acquisition module, configured to acquire first training data and second training data, wherein the first training data is point cloud data obtained from a labeled 3D detection box, and the second training data is point cloud data collected by a radar; A second acquisition module, configured to acquire the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the target quantity is associated with the iteration round of the point cloud 3D detection model; A training module, configured to input the second training data and the first training data with the target quantity corresponding to the iteration round into the point cloud 3D detection model for any iteration round during the training process of the point cloud 3D detection model, so as to train the point cloud 3D detection model; Wherein, the input of the trained point cloud 3D detection model is point cloud data collected by a radar, and the output is a 3D detection box; The second acquisition module includes: A first determination unit, configured to determine that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a first preset quantity when the iteration round is the first N iteration rounds; A second determination unit, configured to determine that the target quantity of the first training data to be input into the point cloud 3D detection model in each iteration round is a second preset quantity when the iteration round is an iteration round after the N-th iteration round; Wherein, the first preset quantity is greater than the second preset quantity, and the value of N is less than the value of the last iteration round of the point cloud 3D detection model.

8. The apparatus according to claim 7, wherein, The second determination unit is further configured to: When the iteration round is an iteration round after the N-th iteration round, determine the (N + 1)-th to (N + n)-th iteration rounds, where the value of n is 1, 2, 3... n, and the value of N + n is less than or equal to the value of the last iteration round of the point cloud 3D detection model; In the (N + 1)-th to (N + n)-th iteration rounds, determine the second preset quantity of the first training data to be input into the point cloud 3D detection model in each iteration round, where the second preset quantity gradually decreases as the value of the iteration round increases.

9. The apparatus according to claim 8, wherein, The second determination unit is further configured to: Acquire the target quantity of the first training data to be input into the point cloud 3D detection model in the N-th iteration round; Determine the quantity of the first training data to be input into the point cloud 3D detection model for the target iteration round according to the target quantity, the value of N+1, and the value of N+n; Wherein, the target iteration round is any iteration round from the (N+1)-th iteration round to the (N+n)-th iteration round.

10. The device according to claim 9, wherein, Among the (N+1)-th iteration round to the (N+n)-th iteration round, the second preset quantity decreases linearly or curvilinearly with respect to the iteration round.

11. The apparatus according to claim 8, wherein, The second preset quantity corresponding to the (N+n)-th iteration round is 0.

12. A point cloud detection device, comprising: A third acquisition module, configured to acquire point cloud data collected by a radar; A fourth acquisition module, configured to input the point cloud data into a point cloud 3D detection model and acquire a 3D detection frame output by the point cloud 3D detection model; Wherein, the point cloud 3D detection model is a model obtained after being trained by a training device of the point cloud 3D detection model according to any one of claims 7-11.

13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

15. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.

16. An autonomous vehicle, comprising the point cloud detection device according to claim 12.

Citation Information

Patent Citations

  • 3D target detection method, model training method, related device and electronic equipment

    CN113674421A

  • 3D target detection method and device

    CN114332845A