Training Method and 3D Detection Method, Device and Equipment of 3D Detection Model

By training the initial three-dimensional detection model, it can perform visual 3D detection with high accuracy during the unmanned truck movement, solving the problem of poor detection effect caused by the fluctuation of the sensor's external parameters, and achieving stable and safe high-precision detection.

CN116030434BActive Publication Date: 2025-05-27BEIJING TRUNK TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211726559.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-05-27
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

During the movement of unmanned trucks, the external parameters of the sensor fluctuate greatly, resulting in poor visual 3D detection effect and it is difficult to obtain accurate external parameters estimates.

Method used

A three-dimensional detection model training method is adopted to obtain detection images marked with camera data and laser point cloud data, and the initial three-dimensional detection model (YOLO5 model) is trained. The model includes two-dimensional detection branches and three-dimensional detection branches. It can output coordinate information, categories and quantities of the target object under the camera and vehicle body coordinate system without relying on internal and external parameters of the camera.

Benefits of technology

It realizes high-precision visual 3D detection during unmanned truck movement, and can identify targets of 200m front and back and 50m left and right, with a relative error accuracy of <3%, with high stability and safety, improving the visual 3D detection effect of unmanned trucks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030434B_ABST
    Figure CN116030434B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a three-dimensional detection model, a three-dimensional detection method, a device, and a device. The training method for the three-dimensional detection model includes: obtaining a detection image, where the detection image is marked with camera data and laser point cloud data of a detected object; inputting the detection image into an initial three-dimensional detection model to train the initial three-dimensional detection model, the initial three-dimensional detection model being a YOLO5 model, the network layer of the YOLO5 model including a two-dimensional detection branch and a three-dimensional detection branch, the two-dimensional detection branch being used to detect camera data in the image, and the three-dimensional detection branch being used to detect laser point cloud data in the image; when a training termination condition is reached, obtaining a three-dimensional detection model, the three-dimensional detection model being used to receive an image to be detected and output a recognition result of an object to be detected in the image to be detected. The method of the present application can solve the problem of how to improve the visual 3D detection effect of an unmanned truck.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to three-dimensional detection technology, and in particular, to a training method for a three-dimensional detection model, a three-dimensional detection method, device, and equipment. Background Art

[0002] In the field of intelligent driving, vision-based monocular three-dimensional (3D) detection remains a highly challenging task. Currently, mainstream monocular 3D detection methods are all based on the camera coordinate system. First, the 3D target ground truth in the vehicle body coordinate system is transformed to the camera coordinate system through the external parameters, and a deep learning model is trained in the camera coordinate system. In the model deployment stage, the model infers the 3D information of the target in the camera coordinate system, and then through the external parameter matrix of the camera, the target is converted to the vehicle body coordinate system for use by the downstream fusion algorithm. In this process, accurate internal and external parameter model parameters are required.

[0003] As an important branch in the field of intelligent driving, driverless trucks have their unique scenario characteristics. That is, for driverless trucks, the perception sensors are generally installed at the front of the truck. Due to the relatively soft suspension of the truck, the external parameters of each sensor fluctuate within a large range during the movement of the truck, so it is difficult to obtain accurate estimates of the external parameters of each sensor. This brings particularly difficult troubles to the vision 3D detection of driverless trucks.

[0004] How to improve the vision 3D detection effect of driverless trucks still needs to be solved. Summary of the Invention

[0005] The present application provides a training method for a three-dimensional detection model, a three-dimensional detection method, device, and equipment to solve the problem of how to improve the vision 3D detection effect of driverless trucks.

[0006] On the one hand, the present application provides a training method for a three-dimensional detection model, including:

[0007] Obtain a detection image, where the detection image is marked with camera data and laser point cloud data of a detected object. The laser point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width, and height of the detected object, as well as the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image;

[0008] Input the detection image into an initial three-dimensional detection model to train the initial three-dimensional detection model. The initial three-dimensional detection model is a YOLO5 model, and the network layer of the YOLO5 model includes a two-dimensional detection branch and a three-dimensional detection branch. The two-dimensional detection branch is used to detect the camera data in the image, and the three-dimensional detection branch is used to detect the laser point cloud data in the image;

[0009] When the training termination condition is reached, a three-dimensional detection model is obtained. The three-dimensional detection model is used to receive the image to be detected and output the recognition result of the target object in the image to be detected. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0010] In one embodiment, the three-dimensional detection branch includes three three-dimensional detection heads. Each three-dimensional detection head includes three prior boxes. Each three-dimensional detection head is used to predict the coordinate information of the detected object in the vehicle body coordinate system, the length, width and height of the detected object, and the first angle and second angle of the detected object. The first angle and second angle of the detected object are used to determine the heading angle of the detected object.

[0011] In one embodiment, the two-dimensional detection branch includes three two-dimensional detection heads. Each two-dimensional detection head includes three prior boxes. Each two-dimensional detection head is used to predict the coordinate information of the detected object in the camera coordinate system, the length and width of the detected object, the category information of the detected object, and the quantity of the detected objects of the target type.

[0012] In one embodiment, the output layer of the initial three-dimensional detection model is used to merge the output results of the two-dimensional detection branch and the three-dimensional detection branch.

[0013] It is also used to filter the background image information in the merged output result according to the output threshold to obtain the detection object recognition result. The detection object recognition result includes the coordinate information of each predicted detection object in the camera coordinate system and the coordinate information of each predicted detection object in the vehicle body coordinate system.

[0014] In one embodiment, the output layer is further used to filter the overlapping boxes in the merged output result based on the non-maximum suppression algorithm.

[0015] In one embodiment, the loss function of the two-dimensional detection branch includes a bounding box regression loss function, and the loss function of the three-dimensional detection branch includes a mean absolute error loss function. The loss weight in the mean absolute error loss function is a preset weight.

[0016] On the other hand, the present application provides a three-dimensional detection method, including:

[0017] Obtain the image to be detected;

[0018] Input the image to be detected into the 3D detection model trained by the training method of the 3D detection model described in the first aspect, and obtain the recognition result of the target object in the image to be detected output by the 3D detection model. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0019] On the other hand, the present application provides a training device for a 3D detection model, including:

[0020] An acquisition module, configured to acquire a detection image, where the detection image is marked with camera data and laser point cloud data of the detected object. The laser point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width and height of the detected object, as well as the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image.

[0021] A training module, configured to input the detection image into an initial 3D detection model to train the initial 3D detection model. The initial 3D detection model is a YOLO5 model, and the network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch. The 2D detection branch is used to detect the camera data in the image, and the 3D detection branch is used to detect the laser point cloud data in the image.

[0022] The acquisition module is further configured to obtain a 3D detection model when a training termination condition is reached. The 3D detection model is used to receive an image to be detected and output the recognition result of the target object in the image to be detected. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0023] On the other hand, the present application provides a 3D detection device, including:

[0024] An acquisition module, configured to acquire an image to be detected;

[0025] A processing module, configured to input the image to be detected into the 3D detection model trained by the training method of the 3D detection model according to any one of claims 1-6, and obtain the recognition result of the target object in the image to be detected output by the 3D detection model. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0026] On the other hand, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0027] The memory stores computer-executable instructions;

[0028] The processor executes the computer-executable instructions stored in the memory to implement the training method of the 3D detection model as described in the first aspect, or to implement the 3D detection method as described in the second aspect.

[0029] On the other hand, the present application provides a computer-readable storage medium storing computer-executable instructions, which when executed, cause a computer to execute the training method of the 3D detection model as described in the first aspect, or to execute the 3D detection method as described in the second aspect.

[0030] On the other hand, the present application provides a computer program product including a computer program, which when executed by a processor, implements the training method of the 3D detection model as described in the first aspect, or executes the 3D detection method as described in the second aspect.

[0031] The training method of the 3D detection model provided by the embodiments of the present application is used to train an initial 3D detection model according to a detection image of camera data and laser point cloud data with detected objects already annotated. The initial 3D detection model is a YOLO5 model, and the network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch. The 2D detection branch is used to detect camera data in the image, and the 3D detection branch is used to detect laser point cloud data in the image. When the training termination condition is reached, a 3D detection model is obtained. The 3D detection model is used to receive a to-be-detected image and output a recognition result of the target object in the to-be-detected image. The recognition result of the target object at least includes coordinate information of the target object in the camera coordinate system, coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0032] In this way, the 3D detection result of the target object can be obtained without relying on the internal and external parameters of the camera, avoiding the problem that it is difficult to obtain accurate estimations of the external parameters of each sensor due to the large fluctuation range of the external parameters of each sensor during the movement of the truck. Through experimental verification, the 3D detection model trained by using the method provided by the embodiments of the present application can identify targets within 200 m in the front and back and 50 m on the left and right, and the relative error accuracy is <3%, having high stability and safety. Therefore, the training method of the 3D detection model provided by the embodiments of the present application can improve the visual 3D detection effect of the driverless truck. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0034] Figure 1A schematic diagram of an application scenario of the training method for the three-dimensional detection model provided by this application;

[0035] Figure 2 A flowchart of the training method for the three-dimensional detection model provided by an embodiment of this application;

[0036] Figure 3 A flowchart of the three-dimensional detection method provided by an embodiment of this application;

[0037] Figure 4 A schematic diagram of the training device for the three-dimensional detection model provided by an embodiment of this application;

[0038] Figure 5 A schematic diagram of the three-dimensional detection device provided by an embodiment of this application;

[0039] Figure 6 A schematic diagram of the electronic device provided by an embodiment of this application.

[0040] Through the above-mentioned drawings, specific embodiments of the present disclosure have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present disclosure in any way, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0041] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0042] In the description of this application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality" means two or more unless otherwise specifically defined.

[0043] In the field of autonomous driving, vision-based monocular three-dimensional (3D) detection remains a highly challenging task. Currently, mainstream monocular 3D detection methods are all based on the camera coordinate system. First, the 3D target ground truth in the vehicle body coordinate system is transformed to the camera coordinate system through the extrinsic parameters, and a deep learning model is trained in the camera coordinate system. At the model deployment stage, the model infers the 3D information of the target in the camera coordinate system, and then through the extrinsic parameter matrix of the camera, the target is transformed to the vehicle body coordinate system for downstream fusion algorithms to use. In this process, accurate intrinsic and extrinsic model parameters are required.

[0044] As an important branch in the field of autonomous driving, driverless trucks have their unique scenario characteristics. That is, for driverless trucks, the perception sensors are generally installed at the front of the truck. Due to the relatively soft suspension of the truck, the extrinsic parameters of each sensor fluctuate within a large range during the movement of the truck, so it is very difficult to obtain accurate estimates of the extrinsic parameters of each sensor. This brings particularly difficult troubles to the vision 3D detection of driverless trucks.

[0045] Based on this, the present application provides a training method for a three-dimensional detection model, a three-dimensional detection method, a device, and a device. The training method for the three-dimensional detection model includes: obtaining a detection image, which is marked with the camera data and the laser point cloud data of the detected object. The laser point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, as well as the length, width, and height of the detected object, and the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, as well as the category information of the detected object and the position information of the detected object in the captured image; inputting the detection image into an initial three-dimensional detection model to train the initial three-dimensional detection model. The initial three-dimensional detection model is the YOLO5 model, and the network layer of the YOLO5 model includes a two-dimensional detection branch and a three-dimensional detection branch. The two-dimensional detection branch is used to detect the camera data in the image, and the three-dimensional detection branch is used to detect the laser point cloud data in the image; when the training termination condition is reached, a three-dimensional detection model is obtained. The three-dimensional detection model is used to receive the image to be detected and output the recognition result of the target object in the image to be detected. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0046] The three-dimensional detection result of the target object can be obtained without relying on the internal and external parameters of the camera, avoiding the problem that it is difficult to obtain an accurate estimation of the external parameters of each sensor due to the large fluctuation range of the external parameters of each sensor during the movement of the truck. Through experimental verification, using the three-dimensional detection model provided by the embodiment of the present application, targets within 200m in the front and rear and 50m on the left and right can be recognized, and the relative error accuracy is <3%, having high stability and security. Therefore, the training method of the three-dimensional detection model provided by the present application can improve the visual 3D detection effect of the driverless truck.

[0047] The training method of the three-dimensional detection model provided by the present application is applied to an electronic device, such as a computer, a server, a control device on a driverless truck, etc. Figure 1 This is an application schematic diagram of the training method of the three-dimensional detection model provided by the present application. In the figure, the electronic device acquires a detection image, and the camera data and lidar point cloud data of the detected object are marked in the detection image. The detection graph is input into the initial three-dimensional detection model to train the initial three-dimensional detection model. When the training termination condition is reached, a three-dimensional detection model is obtained.

[0048] Please refer to Figure 2 , an embodiment of the present application provides a training method for a three-dimensional detection model, including:

[0049] S210, acquire a detection image, in which the camera data and lidar point cloud data of the detected object are marked. The lidar point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width, and height of the detected object, as well as the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image.

[0050] The detection image is acquired by a collection vehicle, and the data marked in the detection image is also acquired by the collection vehicle and then input into the electronic device. The data is marked in the detection image manually by a staff member or by the electronic device.

[0051] The collection vehicle should include at least one camera and at least one lidar. The lidar is mainly used to obtain 3D information of the objects around the vehicle, and the lidar may not be required in actual deployment.

[0052] In one example, 7 cameras and 1 mechanical lidar are installed on the acquisition vehicle. The cameras are distributed as 3 front-facing cameras, 2 left and right cameras, and 1 each at the lower left and lower right. If the camera and lidar installation methods provided by this example are used to collect data and train the initial 3D detection model, when training the initial 3D detection model on other similar vehicles, the installation methods of the cameras and lidars on other similar vehicles should be as identical as possible to those provided by this example. The models of the sensors (sensors include cameras and lidars) at the same position should be kept consistent, and the deviation of the installation position of the sensors should not exceed 2 cm (centimeters), and the angular deviation of the three direction angles should not exceed 2°. In our practice process, the camera and lidar installation methods provided by this example can obtain satisfactory target recognition accuracy and have high stability and security.

[0053] What the camera collects is two-dimensional information, that is, camera data. The camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image. Specifically, the camera data includes the 2D information of the detected object, the category cls of the detected object, and the position information [xmin, ymin, xmax, ymax] of the detected object in the detection image. Among them, [xmin, ymin] is the position of the upper left corner point of the detected object in the detection image, and [xmax, ymax] is the position of the lower right corner of the detected object in the detection image.

[0054] What the lidar collects is three-dimensional information, that is, lidar point cloud data. The lidar point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width and height of the detected object, as well as the heading angle of the detected object. The lidar point cloud data is used to label the 3D information of the detected object in the real world, including [x, y, z, w, h, l, royt]. Among them, [x, y, z] represents the position information of the detected object around the vehicle in the vehicle body coordinate system, [w, h, l] is the information of the length, width and height of the detected object, and royt is the heading angle of the detected object.

[0055] For any detected object, the data collected by each sensor needs to be associated. That is, for any detected object to be recognized, the following information [cls, xmin, ymin, xmax, ymax, x, y, z, w, h, l, royt] should be included.

[0056] The acquisition vehicle should collect as many detection images as possible. The types and quantities of the detected objects in the detection images should be as numerous as possible to better train the initial 3D detection model. The data acquisition scenarios should be as rich as possible, including but not limited to daytime, night, wind, frost, rain, snow, urban roads, and rural roads.

[0057] S220, input the detected image into the initial 3D detection model to train the initial 3D detection model. The initial 3D detection model is the YOLO5 model. The network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch. The 2D detection branch is used to detect camera data in the image, and the 3D detection branch is used to detect laser point cloud data in the image.

[0058] The initial 3D detection model is the YOLO5 model, which is used to extract image features and then output the recognition result of the image.

[0059] The network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch.

[0060] The 2D detection branch is used to detect 2D information in the image, that is, camera data. The 3D detection branch is used to detect 3D information in the image, that is, laser point cloud data. The output of the initial 3D detection model is the result obtained after combining the output results of the 2D detection branch and the 3D detection branch and then processing.

[0061] More specifically, the 3D detection branch includes three 3D detection heads. Among them, each 3D detection head includes three prior boxes. Each 3D detection head is used to predict the coordinate information of the detected object in the vehicle body coordinate system, the length, width and height of the detected object, the first angle and the second angle of the detected object. Among them, the first angle and the second angle of the detected object are used to determine the heading angle of the detected object.

[0062] The 2D detection branch includes three 2D detection heads. Each 2D detection head includes three prior boxes. Each 2D detection head is used to predict the coordinate information of the detected object in the camera coordinate system, the length and width of the detected object, the category information of the detected object, and the number of detected objects of the target type.

[0063] That is, the basic backbone uses the original YOLO5 model, and the output of 3 additional detection heads (head layer) is added to the network layer (Neck layer) of the original YOLO5 model. The 3 additional head layers are used to detect 3D information. The 3 additional head layers are three 3D detection heads, corresponding to the head layer used for 2D detection in the original YOLO5 model.

[0064] The size of the input image of the initial 3D detection model is 640*384. The backbone network selects Yolov5s, and 2 branches are generated in the Neck layer of the network, that is, the 2D detection branch and the 3D detection branch.

[0065] The two-dimensional detection branch generates 3 head layers with sizes of 48*80*(3*(5+cls)), 24*40*(3*(5+cls)), and 12*20*(3*(5+cls)) respectively. Among them, the 5 in the head layer represents the position information [x, y, w, h] of the detected object and the classification information of whether the detected object is the background. The 3 in the head layer represents the number of anchors (detection boxes). Each point in a head layer generates 3 candidate detected object information, and cls represents the number of types of detected objects to be recognized.

[0066] The three-dimensional detection branch also generates 3 head layers with sizes of 48*80*(3*8), 24*40*(3*8), and 12*20*(3*8) respectively. Among them, the 3 in the head layer corresponds to the number of anchors in the two-dimensional detection branch, and the 8 in the head layer is the predicted position information x, y, z, w, h, l, royt_x, royt_y of the candidate detected objects. Among them, x, y, z, w, h, l are the 3D information of the detected object directly regressed by the model in the vehicle body coordinate system without any post-processing. The heading angle is

[0067] The output layer of this initial three-dimensional detection model is used to merge the output results of this two-dimensional detection branch and this three-dimensional detection branch. That is, the two-dimensional detection branch and the three-dimensional detection branch are merged together to obtain 48*80*(3*(5+cls+8)), 24*40*(3*(5+cls+8)), and 12*20*(3*(5+cls+8)).

[0068] The output layer of this initial three-dimensional detection model is also used to filter the background image information in the merged output results according to the output threshold to obtain the detected object recognition results. The detected object recognition results include the coordinate information of each predicted detected object in the camera coordinate system and the coordinate information of each predicted detected object in the vehicle body coordinate system. That is, by setting a threshold (usually selected as 0.5), most of the background image information is filtered out. Finally, n*(5+cls+8) is obtained, where n is the number of detected objects greater than the threshold.

[0069] In an optional embodiment, the output layer is also used to filter the overlapping boxes in the merged output results based on the non-maximum suppression algorithm. That is, after obtaining n*(5+cls+8), some overlapping boxes are filtered out through the non-maximum suppression algorithm (Non-Maximum Suppress, abbreviated as NMS).

[0070] In an optional embodiment, the loss function of the two-dimensional detection branch includes a bounding box regression loss function (Giou loss function), and the category classification selects cross-entropy loss.

[0071] The loss function of the 3D detection branch includes the mean absolute error loss function (L1 loss function), and the loss weight in the mean absolute error loss function is a preset weight. Optionally, the loss weight in the loss function of the 3D detection branch needs to be multiplied by 0.1 to balance the losses of different recognition tasks.

[0072] S230. When the training termination condition is reached, a 3D detection model is obtained. The 3D detection model is used to receive a to-be-detected image and output the recognition result of the target object in the to-be-detected image. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0073] The training termination condition includes any one or more of the following: the number of training times reaches the preset number, the training duration reaches the preset duration, and the error between the output result of the initial 3D detection model and the information of the detected object marked in the detection image is within the preset range.

[0074] When the training termination condition is reached, a 3D detection model is obtained.

[0075] In use, the 3D detection model is used to receive a to-be-detected image and output the recognition result of the target object in the to-be-detected image. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object. The quantity of the target object refers to the quantity of the target object under each category. For example, if the quantity of the target object under the category of driverless truck is 3, then there are 3 driverless trucks in the to-be-detected image. The to-be-detected image may also include target objects of other categories, such as trees, buildings, etc.

[0076] In summary, the training method of the 3D detection model provided in this embodiment is used to train an initial 3D detection model according to the detection image of the camera data and the laser point cloud data with the detected objects already marked. The initial 3D detection model is the YOLO5 model. The network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch. The 2D detection branch is used to detect the camera data in the image, and the 3D detection branch is used to detect the laser point cloud data in the image. When the training termination condition is reached, a 3D detection model is obtained. The 3D detection model is used to receive a to-be-detected image and output the recognition result of the target object in the to-be-detected image. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0077] In this way, the three-dimensional detection result of the target object can be obtained without relying on the internal and external parameters of the camera, avoiding the problem that it is difficult to obtain accurate estimation of the external parameters of each sensor due to the large fluctuation range of the external parameters of each sensor during the movement of the truck. Through experimental verification, the three-dimensional detection model trained by using the method provided in this embodiment can identify targets within 200m in the front and rear and 50m on the left and right, and the relative error accuracy is <3%, having high stability and safety. Therefore, the training method of the three-dimensional detection model provided in this embodiment can improve the visual 3D detection effect of the driverless truck.

[0078] Please refer to Figure 3 , an embodiment of the present application further provides a three-dimensional detection method, including:

[0079] S310, obtain an image to be detected.

[0080] The image to be detected is an image taken in real time during the actual driving of the driverless truck. The two-dimensional information and three-dimensional information of the target object in the image need to be recognized to perform the driving planning of the driverless truck according to the two-dimensional information and three-dimensional information of the target object.

[0081] S320, input the image to be detected into the trained three-dimensional detection model to obtain the target object recognition result output by the three-dimensional detection model for the image to be detected. The target object recognition result at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0082] The three-dimensional detection model is a three-dimensional detection model trained according to the training method of the three-dimensional detection model provided in any one of the above embodiments. The three-dimensional detection model is used to receive the image to be detected and output the target object recognition result for the image to be detected. The target object recognition result at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object. The quantity of the target object refers to the quantity of the target object under each category. For example, if the quantity of the target object under the driverless truck category is 3, then there are 3 driverless trucks in the image to be detected. The image to be detected may also include target objects of other categories, such as trees, buildings, etc.

[0083] After obtaining the image to be detected in the three-dimensional detection method provided in this embodiment, the image to be detected is input into the three-dimensional detection model trained according to the training method of the three-dimensional detection model provided in any one of the above embodiments. The three-dimensional detection model outputs the target object recognition result for the image to be detected. The target object recognition result at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0084] In this way, the three-dimensional detection result of the target object can be obtained without relying on the internal and external parameters of the camera, avoiding the problem that it is difficult to obtain accurate external parameter estimates of each sensor due to the large fluctuation range of the external parameters of each sensor during the movement of the truck. Through experimental verification, using the three-dimensional detection model provided in this embodiment, targets within 200m in the front and rear and 50m on the left and right can be recognized, and the relative error accuracy is <3%, with high stability and safety. Therefore, the three-dimensional detection method provided in this embodiment can improve the visual 3D detection effect of the driverless truck.

[0085] Please refer to Figure 4 , an embodiment of the present application further provides a training device 10 for a three-dimensional detection model, including:

[0086] An acquisition module 11, configured to acquire a detection image, in which the camera data and laser point cloud data of the detected object are marked. The laser point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width, and height of the detected object, as well as the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image.

[0087] A training module 12, configured to input the detection image into an initial three-dimensional detection model to train the initial three-dimensional detection model. The initial three-dimensional detection model is a YOLO5 model, and the network layer of the YOLO5 model includes a two-dimensional detection branch and a three-dimensional detection branch. The two-dimensional detection branch is used to detect the camera data in the image, and the three-dimensional detection branch is used to detect the laser point cloud data in the image.

[0088] The acquisition module 11 is further configured to obtain a three-dimensional detection model when the training termination condition is reached. The three-dimensional detection model is used to receive a to-be-detected image and output the recognition result of the target object in the to-be-detected image. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0089] The three-dimensional detection branch includes three three-dimensional detection heads. Among them, each three-dimensional detection head includes three prior boxes, and each three-dimensional detection head is used to predict the coordinate information of the detected object in the vehicle body coordinate system, the length, width, and height of the detected object, the first angle and the second angle of the detected object; among them, the first angle and the second angle of the detected object are used to determine the heading angle of the detected object.

[0090] The two-dimensional detection branch includes three two-dimensional detection heads. Each two-dimensional detection head includes three prior boxes, and each two-dimensional detection head is used to predict the coordinate information of the detected object in the camera coordinate system, the length and width of the detected object, the category information of the detected object, and the quantity of the detected object of the target type.

[0091] The output layer of the initial three-dimensional detection model is used to merge the output results of the two-dimensional detection branch and the three-dimensional detection branch; it is also used to filter the background image information in the merged output results according to an output threshold to obtain a detection object recognition result, where the detection object recognition result includes the coordinate information of each predicted detection object in the camera coordinate system and the coordinate information of each predicted detection object in the vehicle body coordinate system.

[0092] The output layer is also used to filter overlapping boxes in the merged output results based on the non-maximum suppression algorithm.

[0093] The loss function of the two-dimensional detection branch includes a bounding box regression loss function, and the loss function of the three-dimensional detection branch includes a mean absolute error loss function, where the loss weight in the mean absolute error loss function is a preset weight.

[0094] Please refer to Figure 5 , an embodiment of the present application further provides a three-dimensional detection device 21, including:

[0095] An acquisition module 21, configured to acquire an image to be detected.

[0096] A processing module 22, configured to input the image to be detected into the trained three-dimensional detection model to obtain a recognition result of the target object in the image to be detected output by the three-dimensional detection model, where the recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

[0097] Please refer to Figure 6 , an embodiment of the present application further provides an electronic device 30, including a processor 31 and a memory 32 communicatively connected to the processor 31. The memory 32 stores computer-executable instructions, and the processor 31 executes the computer-executable instructions stored in the memory 32 to implement the training method of the three-dimensional detection model provided in any of the above embodiments, or to implement the three-dimensional detection method provided in any of the above embodiments.

[0098] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the instructions are executed, the computer-executable instructions are used by the processor to implement the training method of the three-dimensional detection model provided in any of the above embodiments, or to implement the three-dimensional detection method provided in any of the above embodiments.

[0099] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the training method of the three-dimensional detection model provided in any of the above embodiments, or implements the three-dimensional detection method provided in any of the above embodiments.

[0100] It should be noted that the above computer-readable storage medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc. It can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0101] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of another identical element in the process, method, article or device comprising such element.

[0102] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0104] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows or blocks Figure 1 or blocks.

[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows or blocks Figure 1 or blocks.

[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows or blocks Figure 1 or blocks.

[0107] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present application.

Claims

1. A training method for a three-dimensional detection model, characterized in that, it includes: Obtain a detection image, in which the camera data and laser point cloud data of the detected object are marked. The laser point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width and height of the detected object, as well as the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image; Input the detection image into the initial three-dimensional detection model to train the initial three-dimensional detection model. The initial three-dimensional detection model is a YOLO5 model. The network layer of the YOLO5 model includes a two-dimensional detection branch and a three-dimensional detection branch. The two-dimensional detection branch is used to detect the camera data in the image, and the three-dimensional detection branch is used to detect the laser point cloud data in the image; When the training termination condition is reached, a three-dimensional detection model is obtained. The three-dimensional detection model is used to receive the image to be detected and output the recognition result of the target object in the image to be detected. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object; The three-dimensional detection branch includes three three-dimensional detection heads. Among them, each three-dimensional detection head includes three prior boxes. Each three-dimensional detection head is used to predict the coordinate information of the detected object in the vehicle body coordinate system in the detection image, the length, width and height of the detected object, the first angle and the second angle of the detected object; among them, the first angle and the second angle of the detected object are used to determine the heading angle of the detected object; The two-dimensional detection branch includes three two-dimensional detection heads. Each two-dimensional detection head includes three prior boxes. Each two-dimensional detection head is used to predict the coordinate information of the detected object in the camera coordinate system in the detection image, the length and width of the detected object, the category information of the detected object, and the quantity of the detected object of the target type; The output layer of the initial three-dimensional detection model is used to merge the output results of the two-dimensional detection branch and the three-dimensional detection branch; It is also used to filter the background image information in the merged output result according to the output threshold to obtain the detection object recognition result. The detection object recognition result includes the coordinate information of each predicted detection object in the camera coordinate system and the coordinate information of each predicted detection object in the vehicle body coordinate system.

2. The method according to claim 1, characterized in that, The output layer is also used to filter the overlapping boxes in the merged output result based on the non-maximum suppression algorithm.

3. The method according to any one of claims 1-2, characterized in that, The loss function of the two-dimensional detection branch includes a bounding box regression loss function, and the loss function of the three-dimensional detection branch includes a mean absolute error loss function. The loss weight in the mean absolute error loss function is a preset weight.

4. A three-dimensional detection method, characterized in that, it includes: Obtain the image to be detected; Input the image to be detected into the 3D detection model trained by the training method of the 3D detection model according to any one of claims 1-3, and obtain the recognition result of the target object in the image to be detected output by the 3D detection model. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object.

5. A training device for a 3D detection model Characterized in that It includes: An acquisition module, configured to acquire a detection image, where the detection image is marked with camera data and lidar point cloud data of a detected object. The lidar point cloud data includes the coordinate information of the detected object in the vehicle body coordinate system, and also includes the length, width and height of the detected object, and the heading angle of the detected object; the camera data includes the coordinate information of the detected object in the camera coordinate system, and also includes the category information of the detected object and the position information of the detected object in the captured image. A training module, configured to input the detection image into an initial 3D detection model to train the initial 3D detection model. The initial 3D detection model is a YOLO5 model. The network layer of the YOLO5 model includes a 2D detection branch and a 3D detection branch. The 2D detection branch is used to detect the camera data in the image, and the 3D detection branch is used to detect the lidar point cloud data in the image. The acquisition module is further configured to obtain a 3D detection model when the training termination condition is reached. The 3D detection model is used to receive the image to be detected and output the recognition result of the target object in the image to be detected. The recognition result of the target object at least includes the coordinate information of the target object in the camera coordinate system, the coordinate information of the target object in the vehicle body coordinate system, the category and quantity of the target object. The 3D detection branch includes three 3D detection heads. Each 3D detection head includes three prior boxes. Each 3D detection head is used to predict the coordinate information of the detected object in the vehicle body coordinate system, the length, width and height of the detected object, the first angle and the second angle of the detected object. The first angle and the second angle of the detected object are used to determine the heading angle of the detected object. The 2D detection branch includes three 2D detection heads. Each 2D detection head includes three prior boxes. Each 2D detection head is used to predict the coordinate information of the detected object in the camera coordinate system, the length and width of the detected object, the category information of the detected object, and the quantity of the detected object of the target type. The output layer of the initial 3D detection model is used to merge the output results of the 2D detection branch and the 3D detection branch. It is also used to filter the background image information in the merged output result according to an output threshold to obtain a detected object recognition result. The detected object recognition result includes the coordinate information of each predicted detected object in the camera coordinate system and the coordinate information of each predicted detected object in the vehicle body coordinate system.

6. A 3D detection device Characterized in that It includes: An acquisition module, configured to acquire an image to be detected; A processing module, configured to input the image to be detected into a three-dimensional detection model obtained by training with the training method of the three-dimensional detection model according to any one of claims 1-3, and obtain an object recognition result in the image to be detected output by the three-dimensional detection model, where the object recognition result at least includes coordinate information of the object in the camera coordinate system, coordinate information of the object in the vehicle body coordinate system, the category and quantity of the object.

7. An electronic device, characterized in that it includes: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the training method of the three-dimensional detection model according to any one of claims 1 to 3, or to implement the three-dimensional detection method according to claim 4.

Citation Information

Patent Citations

  • Point cloud prediction model generation method, pose estimation method and pose estimation device

    CN112652016A

  • Training data generation method and device, model training method and device, model detection method and device and electronic equipment

    CN114997264A