Training of 3D Object Detection Model and 3D Object Detection Method and Device

By combining the knowledge distillation technology of teacher networks and student networks, the loss function value is calculated and parameters are adjusted. The trained three-dimensional object detection model improves the accuracy of three-dimensional object detection in autonomous driving and intelligent robots, and solves the problem of insufficient detection accuracy in the existing technology.

CN115880684BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211259102.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2025-07-22
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

The existing three-dimensional object detection technology has the problem of insufficient detection accuracy in autonomous driving and intelligent robots, especially when combining point cloud data and binocular images, it is difficult to effectively improve the accuracy of the detection model.

Method used

By obtaining the detection box labeling information of multiple data pairs, combining the teacher network and the student network, knowledge distillation technology is performed, loss function values are calculated, parameters of the student network are adjusted, and a three-dimensional object detection model is trained.

Benefits of technology

The detection accuracy of the three-dimensional object detection model in binocular images is improved, especially at the output level and feature level combined with knowledge distillation, which improves the accuracy of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880684B_ABST
    Figure CN115880684B_ABST
Patent Text Reader

Abstract

The present disclosure provides a training method for a three-dimensional object detection model and a three-dimensional object detection method, which relate to the fields of artificial intelligence technologies such as computer vision, image processing, deep learning, and augmented reality, and can be applied to scenarios such as autonomous driving and smart cities. The training method for the three-dimensional object detection model includes: obtaining a plurality of data pairs and detection box annotation information of the plurality of data pairs; obtaining a teacher network and a student network; inputting the point cloud data in the data pairs into the teacher network and inputting the binocular images into the student network to obtain first detection box information and second detection box information; obtaining target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pairs; obtaining a first loss function value according to the target detection box information and the second detection box information, and obtaining a second loss function value according to the second detection box information and the detection box annotation information; adjusting the parameters of the student network according to the loss function values to obtain a three-dimensional object detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, image processing, deep learning, and augmented reality, and can be applied to scenarios such as autonomous driving and smart cities. A training method for a three-dimensional object detection model, a three-dimensional object detection method, device, system, electronic device, and readable storage medium are provided. Background Art

[0002] With the development of artificial intelligence technologies, three-dimensional object detection technologies are widely applied in various fields. For example, during the movement of an intelligent robot or an autonomous driving vehicle, three-dimensional object detection technologies can be used to detect surrounding obstacles, so as to avoid the obstacles. Summary of the Invention

[0003] According to a first aspect of the present disclosure, a training method for a three-dimensional object detection model is provided, including: obtaining a plurality of data pairs and detection box annotation information of the plurality of data pairs, each data pair including point cloud data and a binocular image corresponding to the same scene; obtaining a teacher network and a student network, where the teacher network is a detection network corresponding to the point cloud data, and the student network is a detection network corresponding to the binocular image; inputting the point cloud data in the data pair into the teacher network and the binocular image into the student network to obtain first detection box information output by the teacher network and second detection box information output by the student network; obtaining target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair; obtaining a first loss function value according to the target detection box information and the second detection box information, and obtaining a second loss function value according to the second detection box information and the detection box annotation information; adjusting parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model.

[0004] According to a second aspect of the present disclosure, a three-dimensional object detection method is provided, including: obtaining a binocular image to be detected; inputting the binocular image to be detected into the three-dimensional object detection model, and obtaining a detection result of the binocular image to be detected according to an output result of the three-dimensional object detection model.

[0005] According to a third aspect of the present disclosure, there is provided a training device for a three-dimensional object detection model, including: a first acquisition unit configured to acquire a plurality of data pairs and detection box annotation information of the plurality of data pairs, each data pair including point cloud data and a binocular image corresponding to the same scene; a second acquisition unit configured to acquire a teacher network and a student network, the teacher network being a detection network corresponding to the point cloud data, and the student network being a detection network corresponding to the binocular image; a first processing unit configured to input the point cloud data in the data pair into the teacher network and input the binocular image into the student network to obtain first detection box information output by the teacher network and second detection box information output by the student network; a second processing unit configured to obtain target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair; a calculation unit configured to obtain a first loss function value according to the target detection box information and the second detection box information, and obtain a second loss function value according to the second detection box information and the detection box annotation information; and a training unit configured to adjust parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model.

[0006] According to a fourth aspect of the present disclosure, there is provided a three-dimensional object detection device, including: a third acquisition unit configured to acquire a binocular image to be detected; and a detection unit configured to input the binocular image to be detected into the three-dimensional object detection model and obtain a detection result of the binocular image to be detected according to an output result of the three-dimensional object detection model.

[0007] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0008] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0009] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0010] As can be seen from the above technical solutions, when training the 3D detection model, the present disclosure combines the target detection frame information obtained from the first detection frame information, the second detection frame information, and the detection frame annotation information to calculate the loss function, thereby completing the knowledge distillation at the output level, and improving the accuracy of the trained 3D object detection model in performing 3D detection on binocular images.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0014] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;

[0016] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0017] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0018] Figure 6 is a schematic diagram according to the sixth embodiment of the present disclosure;

[0019] Figure 7 is a block diagram of an electronic device for implementing the training of the 3D object detection model and / or the 3D object detection method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and mechanisms are omitted below.

[0021] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. As Figure 1 shown, the training method of the 3D object detection model in this embodiment specifically includes the following steps:

[0022] S101. Obtain multiple data pairs and the detection box annotation information of the multiple data pairs. Each data pair includes point cloud data and binocular images corresponding to the same scene.

[0023] S102. Obtain a teacher network and a student network. The teacher network is a detection network for the corresponding point cloud data, and the student network is a detection network for the corresponding binocular images.

[0024] S103. Input the point cloud data in the data pair into the teacher network and the binocular images into the student network to obtain the first detection box information output by the teacher network and the second detection box information output by the student network.

[0025] S104. Obtain the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair.

[0026] S105. Obtain a first loss function value according to the target detection box information and the second detection box information, and obtain a second loss function value according to the second detection box information and the detection box annotation information.

[0027] S106. Adjust the parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model.

[0028] The training method of the three-dimensional object detection model in this embodiment is based on the knowledge distillation technology. Through the teacher network corresponding to the point cloud data, a model for three-dimensional object detection of binocular images is distilled. Since in this embodiment, when training the three-dimensional detection model, the target detection box information obtained from the first detection box information, the second detection box information, and the detection box annotation information is combined to calculate the damage function, thus completing the knowledge distillation at the output level, which can improve the accuracy of the trained three-dimensional object detection model when performing three-dimensional detection on binocular images.

[0029] Each data pair obtained by this embodiment in S101 includes point cloud data and binocular images corresponding to the same scene. Among them, the detection box annotation information of the data pair includes the center point of the detection box, the size (length, width, and height) of the detection box, and the orientation angle of the detection box. In addition, this embodiment can also obtain the category annotation information of the objects included in the scene corresponding to each data pair.

[0030] After this embodiment executes S101 to obtain multiple data pairs and the detection box annotation information of the multiple data pairs, it executes S102 to obtain the teacher network and the student network.

[0031] The teacher network obtained by executing S102 in this embodiment is a detection network for corresponding point cloud data, which can output the detection box information of the objects included in the input point cloud data, and can further output the category information of the objects and their corresponding category confidence scores.

[0032] The teacher network obtained by executing S102 in this embodiment sequentially includes a three-dimensional feature extraction layer, a bird's-eye view feature extraction layer, and a three-dimensional head network. Among them, the three-dimensional feature extraction layer (such as LiDAR Backbone, a lidar backbone network) is used to obtain three-dimensional features from the point cloud data, the bird's-eye view (Bird’s Eye View, BEV) feature extraction layer (such as BEV Encoder, a BEV encoder) is used to obtain bird's-eye view features from the three-dimensional features, and the three-dimensional head network (such as a multi-layer perceptron MLP) is used to obtain the detection box information, category information, and category confidence scores of the objects from the bird's-eye view features.

[0033] The student network obtained by executing S102 in this embodiment is a detection network for corresponding binocular images, which can output the detection box information of the objects included in the input binocular images, and can further output the category information of the objects and their corresponding category confidence scores.

[0034] The student network obtained by executing S102 in this embodiment sequentially includes a two-dimensional feature extraction layer, a bird's-eye view feature extraction layer, and a three-dimensional head network. Among them, the two-dimensional feature extraction layer (such as Stereo Backbone, a stereo backbone network) is used to obtain three-dimensional features from the binocular images, the bird's-eye view (Bird’s Eye View, BEV) feature extraction layer (such as BEV Encoder) is used to obtain bird's-eye view features from the three-dimensional features, and the three-dimensional head network (such as a multi-layer perceptron MLP) is used to obtain the detection box information, category information, and category confidence scores of the objects from the bird's-eye view features.

[0035] After obtaining the teacher network and the student network by executing S102 in this embodiment, S103 is executed to input the point cloud data in the data pair into the teacher network, and input the binocular image in the data pair into the student network, to obtain the first detection box information output by the teacher network and the second detection box information output by the student network.

[0036] When executing S103 in this embodiment, the point cloud data and the binocular image corresponding to the same data pair are respectively input into the teacher network and the student network, so as to obtain the first detection box information output by the teacher network for the point cloud data and the second detection box information output by the student network for the binocular image.

[0037] When executing S103 in this embodiment, the category information and category confidence scores of the objects can further be obtained.

[0038] After obtaining the first detection box information and the second detection box information in S103, this embodiment executes S104 to obtain the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair.

[0039] When this embodiment executes S104 to obtain the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair, an optional implementation method that can be adopted is: according to the first detection box information, the second detection box information, and the detection box annotation information, obtain a detection box information group corresponding to the same object, and the generated detection box information group includes the detection box information corresponding to the object in the first detection box information, the detection box information corresponding to the object in the second detection box information, and the detection box information corresponding to the object in the detection box annotation information; according to the detection box information group, obtain candidate target detection box information; according to the candidate target detection box information, obtain the target detection box information.

[0040] That is to say, this embodiment obtains the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information, achieving the purpose of knowledge distillation at the output level of the network, so that when training the student network, the parameter adjustment can be carried out in combination with the loss function value obtained according to the target detection box information, thereby improving the training effect of the trained 3D object detection model.

[0041] Among them, when this embodiment executes S104 to obtain a detection box information group corresponding to the same object according to the first detection box information, the second detection box information, and the detection box annotation information, an optional implementation method that can be adopted is: calculate the first intersection over union (IoU, Intersection over Union) between the detection box annotation information and the first detection box information, and the second intersection over union between the detection box annotation information and the second detection box information; in the case where it is determined that the first intersection over union is less than the first threshold and the second intersection over union is less than the second threshold, use the first detection box information, the second detection box information, and the detection box annotation information as the detection box information group corresponding to the same object.

[0042] That is to say, this embodiment is based on the intersection over union between the detection box information to find the detection box information group corresponding to the same object, and further realizes the purpose of distillation at the output level according to the detection box information group.

[0043] When this embodiment executes S104 to obtain candidate target detection box information based on the detection box information group, it may further include the following content: obtaining a first vector according to the center of the first detection box in the first detection box information and the center of the second detection box in the second detection box information; obtaining a second vector according to the center of the second detection box in the second detection box information and the center of the detection box annotation in the detection box annotation information; in the case where the included angle between the first vector and the second vector is an acute angle, obtaining candidate target detection box information according to the first detection box information, the second detection box information and the detection box annotation information in the detection box information group.

[0044] That is to say, before generating the target detection box information in this embodiment, it may also determine whether the obtained detection box information group is accurate according to the center point of the first detection box, the center point of the second detection box and the center point of the labeled detection box, so as to improve the accuracy of the obtained target detection box information.

[0045] When this embodiment executes S104 to obtain candidate target detection box information according to the first detection box information, the second detection box information and the detection box annotation information in the detection box information group, an optional implementation method that can be adopted is: obtaining the center point of the candidate target detection box according to the center point of the first detection box in the first detection box information, the center point of the second detection box in the second detection box information and the center point of the detection box annotation in the detection box annotation information; obtaining the size of the candidate target detection box according to the size of the first detection box in the first detection box information, the size of the second detection box in the second detection box information and the size of the detection box annotation in the detection box annotation information; obtaining the orientation angle of the candidate target detection box according to the orientation angle of the first detection box in the first detection box information, the orientation angle of the second detection box in the second detection box information and the orientation angle of the detection box annotation in the detection box annotation information; obtaining candidate target detection box information according to the center point of the candidate target detection box, the size of the candidate target detection box and the orientation angle of the candidate target detection box.

[0046] Among them, when this embodiment executes S104 to obtain the center point of the candidate target detection box according to the center point of the first detection box in the first detection box information, the center point of the second detection box in the second detection box information and the center point of the detection box annotation in the detection box annotation information, the following calculation formula can be adopted:

[0047] M center =(T center -S center )*(G center -S center )

[0048] In the formula: M center represents the center point of the candidate target detection box; T center represents the center point of the first detection box; G center represents the center point of the detection box annotation; Scenter Denote the center point of the second detection box.

[0049] In this embodiment, when performing S104 to obtain the candidate target detection box size according to the first detection box size in the first detection box information, the second detection box size in the second detection box information, and the detection box annotation size in the detection box annotation information, the following calculation formula can be used:

[0050] M size =(T size -S size )*(G size -S size )

[0051] In the formula: M size Denotes the candidate target detection box size; T size Denotes the first detection box size; G center Denotes the detection box annotation size; S size Denotes the second detection box size.

[0052] In this embodiment, when performing S104 to obtain the candidate target detection box orientation angle according to the first detection box orientation angle in the first detection box information, the second detection box orientation angle in the second detection box information, and the detection box annotation orientation angle in the detection box annotation information, the following calculation formula can be used:

[0053] M angle =(T angle -S angle )*(G angle -S angle )

[0054] In the formula: M angle Denotes the candidate target detection box orientation angle; T angle Denotes the first detection box orientation angle; G angle Denotes the detection box annotation orientation angle; S angle Denotes the second detection box orientation angle.

[0055] In this embodiment, when performing S104 to obtain the target detection box information according to the candidate target detection box information, an optional implementation method that can be adopted is: when it is determined that the candidate detection box center point in the candidate target detection box information is greater than or equal to the first threshold, use the first detection box center point in the first detection box information as the target detection box center point; otherwise, use the annotation detection box center point as the target detection box center point.

[0056] When this embodiment executes S104 to obtain the target detection frame information based on the obtained candidate target detection frame information, an optional implementation method that can be adopted is as follows: when it is determined that the size of the candidate detection frame in the candidate target detection frame information is greater than or equal to the second threshold, the size of the first detection frame in the first detection frame information is used as the size of the target detection frame; otherwise, the size of the labeled detection frame is used as the size of the target detection frame.

[0057] When this embodiment executes S104 to obtain the target detection frame information based on the obtained candidate target detection frame information, an optional implementation method that can be adopted is as follows: when it is determined that the orientation angle of the candidate detection frame in the candidate target detection frame information is greater than or equal to the third threshold, the orientation angle of the first detection frame in the first detection frame information is used as the orientation angle of the target detection frame; otherwise, the orientation angle of the labeled detection frame is used as the orientation angle of the target detection frame.

[0058] It can be understood that the first threshold, the second threshold, and the third threshold in this embodiment can be the same. For example, all three thresholds are 0; or three thresholds of different sizes can be set according to actual needs.

[0059] After this embodiment executes S104 to obtain the target detection frame information, it executes S105 to obtain the first loss function value according to the target detection frame information and the second detection frame information, and obtains the second loss function value according to the second detection frame information and the detection frame annotation information.

[0060] When this embodiment executes S105 to obtain the loss function value, the calculation methods of the cross-entropy loss function can be used to calculate the first loss function value and the second loss function value; it can be understood that when this embodiment executes S105, in addition to using the detection frame information of the object, the category information of the object corresponding to the detection frame can also be combined to calculate the loss function.

[0061] After this embodiment executes S105 to obtain the first loss function value and the second loss function value, it executes S106 to adjust the parameters of the student network according to the first loss function value and the second loss function value to obtain a three-dimensional object detection model.

[0062] When this embodiment executes S106 to adjust the parameters of the student network according to the first loss function value and the second loss function value, it can stop adjusting the parameters of the student network when it is determined that the first loss function value and the second loss function value converge simultaneously, and use the student network when the loss function value converges as the three-dimensional object detection model.

[0063] The three-dimensional object detection model obtained by this embodiment executing S106 outputs the detection frame information of the object in the scene corresponding to the input binocular image, and can also output the category information of the object.

[0064] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. As Figure 2 shown, when performing S106 "adjust the parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model" in this embodiment, the following steps are specifically included:

[0065] S201. Obtain a first class confidence score corresponding to the target object according to the teacher network, and obtain a second class confidence score corresponding to the target object according to the student network;

[0066] S202. Obtain a target weight according to the first class confidence score and the second class confidence score;

[0067] S203. Obtain a third loss function value according to the first target feature corresponding to the target object obtained by the teacher network, the second target feature corresponding to the target object obtained by the student network, and the target weight;

[0068] S204. Adjust the parameters of the student network according to the first loss function value, the second loss function value and the third loss function to obtain the three-dimensional object detection model.

[0069] That is to say, when adjusting the parameters of the student network in this embodiment, in addition to performing knowledge distillation at the output level, knowledge distillation at the feature level can also be combined, so as to improve the accuracy when adjusting the student network, and correspondingly improve the training effect of the trained three-dimensional object detection model, so that the three-dimensional object detection model can perform three-dimensional object detection more accurately.

[0070] It can be understood that when performing S201 in this embodiment, the target object can be preset, for example, the object corresponding to the target class is used as the target object, or the target object can also be randomly selected; the number of target objects can be one or multiple.

[0071] When this embodiment performs S202 to obtain the target weight according to the first class confidence score and the second class confidence score, the following calculation formula can be used:

[0072]

[0073] In the formula: M d represents the target weight; T represents a preset parameter; P S represents the second class confidence score; P T denotes the first class confidence score.

[0074] In this embodiment, the first target feature corresponding to the target object obtained according to the teacher network may be the three-dimensional feature output by the three-dimensional feature extraction layer in the teacher network, or the bird's-eye view feature output by the bird's-eye view feature extraction layer in the teacher network, or the three-dimensional feature and the bird's-eye view feature output by the teacher network.

[0075] In this embodiment, the second target feature corresponding to the target object obtained according to the student network may be the three-dimensional feature output by the two-dimensional feature extraction layer in the student network, or the bird's-eye view feature output by the bird's-eye view feature extraction layer in the student network, or the three-dimensional feature and the bird's-eye view feature output by the student network.

[0076] When this embodiment executes S204 to obtain the third loss function value according to the first target feature, the second target feature, and the target weight, the following calculation formula may be adopted:

[0077]

[0078] In the formula: L represents the third loss function value; N represents the number of pixels included in the binocular image; F 3D represents that the target feature is a three-dimensional feature; F BEV represents that the target feature is a bird's-eye view feature; f S represents the second target feature; f T represents the first target feature.

[0079] That is to say, this embodiment combines the category confidence score and obtains the third loss function value according to the target feature, which can improve the accuracy of the third loss function value, and further improve the accuracy when adjusting the parameters of the student network.

[0080] When this embodiment executes S204 to adjust the parameters of the student network according to the first loss function value, the second loss function value, and the third loss function value to obtain the three-dimensional object detection model, it may stop adjusting the parameters of the student network when it is determined that the first loss function value, the second loss function value, and the third loss function value converge simultaneously, and use the student network when the loss function value converges as the three-dimensional object detection model.

[0081] Figure 3 It is a schematic diagram according to the third embodiment of the present disclosure. Figure 3 The flowchart of this embodiment in training the three-dimensional object detection model is shown: Figure 3The upper part is the student network, and the lower part is the teacher network. When obtaining the final 3D object detection model based on the teacher network in the way of knowledge distillation, in addition to distilling at the feature level according to the 3D features and bird's-eye view features, output-level distillation is also performed according to the 3D detection box information output by the two networks, thereby improving the accuracy of the obtained 3D object detection model when performing 3D object detection.

[0082] Figure 4 It is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 4 shown, the 3D object detection method of this embodiment specifically includes the following steps:

[0083] S401. Obtain the binocular image to be detected;

[0084] S402. Input the binocular image to be detected into the 3D object detection model, and obtain the detection result of the binocular image to be detected according to the output result of the 3D object detection model.

[0085] That is to say, in this embodiment, the binocular image is detected by the pre-trained 3D object detection model, which can improve the accuracy of the obtained detection result.

[0086] When this embodiment executes S401 to obtain the binocular image to be detected, the binocular image captured in real time by the input end (such as an autonomous vehicle) can be used as the binocular image to be detected.

[0087] In the detection result obtained by this embodiment executing S402, in addition to the detection box information of the objects in the binocular image to be detected, it may further include the category information and category confidence score of the objects in the binocular image to be detected.

[0088] Figure 5 It is a schematic diagram according to the fifth embodiment of the present disclosure. As Figure 5 shown, the training device 500 of the 3D object detection model of this embodiment includes:

[0089] The first acquisition unit 501 is used to acquire a plurality of data pairs and the detection box annotation information of the plurality of data pairs, and each data pair includes the point cloud data and the binocular image corresponding to the same scene;

[0090] The second acquisition unit 502 is used to acquire the teacher network and the student network, where the teacher network is the detection network corresponding to the point cloud data, and the student network is the detection network corresponding to the binocular image;

[0091] The first processing unit 503 is configured to input the point cloud data in the data pair into the teacher network, input the binocular image into the student network, and obtain the first detection box information output by the teacher network and the second detection box information output by the student network;

[0092] The second processing unit 504 is configured to obtain the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair;

[0093] The calculation unit 505 is configured to obtain a first loss function value according to the target detection box information and the second detection box information, and obtain a second loss function value according to the second detection box information and the detection box annotation information;

[0094] The training unit 506 is configured to adjust the parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model.

[0095] Each data pair obtained by the first acquisition unit 501 includes point cloud data and a binocular image corresponding to the same scene; wherein, the detection box annotation information of the data pair includes the center point of the detection box, the size (length, width, and height) of the detection box, and the orientation angle of the detection box; in addition, the first acquisition unit 501 can also acquire the category annotation information of the objects included in the scene corresponding to each data pair.

[0096] In this embodiment, after the first acquisition unit 501 acquires a plurality of data pairs and the detection box annotation information of the plurality of data pairs, the second acquisition unit 502 acquires the teacher network and the student network.

[0097] The teacher network acquired by the second acquisition unit 502 is a detection network for the corresponding point cloud data, which can output the detection box information of the objects included in the input point cloud data; it can further output the category information of the objects and the corresponding category confidence scores.

[0098] The student network acquired by the second acquisition unit 502 is a detection network for the corresponding binocular image, which can output the detection box information of the objects included in the input binocular image; it can further output the category information of the objects and the corresponding category confidence scores.

[0099] In this embodiment, after the second acquisition unit 502 acquires the teacher network and the student network, the first processing unit 503 inputs the point cloud data in the data pair into the teacher network and inputs the binocular image in the data pair into the student network, and obtains the first detection box information output by the teacher network and the second detection box information output by the student network.

[0100] The first processing unit 503 inputs the point cloud data and the binocular image corresponding to the same data pair into the teacher network and the student network respectively, and then obtains the first detection box information output by the teacher network for the point cloud data and the second detection box information output by the student network for the binocular image.

[0101] The first processing unit 503 can further obtain the category information and category confidence score of the object.

[0102] In this embodiment, after the first processing unit 503 obtains the first detection box information and the second detection box information, the second processing unit 504 obtains the target detection box information according to the first detection box information, the second detection box information and the detection box annotation information of the data pair.

[0103] When the second processing unit 504 obtains the target detection box information according to the first detection box information, the second detection box information and the detection box annotation information of the data pair, the optional implementation method that can be adopted is: obtaining a detection box information group corresponding to the same object according to the first detection box information, the second detection box information and the detection box annotation information; obtaining candidate target detection box information according to the detection box information group; and obtaining the target detection box information according to the candidate target detection box information.

[0104] That is to say, the second processing unit 504 obtains the target detection box information according to the first detection box information, the second detection box information and the detection box annotation information, achieving the purpose of knowledge distillation at the output level of the network, so that when training the student network, the parameter adjustment can be carried out in combination with the loss function value obtained according to the target detection box information, thereby improving the training effect of the trained 3D object detection model.

[0105] Among them, when the second processing unit 504 obtains a detection box information group corresponding to the same object according to the first detection box information, the second detection box information and the detection box annotation information, the optional implementation method that can be adopted is: calculating the first intersection over union between the detection box annotation information and the first detection box information, and the second intersection over union between the detection box annotation information and the second detection box information; and in the case where it is determined that the first intersection over union is less than the first threshold and the second intersection over union is less than the second threshold, taking the first detection box information, the second detection box information and the detection box annotation information as the detection box information group corresponding to the same object.

[0106] When the second processing unit 504 obtains the candidate target detection frame information according to the detection frame information group, the following content may also be included: obtaining a first vector according to the center of the first detection frame in the first detection frame information and the center of the second detection frame in the second detection frame information; obtaining a second vector according to the center of the second detection frame in the second detection frame information and the center of the detection frame annotation in the detection frame annotation information; and when it is determined that the included angle between the first vector and the second vector is an acute angle, obtaining the candidate target detection frame information according to the first detection frame information, the second detection frame information, and the detection frame annotation information in the detection frame information group.

[0107] That is to say, before generating the target detection frame information, the second processing unit 504 can also determine whether the obtained detection frame information group is accurate according to the center point of the first detection frame, the center point of the second detection frame, and the center point of the annotated detection frame, so as to improve the accuracy of the obtained target detection frame information.

[0108] When the second processing unit 504 obtains the candidate target detection frame information according to the first detection frame information, the second detection frame information, and the detection frame annotation information in the detection frame information group, the optional implementation method that can be adopted is: obtaining the candidate target detection frame center point according to the center point of the first detection frame in the first detection frame information, the center point of the second detection frame in the second detection frame information, and the center point of the detection frame annotation in the detection frame annotation information; obtaining the candidate target detection frame size according to the size of the first detection frame in the first detection frame information, the size of the second detection frame in the second detection frame information, and the size of the detection frame annotation in the detection frame annotation information; obtaining the candidate target detection frame orientation angle according to the orientation angle of the first detection frame in the first detection frame information, the orientation angle of the second detection frame in the second detection frame information, and the orientation angle of the detection frame annotation in the detection frame annotation information; and obtaining the candidate target detection frame information according to the candidate target detection frame center point, the candidate target detection frame size, and the candidate target detection frame orientation angle.

[0109] When the second processing unit 504 obtains the target detection frame information according to the candidate target detection frame information, the optional implementation method that can be adopted is: when it is determined that the candidate detection frame center point in the candidate target detection frame information is greater than or equal to the first threshold, using the center point of the first detection frame in the first detection frame information as the target detection frame center point; otherwise, using the center point of the annotated detection frame as the target detection frame center point.

[0110] When the second processing unit 504 obtains the target detection frame information according to the obtained candidate target detection frame information, the optional implementation method that can be adopted is: when it is determined that the candidate detection frame size in the candidate target detection frame information is greater than or equal to the second threshold, using the size of the first detection frame in the first detection frame information as the target detection frame size; otherwise, using the size of the annotated detection frame as the target detection frame size.

[0111] When the second processing unit 504 obtains the target detection box information based on the obtained candidate target detection box information, an optional implementation method that can be adopted is as follows: when it is determined that the candidate detection box orientation angle in the candidate target detection box information is greater than or equal to the third threshold, the first detection box orientation angle in the first detection box information is used as the target detection box orientation angle; otherwise, the labeled detection box orientation angle is used as the target detection box orientation angle.

[0112] It can be understood that the first threshold, the second threshold, and the third threshold in this embodiment can be the same. For example, all three thresholds are 0; or three thresholds of different sizes can be set according to actual needs.

[0113] After the second processing unit 504 obtains the target detection box information in this embodiment, the calculation unit 505 obtains the first loss function value based on the target detection box information and the second detection box information, and obtains the second loss function value based on the second detection box information and the detection box annotation information.

[0114] When the calculation unit 505 obtains the loss function value, it can use the calculation method of the cross-entropy loss function to calculate the first loss function value and the second loss function value; it can be understood that in addition to using the detection box information of the object, the calculation unit 505 can also combine the category information of the object corresponding to the detection box to calculate the loss function.

[0115] After the calculation unit 505 obtains the first loss function value and the second loss function value in this embodiment, the training unit 506 adjusts the parameters of the student network according to the first loss function value and the second loss function value to obtain a three-dimensional object detection model.

[0116] When the training unit 506 adjusts the parameters of the student network according to the first loss function value and the second loss function value, it can stop adjusting the parameters of the student network when it is determined that the first loss function value and the second loss function value converge simultaneously, and use the student network when the loss function value converges as the three-dimensional object detection model.

[0117] When the training unit 506 adjusts the parameters of the student network according to the first loss function value and the second loss function value to obtain a three-dimensional object detection model, it can also include the following content: obtaining the first category confidence score corresponding to the target object according to the teacher network, and obtaining the second category confidence score corresponding to the target object according to the student network; obtaining the target weight according to the first category confidence score and the second category confidence score; obtaining the third loss function value according to the first target feature corresponding to the target object obtained by the teacher network, the second target feature corresponding to the target object obtained by the student network, and the target weight; adjusting the parameters of the student network according to the first loss function value, the second loss function value, and the third loss function to obtain a three-dimensional object detection model.

[0118] That is to say, when adjusting the parameters of the student network, the training unit 506 can, in addition to performing knowledge distillation at the output level, also combine knowledge distillation at the feature level, thereby improving the accuracy when adjusting the student network, and correspondingly improving the training effect of the trained 3D object detection model, enabling the 3D object detection model to perform 3D object detection more accurately.

[0119] It can be understood that the target object in the training unit 506 can be preset, for example, the object corresponding to the target category is used as the target object, and the target object can also be randomly selected; the number of target objects can be one or multiple.

[0120] The first target feature corresponding to the target object obtained by the training unit 506 according to the teacher network can be the 3D feature output by the 3D feature extraction layer in the teacher network, can also be the bird's-eye view feature output by the bird's-eye view feature extraction layer in the teacher network, and can also be the 3D feature and the bird's-eye view feature output by the teacher network.

[0121] The second target feature corresponding to the target object obtained by the training unit 506 according to the student network can be the 3D feature output by the 2D feature extraction layer in the student network, can also be the bird's-eye view feature output by the bird's-eye view feature extraction layer in the student network, and can also be the 3D feature and the bird's-eye view feature output by the student network.

[0122] That is to say, the training unit 506 combines the class confidence score and obtains the third loss function value according to the target feature, which can improve the accuracy of the third loss function value, and further improve the accuracy when adjusting the parameters of the student network.

[0123] When the training unit 506 adjusts the parameters of the student network according to the first loss function value, the second loss function value, and the third loss function value to obtain the 3D object detection model, it can stop adjusting the parameters of the student network when it is determined that the first loss function value, the second loss function value, and the third loss function value converge simultaneously, and use the student network when the loss function value converges as the 3D object detection model.

[0124] Figure 6 It is a schematic diagram according to the sixth embodiment of the present disclosure. As Figure 6 shown, the 3D object detection device 600 of this embodiment includes:

[0125] A third acquisition unit 601, configured to acquire a binocular image to be detected;

[0126] A detection unit 602, configured to input the binocular image to be detected into the 3D object detection model, and obtain the detection result of the binocular image to be detected according to the output result of the 3D object detection model.

[0127] When the third acquisition unit 601 acquires the binocular image to be detected, it may use the binocular image captured in real time by the input end (such as an autonomous vehicle) as the binocular image to be detected.

[0128] In the detection result obtained by the detection unit 602, in addition to the detection box information of the object in the binocular image to be detected, it may further include the category information and category confidence score of the object in the binocular image to be detected.

[0129] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0130] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0131] As Figure 7 shown, it is a block diagram of an electronic device for training a three-dimensional object detection model and a three-dimensional object detection method according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0132] As Figure 7 shown, the device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0133] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as a keyboard, a mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a disk, an optical disc, etc.; and communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0134] Computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 701 executes the various methods and processes described above, such as the training of a three-dimensional object detection model and the three-dimensional object detection method. For example, in some embodiments, the training of a three-dimensional object detection model and the three-dimensional object detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708.

[0135] In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the training of a three-dimensional object detection model and the three-dimensional object detection method described above can be executed. Alternatively, in other embodiments, computing unit 701 can be configured to execute the training of a three-dimensional object detection model and the three-dimensional object detection method by any other suitable means (e.g., by means of firmware).

[0136] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a general-purpose computer, a special-purpose computer, or a processor or controller of other programmable three-dimensional object detection model training and three-dimensional object detection devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0138] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for presenting information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0140] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0141] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on respective computers and having a client - server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (“Virtual Private Server”, or simply “VPS”). The server can also be a server of a distributed system, or a server combined with blockchain.

[0142] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0143] The above - described specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A training method for a three-dimensional object detection model, comprising: Obtaining a plurality of data pairs and detection box annotation information of the plurality of data pairs, each data pair including point cloud data and binocular images corresponding to the same scene; Obtaining a teacher network and a student network, where the teacher network is a detection network for the corresponding point cloud data, and the student network is a detection network for the corresponding binocular images; Inputting the point cloud data in the data pair into the teacher network and inputting the binocular images into the student network to obtain first detection box information output by the teacher network and second detection box information output by the student network; Obtaining target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair; Obtaining a first loss function value according to the target detection box information and the second detection box information, and obtaining a second loss function value according to the second detection box information and the detection box annotation information; Adjusting the parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model; Wherein, the obtaining of the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair includes: Obtaining a detection box information group corresponding to the same object according to the first detection box information, the second detection box information, and the detection box annotation information; Obtaining candidate target detection box information according to the detection box information group; Obtaining the target detection box information according to the candidate target detection box information; The obtaining of the candidate target detection box information according to the detection box information group includes: Obtaining a first vector according to a first detection box center in the first detection box information and a second detection box center in the second detection box information; Obtaining a second vector according to the second detection box center in the second detection box information and a detection box annotation center in the detection box annotation information; When it is determined that the included angle between the first vector and the second vector is an acute angle, obtaining the candidate target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information in the detection box information group.

2. The method according to claim 1, wherein the obtaining of the detection box information group corresponding to the same object according to the first detection box information, the second detection box information, and the detection box annotation information includes: Calculating a first intersection over union between the detection box annotation information and the first detection box information, and a second intersection over union between the detection box annotation information and the second detection box information; When it is determined that the first intersection over union is less than a first threshold and the second intersection over union is less than a second threshold, taking the first detection box information, the second detection box information, and the detection box annotation information as the detection box information group corresponding to the same object.

3. The method according to claim 1, wherein The obtaining of the candidate target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information in the detection box information group includes: Obtain the candidate target detection frame center point based on the first detection frame center point in the first detection frame information, the second detection frame center point in the second detection frame information, and the detection frame annotation center point in the detection frame annotation information; Obtain the candidate target detection frame size based on the first detection frame size in the first detection frame information, the second detection frame size in the second detection frame information, and the detection frame annotation size in the detection frame annotation information; Obtain the candidate target detection frame orientation angle based on the first detection frame orientation angle in the first detection frame information, the second detection frame orientation angle in the second detection frame information, and the detection frame annotation orientation angle in the detection frame annotation information; Obtain the candidate target detection frame information based on the candidate target detection frame center point, the candidate target detection frame size, and the candidate target detection frame orientation angle; 4. The method according to any one of claims 1 to 3, wherein, The obtaining the target detection frame information according to the candidate target detection frame information includes: When it is determined that the candidate detection frame center point in the candidate target detection frame information is greater than or equal to the first threshold, use the first detection frame center point in the first detection frame information as the target detection frame center point in the target detection frame information; Otherwise, use the annotation detection frame center point as the target detection frame center point in the target detection frame information.

5. The method according to any one of claims 1-4, wherein The obtaining the target detection frame information according to the candidate target detection frame information includes: When it is determined that the candidate detection frame size in the candidate target detection frame information is greater than or equal to the second threshold, use the first detection frame size in the first detection frame information as the target detection frame size in the target detection frame information; Otherwise, use the annotation detection frame size as the target detection frame size in the target detection frame information.

6. The method according to any one of claims 1-5, wherein The obtaining the target detection frame information according to the candidate target detection frame information includes: When it is determined that the candidate detection frame orientation angle in the candidate target detection frame information is greater than or equal to the third threshold, use the first detection frame orientation angle in the first detection frame information as the target detection frame orientation angle in the target detection frame information; Otherwise, use the annotation detection frame orientation angle as the target detection frame orientation angle in the target detection frame information.

7. The method according to any one of claims 1-6, wherein, The obtaining the three-dimensional object detection model by adjusting the parameters of the student network according to the first loss function value and the second loss function value includes: Obtain the first category confidence score corresponding to the target object according to the teacher network, and obtain the second category confidence score corresponding to the target object according to the student network; Obtain the target weight according to the first category confidence score and the second category confidence score; Obtain the third loss function value according to the first target feature corresponding to the target object obtained by the teacher network, the second target feature corresponding to the target object obtained by the student network, and the target weight; Adjust the parameters of the student network according to the first loss function value, the second loss function value, and the third loss function to obtain the three-dimensional object detection model.

8. A three-dimensional object detection method, including: Obtain the binocular image to be detected; Input the binocular image to be detected into a 3D object detection model, and obtain the detection result of the binocular image to be detected according to the output result of the 3D object detection model; Among them, the 3D object detection model is trained according to the method described in any one of claims 1-7.

9. A training device for a 3D object detection model, comprising: A first acquisition unit, configured to acquire a plurality of data pairs and the detection box annotation information of the plurality of data pairs, and each data pair includes point cloud data and a binocular image corresponding to the same scene; A second acquisition unit, configured to acquire a teacher network and a student network, where the teacher network is a detection network for the corresponding point cloud data, and the student network is a detection network for the corresponding binocular image; A first processing unit, configured to input the point cloud data in the data pair into the teacher network and the binocular image into the student network, and obtain the first detection box information output by the teacher network and the second detection box information output by the student network; A second processing unit, configured to obtain target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair; A calculation unit, configured to obtain a first loss function value according to the target detection box information and the second detection box information, and obtain a second loss function value according to the second detection box information and the detection box annotation information; A training unit, configured to adjust the parameters of the student network according to the first loss function value and the second loss function value to obtain the 3D object detection model; Among them, when the second processing unit obtains the target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information of the data pair, it specifically performs: Obtain a detection box information group corresponding to the same object according to the first detection box information, the second detection box information, and the detection box annotation information; Obtain candidate target detection box information according to the detection box information group; Obtain the target detection box information according to the candidate target detection box information; When the second processing unit obtains the candidate target detection box information according to the detection box information group, it specifically performs: Obtain a first vector according to the first detection box center in the first detection box information and the second detection box center in the second detection box information; Obtain a second vector according to the second detection box center in the second detection box information and the detection box annotation center in the detection box annotation information; In the case where it is determined that the included angle between the first vector and the second vector is an acute angle, obtain the candidate target detection box information according to the first detection box information, the second detection box information, and the detection box annotation information in the detection box information group.

10. The device according to claim 9, when the second processing unit obtains a detection box information group corresponding to the same object according to the first detection box information, the second detection box information, and the detection box annotation information, it specifically performs: Calculate the first intersection over union (IoU) between the detected bounding box annotation information and the first detected bounding box information, and the second IoU between the detected bounding box annotation information and the second detected bounding box information; When it is determined that the first IoU is less than the first threshold and the second IoU is less than the second threshold, use the first detected bounding box information, the second detected bounding box information, and the detected bounding box annotation information as the detected bounding box information group corresponding to the same object.

11. The device according to claim 9, wherein When obtaining the candidate target detected bounding box information according to the first detected bounding box information, the second detected bounding box information, and the detected bounding box annotation information in the detected bounding box information group, the second processing unit specifically performs: Obtain the candidate target detected bounding box center point according to the first detected bounding box center point in the first detected bounding box information, the second detected bounding box center point in the second detected bounding box information, and the detected bounding box annotation center point in the detected bounding box annotation information; Obtain the candidate target detected bounding box size according to the first detected bounding box size in the first detected bounding box information, the second detected bounding box size in the second detected bounding box information, and the detected bounding box annotation size in the detected bounding box annotation information; Obtain the candidate target detected bounding box orientation angle according to the first detected bounding box orientation angle in the first detected bounding box information, the second detected bounding box orientation angle in the second detected bounding box information, and the detected bounding box annotation orientation angle in the detected bounding box annotation information; Obtain the candidate target detected bounding box information according to the candidate target detected bounding box center point, the candidate target detected bounding box size, and the candidate target detected bounding box orientation angle.

12. The device according to any one of claims 9-11, wherein, When obtaining the target detected bounding box information according to the candidate target detected bounding box information, the second processing unit specifically performs: When it is determined that the candidate detected bounding box center point in the candidate target detected bounding box information is greater than or equal to the first threshold, use the first detected bounding box center point in the first detected bounding box information as the target detected bounding box center point in the target detected bounding box information; Otherwise, use the annotated detected bounding box center point as the target detected bounding box center point in the target detected bounding box information.

13. The device according to any one of claims 9 - 12, wherein, When obtaining the target detected bounding box information according to the candidate target detected bounding box information, the second processing unit specifically performs: When it is determined that the candidate detected bounding box size in the candidate target detected bounding box information is greater than or equal to the second threshold, use the first detected bounding box size in the first detected bounding box information as the target detected bounding box size in the target detected bounding box information; Otherwise, use the annotated detected bounding box size as the target detected bounding box size in the target detected bounding box information.

14. The apparatus according to any one of claims 9 - 13, wherein, When obtaining the target detected bounding box information according to the candidate target detected bounding box information, the second processing unit specifically performs: When it is determined that the candidate detected bounding box orientation angle in the candidate target detected bounding box information is greater than or equal to the third threshold, use the first detected bounding box orientation angle in the first detected bounding box information as the target detected bounding box orientation angle in the target detected bounding box information; Otherwise, use the annotated detected bounding box orientation angle as the target detected bounding box orientation angle in the target detected bounding box information.

15. The device according to any one of claims 9 - 14, wherein, When adjusting the parameters of the student network according to the first loss function value and the second loss function value to obtain the three-dimensional object detection model, the training unit specifically performs the following: Obtaining a first class confidence score corresponding to the target object according to the teacher network, and obtaining a second class confidence score corresponding to the target object according to the student network; Obtaining a target weight according to the first class confidence score and the second class confidence score; Obtaining a third loss function value according to the first target feature corresponding to the target object obtained by the teacher network, the second target feature corresponding to the target object obtained by the student network, and the target weight; Adjusting the parameters of the student network according to the first loss function value, the second loss function value and the third loss function to obtain the three-dimensional object detection model.

16. A three-dimensional object detection device, comprising: A third acquisition unit, configured to acquire a binocular image to be detected; A detection unit, configured to input the binocular image to be detected into the three-dimensional object detection model, and obtain a detection result of the binocular image to be detected according to an output result of the three-dimensional object detection model; Wherein, the three-dimensional object detection model is trained according to the device described in any one of claims 9-15.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-8.

19. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Model training method and device, equipment, storage medium and image detection method

    CN113920307A

  • Three-dimensional vehicle detection method, system and device and medium

    CN114359891A