Target detection method, target detection device, equipment and computer storage medium
By combining monocular 2D detection and monocular 3D detection, and using two-dimensional truth value to train a three-dimensional detection model, the problems of unstable ranging and missed detection and missed detection in the existing technology are solved, and the reliability of target detection results is improved.
Patent Information
- Application Number
- CN202510281475.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing intelligent assisted driving method based on visual images lacks 3D information, resulting in unstable distance measurement. The BEV-based deep learning method is limited by the short perceived distance, and the monocular 3D method has missed detection and misdetection problems.
Combining monocular 2D detection and monocular 3D detection, the target two-dimensional information is obtained through the two-dimensional detection model, and the target three-dimensional information is obtained through the three-dimensional detection model, and the three-dimensional truth value is used to train the three-dimensional detection model to improve the reliability of target detection.
While ensuring recall and accuracy, the three-dimensional information of the target is predicted, improving the reliability of the target detection results.
Smart Images

Figure CN120339574A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of object detection technology, and particularly to an object detection method, an object detection device, an object detection equipment, and a computer storage medium. Background Art
[0002] The perception based on visual images lacks 3D information and cannot be directly used for intelligent assisted driving functions.
[0003] Currently, the mainstream methods are divided into three points. First, based on the traditional convolutional network, the pseudo 3D information of the obstacle is predicted, including the 2D bounding box of the image, the front or rear frame of the vehicle, the direction line of the vehicle, and the sub-category attributes of the vehicle (sedan, truck, bus). Then, based on the plane hypothesis, using the image ground contact point (the bottom of the 2D bounding box) or the width and height information of the obstacle, the distance measurement is carried out according to the principle of similar triangles. This method relies on strong prior assumptions and the distance measurement is unstable.
[0004] Second, the deep learning method based on BEV (Bird's Eye View) directly predicts the position, size, and heading angle information of the obstacle in the camera coordinate system. This method is limited by the short perception distance, and at the same time, the computing power of the in-vehicle chip and the support degree of the operator are low, and the application range is severely restricted.
[0005] Third, the method based on monocular 3D is a compromise between the first technical route and the second route. It completes the prediction of 3D information based on the basic convolutional operator and has good portability. However, it is limited by the quality of the ground truth. The ground truth is generally obtained from the point cloud obtained by the lidar sensor and obtained in the vision algorithm with the point cloud as the input. There are inevitably problems of missing and misdetection, so the algorithm trained based on this ground truth inevitably has problems of missing and misdetection. Summary of the Invention
[0006] To solve the above technical problems, this application proposes an object detection method, an object detection device, an object detection equipment, and a computer storage medium.
[0007] To solve the above technical problems, this application proposes an object detection method, and the object detection method includes:
[0008] Obtain the image to be detected;
[0009] Input the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; wherein, the two-dimensional detection information includes target two-dimensional information;
[0010] Input the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information;
[0011] Based on the position information of the target two-dimensional information in the two-dimensional detection information, obtain the target three-dimensional information of the image to be detected in the three-dimensional detection information;
[0012] Based on the target two-dimensional information and the target three-dimensional information, obtain the target detection result in the image to be detected.
[0013] Among them, the position information of the target two-dimensional information in the two-dimensional detection information, that is, the center point position in the target two-dimensional information.
[0014] Among them, the two-dimensional detection model includes a two-dimensional feature extraction network and a two-dimensional detection head;
[0015] The step of inputting the image to be detected into the two-dimensional detection model to obtain two-dimensional detection information includes:
[0016] Input the image to be detected into the two-dimensional feature extraction network to extract two-dimensional features;
[0017] Input the two-dimensional features into the two-dimensional detection head to obtain the two-dimensional detection information.
[0018] Among them, the three-dimensional detection model includes a three-dimensional feature extraction network and a three-dimensional detection head;
[0019] The step of inputting the image to be detected into the three-dimensional detection model to obtain three-dimensional detection information includes:
[0020] Input the image to be detected into the three-dimensional feature extraction network to extract three-dimensional features;
[0021] Concatenate the three-dimensional features and the two-dimensional features to obtain concatenated features;
[0022] Input the concatenated features into the three-dimensional detection head to obtain the three-dimensional detection information.
[0023] Among them, the target detection method further includes:
[0024] Obtain the two-dimensional ground truth in the first image to be trained;
[0025] Use the two-dimensional ground truth to train the two-dimensional detection model;
[0026] After the two-dimensional detection model is trained, obtain the two-dimensional prediction value output by the two-dimensional detection model for detecting the second image to be trained;
[0027] Use the two-dimensional prediction value and the three-dimensional ground truth of the second point cloud to be trained to train the three-dimensional detection model;
[0028] Among them, the second image to be trained and the second point cloud to be trained are collected simultaneously for the same region.
[0029] Among them, training the 3D detection model by using the 2D prediction value and the 3D ground truth of the second image to be trained includes:
[0030] Matching the 2D prediction value and the projection of the 3D ground truth on the second image to be trained to obtain a matching 2D prediction value;
[0031] Based on the matching 2D prediction value, obtaining the 3D prediction value at the position corresponding to the 2D prediction value in the output layer of the 3D detection head;
[0032] Calculating a loss value by using the 3D prediction value and the 3D ground truth;
[0033] Training the 3D detection model by using the loss value.
[0034] Among them, matching the 2D prediction value and the projection of the 3D ground truth on the second image to be trained to obtain a matching 2D prediction value includes:
[0035] Obtaining the projected ground truth box of the 3D ground truth on the second image to be trained;
[0036] Obtaining the intersection over union (IoU) between the projected ground truth box and a number of prediction boxes of the 2D prediction value;
[0037] Obtaining the 2D prediction value corresponding to the prediction box with the largest IoU as the matching 2D prediction value.
[0038] To solve the above technical problems, the present application also proposes an object detection device, which includes: an image acquisition module, a 2D detection module, a 3D detection module, and an object detection module; among them,
[0039] The image acquisition module is used to acquire an image to be detected;
[0040] The 2D detection module is used to input the image to be detected into a 2D detection model to obtain 2D detection information; among them, the 2D detection information includes target 2D information;
[0041] The 3D detection module is used to input the image to be detected into a 3D detection model to obtain 3D detection information;
[0042] The 3D detection module is used to obtain the target 3D information of the image to be detected in the 3D detection information based on the position information of the target 2D information in the 2D detection information;
[0043] The object detection module is used to obtain the object detection result in the image to be detected based on the target 2D information and the target 3D information.
[0044] To solve the above technical problems, the present application also proposes an object detection device, which includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the object detection method as described above.
[0045] To solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the above object detection method.
[0046] Compared with the prior art, the beneficial effect of the present application is that the object detection device acquires an image to be detected; inputs the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; wherein, the two-dimensional detection information includes object two-dimensional information; inputs the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information; based on the position information of the object two-dimensional information in the two-dimensional detection information, obtains the object three-dimensional information of the image to be detected in the three-dimensional detection information; and based on the object two-dimensional information and the object three-dimensional information, obtains the object detection result in the image to be detected. Through the above object detection method, the three-dimensional detection is coupled with the two-dimensional detection, and while ensuring recall and accuracy, the three-dimensional information of the object is predicted, improving the reliability of the object detection result. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0048] Among them:
[0049] Figure 1 is a schematic flowchart of the first embodiment of the object detection method provided by the present application;
[0050] Figure 2 is a schematic flowchart of the object detection method provided by the present application in the prediction stage;
[0051] Figure 3 is a schematic diagram of an embodiment of the object two-dimensional information provided by the present application;
[0052] Figure 4 is a schematic diagram of the projection of the object three-dimensional information provided by the present application on a two-dimensional image;
[0053] Figure 5 is a schematic flowchart of the second embodiment of the object detection method provided by the present application;
[0054] Figure 6 It is a schematic flowchart of the object detection method provided by this application in the training stage;
[0055] Figure 7 It is a schematic flowchart of the third embodiment of the object detection method provided by this application;
[0056] Figure 8 It is a schematic structural diagram of an embodiment of the object detection device provided by this application;
[0057] Figure 9 It is a schematic structural diagram of an embodiment of the object detection device provided by this application;
[0058] Figure 10 It is a schematic structural diagram of an embodiment of the computer storage medium provided by this application. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0060] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0061] In autonomous driving technology, in order to implement the vehicle's assisted driving functions, such as adaptive cruise control, highway pilot, urban pilot, urban commuting assistance, etc., it is necessary to identify dynamic and static obstacles in the traffic scene, measure information such as the position, size, heading angle, speed, etc. of the dynamic and static obstacles, and after obtaining this information, output it to the downstream planning and control system, and the control system controls the vehicle based on this information.
[0062] Currently, monocular cameras, binocular cameras, and multi-line lidar are mainly used to identify dynamic and static obstacles in traffic scenes. The task of 3D object detection using a monocular camera can be completed by a monocular 3D object detection model. Since the cost of a monocular camera is low, the economic effect is better. The specific approach is to input a single RGB image into the monocular 3D object detection model, which can predict the object type and its 3D position information in the image. The 3D position information includes the height h, width w, length l of the object, the position coordinates (x, y, z) of the center point of the object, and the heading angle theta.
[0063] Because the monocular 3D object detection scheme is supervised learning based on ground truth, and the current ground truth is obtained based on the lidar point cloud object detection algorithm. The ground truth system algorithm based on point cloud inevitably has the defects of missed detection and false detection. Then, the monocular 3D object detection algorithm directly based on this ground truth will have the problems of missed detection and false detection. However, the 3D information of the ground truth obtained by the ground truth system is relatively accurate and reliable. For the algorithm based on 2D detection, its supervision signal is based on manual annotation. The accuracy of manual annotation is approximate to the real physical world and does not have the defects of the ground truth system. However, the annotated ground truth based on images does not have the three-dimensional information such as length, width, height, heading angle, and depth of the physical world.
[0064] To solve or partially solve the problems existing in the related technologies, the present application provides an object detection method, device, and equipment for autonomous driving, which couples monocular 3D detection and monocular 2D detection, and predicts the three-dimensional information of the object while ensuring recall and accuracy, thereby improving the reliability of the object detection result.
[0065] For details, please refer to Figure 1 and Figure 2 , Figure 1 which are the schematic flowcharts of the first embodiment of the object detection method provided by the present application, Figure 2 and which is the schematic flowchart of the object detection method provided by the present application in the prediction stage.
[0066] The object detection method of the present application is applied to an object detection device. Among them, the object detection device of the present application can be a server, or a terminal device, or a system in which the server and the terminal device cooperate with each other. Correspondingly, each part included in the object detection device, such as each unit, subunit, module, and submodule, can be all set in the server, or all set in the terminal device, or can be respectively set in the server and the terminal device.
[0067] Further, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module, which is not specifically limited herein.
[0068] As Figure 1 shown, the specific steps are as follows:
[0069] Step S11: Obtain the image to be detected.
[0070] In the embodiment of the present application, the target detection device acquires the image to be detected through a monocular vehicle-mounted camera. Among them, the image to be detected can be a real-time image acquired during the vehicle driving process, or an image acquired when the vehicle is parked or when the vehicle activates the sentry mode, etc.
[0071] Step S12: Input the image to be detected into the two-dimensional detection model to obtain two-dimensional detection information; among them, the two-dimensional detection information includes target two-dimensional information.
[0072] In the embodiment of the present application, as Figure 2 shown, the target detection device inputs the image to be detected into the two-dimensional detection model and the three-dimensional detection model respectively for target detection.
[0073] Among them, the two-dimensional detection model includes a 2D feature extraction network and a 2D detection head as Figure 2 shown. The detection process of the two-dimensional detection model is as follows:
[0074] The target detection device inputs the image to be detected into the 2D feature extraction network to extract the two-dimensional features of the image to be detected, and then inputs the two-dimensional features into the 2D detection head to predict the 2D information of the image to be detected.
[0075] Among them, the detection models provided in the present application are all target detection models that generate prediction boxes. Therefore, the 2D information of the image to be detected output by the 2D detection model is specifically 2D prediction box information, that is, it includes attribute information such as the center point position of the 2D prediction box, the prediction box size, and the target prediction category.
[0076] It should be noted that all neural networks in the two-dimensional detection model and the three-dimensional detection model provided in the present application are composed of conventional network layers such as convolutional layers, normalization layers, and activation layers.
[0077] In a specific implementation manner, the 2D detection head can be the FCOS (Fully Convolutional One-Stage Object Detection) target detection algorithm, and the 3D detection head can be the FCOS3D detection model.
[0078] The feature extraction modules of the two branches are residual structure networks that do not share with each other.
[0079] Step S13: Input the image to be detected into the 3D detection model to obtain 3D detection information.
[0080] In the embodiment of the present application, the 3D detection model includes a 3D feature extraction network and a 3D detection head as shown in Figure 2 . The detection process of the 3D detection model is as follows:
[0081] The target detection device inputs the image to be detected into the 3D feature extraction network to extract the 3D features of the image to be detected, and then inputs the 2D features into the 3D detection head to predict the 3D information of the image to be detected.
[0082] Among them, the detection models provided in the present application are all object detection models that generate prediction boxes. Therefore, the 3D information of the image to be detected output by the 3D detection model is specifically 3D prediction box information, that is, it includes the center point position of the 3D prediction box, the prediction box size and other information. Compared with the 2D information in step S12, the prediction box size of the 3D information is composed of information in three dimensions of length, width and height. In addition, the 3D information also includes the predicted heading angle corresponding to the target. Among them, the two-dimensional size in the 3D prediction box size is the same as the 2D prediction box size.
[0083] Furthermore, as shown in Figure 2 , in order to improve the prediction accuracy of the 3D detection head, the target detection device can also splice the 2D features and the 3D features to obtain spliced features, and then input the spliced features into the 3D detection head for prediction to obtain 3D information. By fusing the 2D features and the 3D features, more abundant feature information of the image to be detected can be introduced, thereby improving the prediction accuracy of the 3D information.
[0084] Step S14: Based on the position information of the target two-dimensional information in the two-dimensional detection information, obtain the target three-dimensional information of the image to be detected in the three-dimensional detection information.
[0085] In the embodiment of the present application, the target detection device correlates the target two-dimensional information and the three-dimensional information of the target category to extract the associated target two-dimensional information and target three-dimensional information.
[0086] Among them, according to the two-dimensional information and three-dimensional information of the same size, the above correlation method can be: obtain the position information of the target two-dimensional information, and determine the target three-dimensional information of the same position information according to this position information.
[0087] The position information can be the center point position or the corner point position of the prediction box. The following takes the center point position as an example for illustration: Please refer specifically to Figure 3 and Figure 4 ,Figure 3 It is a schematic diagram of an embodiment of the target two-dimensional information provided by this application. Figure 4 It is a schematic diagram of the projection of the target three-dimensional information provided by this application on a two-dimensional image.
[0088] As Figure 3 shown, the target detection device determines the center point position of the target two-dimensional prediction box as the index of the two-dimensional prediction box according to the target two-dimensional information, and then obtains the target three-dimensional information at the same position according to this index. Among them, the target three-dimensional information is specifically the information of the three-dimensional prediction box with the same center point position, including but not limited to: the length, width and height of the three-dimensional prediction target in the physical world, the three-dimensional coordinates of the center point, and the three-dimensional information of the heading angle, etc.
[0089] Furthermore, since the number of two-dimensional prediction boxes output by the two-dimensional detection model is large and the overlapping situation of the two-dimensional prediction boxes is many, therefore, before obtaining the target three-dimensional information, the target detection device can first perform a non-maximum suppression operation on the two-dimensional information output by the two-dimensional detection model, so as to screen the two-dimensional prediction boxes and reduce the target two-dimensional information.
[0090] Step S15: Based on the target two-dimensional information and the target three-dimensional information, obtain the target detection result in the image to be detected.
[0091] In the embodiment of this application, the target detection device outputs the target two-dimensional information and the target three-dimensional information of the target on the image to be detected, thereby forming the target detection result, including but not limited to: the 2D box on the image, other attributes such as the category and sub-category of the object, the length, width and height in the physical world, the three-dimensional coordinates of the center point, the three-dimensional information of the heading angle, etc.
[0092] In this application, the target detection device obtains the image to be detected; inputs the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; wherein, the two-dimensional detection information includes target two-dimensional information; inputs the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information; based on the position information of the target two-dimensional information in the two-dimensional detection information, obtain the target three-dimensional information of the image to be detected in the three-dimensional detection information; based on the target two-dimensional information and the target three-dimensional information, obtain the target detection result in the image to be detected. Through the above target detection method, the three-dimensional detection and the two-dimensional detection are coupled, and the three-dimensional information of the target is predicted while ensuring the recall and accuracy, improving the reliability of the target detection result.
[0093] The data required for this application includes monocular images obtained by an in-vehicle camera, the internal parameters of the camera, the joint calibration parameters of the lidar and the camera, 2D ground truth based on in-vehicle images, 3D ground truth based on lidar, and the joint calibration parameters of the lidar and the camera are used to transfer the ground truth based on lidar to the camera coordinate system. The internal parameters are used to decode 3D information and for the conversion from the camera coordinate system to the pixel coordinate system. The 2D ground truth is manually labeled and used to train the 2D network. The 3D ground truth is obtained by a ground truth algorithm based on point cloud and is used to train the 3D network.
[0094] Please continue to refer to 5 and Figure 6 , Figure 5 FIG. is a schematic flowchart of the second embodiment of the object detection method provided by this application, Figure 6 FIG. is a schematic flowchart of the object detection method provided by this application in the training phase.
[0095] As Figure 5 shown, the specific steps are as follows:
[0096] Step S21: Obtain the 2D ground truth in the first image to be trained.
[0097] In the embodiment of this application, the object detection device extracts the 2D ground truth of the image to be trained. Among them, the 2D ground truth is specifically the annotation box of the object and is pre-annotated by the staff.
[0098] Step S22: Train the 2D detection model using the 2D ground truth.
[0099] In the embodiment of this application, the object detection device trains the 2D detection model using the 2D ground truth of the image to be trained. The process is basically the same as the training process of the 2D object detection model in the prior art and will not be elaborated here.
[0100] Step S23: After the 2D detection model is trained, obtain the 2D prediction value output by the 2D detection model for detecting the second image to be trained.
[0101] In the embodiment of this application, after the training of the 2D detection model in step S22 is completed, the object detection device continues to input another image to be trained into the trained 2D detection model for detection to obtain the 2D prediction value.
[0102] Step S24: Train the 3D detection model using the 2D prediction value and the 3D ground truth of the second point cloud to be trained.
[0103] In the embodiment of this application, the object detection device simultaneously collects the image to be trained and the point cloud to be trained in the same area. Among them, the image to be trained is input into the 2D detection model in step S23 to obtain the 2D prediction value, and the point cloud to be trained is used to jointly train the 3D detection model with the 2D prediction value.
[0104] Specifically, for the process of jointly training the 3D detection model with the 2D prediction values and the point cloud to be trained, please continue to refer to Figure 7 , Figure 7 which is a schematic flowchart of the third embodiment of the object detection method provided by this application.
[0105] As Figure 7 shown, the specific steps are as follows:
[0106] Step S31: Match the 2D prediction values with the projection of the 3D ground truth on the second image to be trained, and obtain the matching 2D prediction values.
[0107] In the embodiments of this application, the object detection device uses the final prediction result of the 2D detection model to match with the projection of the 3D ground truth on the 2D image, obtains the index corresponding to the 2D box, and uses the index to obtain the 3D information prediction value.
[0108] Specifically, the matching of the projection of the 3D ground truth on the 2D image uses the intersection over union (IoU) of two rectangular boxes. When the IoU of the two is greater than A, it is considered a successful match. Otherwise, it is a failure. For a 2D prediction value with multiple successfully matched 3D ground truths, the one with the largest IoU is taken as the matching object. For objects with the same IoU, the one with the closest Euclidean distance of the center points is taken as the matching object, that is, the matching 2D prediction value.
[0109] Step S32: Based on the matching 2D prediction values, obtain the 3D prediction values.
[0110] In the embodiments of this application, the object detection device remaps the matching 2D prediction values to the 3D image, that is, the 3D point cloud, to obtain the 3D prediction values.
[0111] Step S33: Use the 3D prediction values and the 3D ground truth to calculate the loss value.
[0112] In the embodiments of this application, the object detection device calculates the loss value according to the 3D prediction values and the 3D ground truth for training the 3D detection model.
[0113] Step S34: Use the loss value to train the 3D detection model.
[0114] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0115] To implement the above object detection method, this application also proposes an object detection device. Specifically, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an embodiment of the object detection device provided by this application.
[0116] The target detection device 500 in this embodiment includes: an image acquisition module 51, a two-dimensional detection module 52, a three-dimensional detection module 53, and a target detection module 54.
[0117] Among them, the image acquisition module 51 is used to acquire the image to be detected.
[0118] The two-dimensional detection module 52 is used to input the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; among them, the two-dimensional detection information includes target two-dimensional information.
[0119] The three-dimensional detection module 53 is used to input the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information.
[0120] The three-dimensional detection module 53 is used to obtain the target three-dimensional information of the image to be detected in the three-dimensional detection information based on the position information of the target two-dimensional information in the two-dimensional detection information.
[0121] The target detection module 54 is used to obtain the target detection result in the image to be detected based on the target two-dimensional information and the target three-dimensional information.
[0122] To implement the above target detection method, the present application also proposes a target detection device. For details, please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an embodiment of the target detection device provided by the present application.
[0123] The target detection device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.
[0124] The processor 41, the memory 42, and the input / output device 43 are respectively connected to the bus 44. Program data is stored in the memory 42, and the processor 41 is used to execute the program data to implement the target detection method described in the above embodiment.
[0125] In an embodiment of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.
[0126] The present application also provides a computer storage medium. Please continue to refer to Figure 10 , Figure 10 FIG. is a schematic structural diagram of an embodiment of the computer storage medium provided by the present application. The computer storage medium 600 stores a computer program 61, and when the computer program 61 is executed by a processor, it is used to implement the object detection method of the above embodiment.
[0127] When the embodiments of the present application are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0128] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A target detection method, characterized in that, The target detection method includes: Obtain the image to be detected; Input the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; wherein, the two-dimensional detection information includes target two-dimensional information; Input the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information; Based on the position information of the target two-dimensional information in the two-dimensional detection information, obtain the target three-dimensional information of the image to be detected in the three-dimensional detection information; Based on the target two-dimensional information and the target three-dimensional information, obtain the target detection result in the image to be detected.
2. The target detection method according to claim 1, wherein The position information of the target two-dimensional information in the two-dimensional detection information, that is, the center point position in the target two-dimensional information.
3. The target detection method according to claim 1, wherein The two-dimensional detection model includes a two-dimensional feature extraction network and a two-dimensional detection head; The step of inputting the image to be detected into the two-dimensional detection model to obtain two-dimensional detection information includes: Input the image to be detected into the two-dimensional feature extraction network to extract two-dimensional features; Input the two-dimensional features into the two-dimensional detection head to obtain the two-dimensional detection information.
4. The target detection method according to claim 3, wherein The three-dimensional detection model includes a three-dimensional feature extraction network and a three-dimensional detection head; The step of inputting the image to be detected into the three-dimensional detection model to obtain three-dimensional detection information includes: Input the image to be detected into the three-dimensional feature extraction network to extract three-dimensional features; Concatenate the three-dimensional features with the two-dimensional features to obtain concatenated features; Input the concatenated features into the three-dimensional detection head to obtain the three-dimensional detection information.
5. The target detection method according to claim 1, wherein The target detection method further includes: Obtain the two-dimensional ground truth in the first image to be trained; Train the two-dimensional detection model using the two-dimensional ground truth; After the two-dimensional detection model is trained, obtain the two-dimensional prediction value output by the two-dimensional detection model for detecting the second image to be trained; Train the three-dimensional detection model using the two-dimensional prediction value and the three-dimensional ground truth of the second point cloud to be trained; Wherein, the second image to be trained and the second point cloud to be trained are collected simultaneously for the same area.
6. The target detection method according to claim 5, wherein The step of training the three-dimensional detection model using the two-dimensional prediction value and the three-dimensional ground truth of the second image to be trained includes: Match the two-dimensional prediction value with the projection of the three-dimensional ground truth in the second image to be trained to obtain a matched two-dimensional prediction value; Based on the matched two-dimensional prediction value, obtain the three-dimensional prediction value at the position corresponding to the two-dimensional prediction value in the output layer of the three-dimensional detection head; Calculate the loss value using the three-dimensional prediction value and the three-dimensional ground truth; Train the three-dimensional detection model using the loss value.
7. The target detection method according to claim 6, wherein The step of matching the two-dimensional prediction value with the projection of the three-dimensional ground truth in the second image to be trained to obtain a matched two-dimensional prediction value includes: Obtain the projected ground truth box of the three-dimensional ground truth in the second image to be trained; Obtain the intersection over union (IoU) between the projected ground truth box and a number of predicted boxes of the two-dimensional prediction value; Obtain the two-dimensional prediction value corresponding to the predicted box with the largest IoU as the matching two-dimensional prediction value.
8. A target detection device, characterized in that, The object detection device includes: an image acquisition module, a two-dimensional detection module, a three-dimensional detection module, and an object detection module; wherein, The image acquisition module is used to acquire the image to be detected; The two-dimensional detection module is used to input the image to be detected into a two-dimensional detection model to obtain two-dimensional detection information; wherein, the two-dimensional detection information includes object two-dimensional information; The three-dimensional detection module is used to input the image to be detected into a three-dimensional detection model to obtain three-dimensional detection information; The three-dimensional detection module is used to obtain the object three-dimensional information of the image to be detected in the three-dimensional detection information based on the position information of the object two-dimensional information in the two-dimensional detection information; The object detection module is used to obtain the object detection result in the image to be detected based on the object two-dimensional information and the object three-dimensional information.
9. A target detection device, characterized in that, The object detection device includes a memory and a processor coupled to the memory; Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the object detection method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the object detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Target detection method and electronic equipment
CN120612554A