Candidate box optimization method and device for three-dimensional object detection
By using pre-trained target neural networks and specific processing methods, the candidate boxes in three-dimensional object detection are optimized, and the detection inaccuracy caused by redundant candidate boxes is solved, and more accurate target detection and positioning is achieved.
Patent Information
- Application Number
- CN202510043445.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-10
AI Technical Summary
During the three-dimensional object detection process, the three-dimensional detection object is susceptible to occlusion or sensor noise, resulting in excessive redundant candidate boxes, which in turn makes the detection result in inaccurate.
By using the three-dimensional characteristics of the object to detect the object, the position confidence of the candidate box is predicted using the pre-trained target neural network, and these confidences are processed through the segmentation processing method and the normalization processing method, the position determination probability value of the candidate box is obtained, and the position confidence of the candidate box is finally suppressed to obtain the optimal candidate box.
Effectively remove redundant candidate boxes, improve the accuracy of detection results, and achieve accurate positioning of target detection objects.
Smart Images

Figure CN119992047A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a candidate frame optimization method and device for three-dimensional detection of objects. Background Art
[0002] Object detection not only needs to detect the objects in the image and give the corresponding object category, but also needs to give the location of the object in the form of a minimum bounding box (candidate box). With the development of object detection technology, three-dimensional object detection has received widespread attention and application in the field of image detection.
[0003] In the related art, because the three-dimensional detection object is easily affected by occlusion or sensor noise during the three-dimensional target detection process, too many redundant candidate boxes are generated during detection, which ultimately leads to inaccurate detection results of the three-dimensional detection object. Summary of the invention
[0004] In view of this, the present invention provides a candidate box optimization method and device for three-dimensional detection objects to solve the problem that too many redundant candidate boxes may easily lead to inaccurate detection results of three-dimensional detection objects during the three-dimensional target detection process.
[0005] According to a first aspect, an embodiment of the present disclosure provides a method for optimizing a candidate box of a three-dimensional detection object, the method comprising:
[0006] Based on the three-dimensional features of the target detection object, the target neural network is used to predict the position confidence of multiple candidate boxes of the target detection object, wherein the target neural network is pre-trained;
[0007] Based on the position confidence of multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in sequence by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames;
[0008] Based on the positioning certainty probability values of the multiple candidate boxes, a target detection area of the target detection object is obtained, and the position confidences of the multiple candidate boxes in the target detection area are obtained;
[0009] The position confidence of multiple candidate boxes in the target detection area is suppressed to obtain the optimal candidate box of the target detection object.
[0010] The disclosed embodiment predicts the position confidences of multiple candidate frames of the target detection object by using a target neural network based on the three-dimensional features of the target detection object, wherein the target neural network is pre-trained; based on the position confidences of the multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in turn by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames; based on the positioning certainty probability values of the multiple candidate frames, a target detection area of the target detection object is obtained, and the position confidences of the multiple candidate frames in the target detection area are obtained; the position confidences of the multiple candidate frames in the target detection area are suppressed to obtain the optimal candidate frame of the target detection object. Finally, the disclosed embodiment is not only conducive to accurately removing redundant candidate frames of the target detection object, but also conducive to accurately positioning the target detection object.
[0011] In some optional implementations, obtaining a target detection area of a target detection object based on the positioning certainty probability values of multiple candidate boxes includes:
[0012] Determine whether a positioning certainty probability value of a target candidate frame of the target detection object is greater than a preset threshold, wherein the target candidate frame is any candidate frame among the multiple candidate frames;
[0013] If the positioning certainty probability value of the target candidate box of the target detection object is greater than a preset threshold, it is determined that there are voxels in the area where the target candidate box is located;
[0014] If the positioning certainty probability value of the target candidate frame of the target detection object is not greater than the preset threshold, the target candidate frame is deleted;
[0015] The entire region where the voxels exist is regarded as the target detection region of the target detection object.
[0016] The disclosed embodiment facilitates accurate removal of redundant candidate frames through the above method, while more accurately locking the target detection area of the target detection object.
[0017] In some optional implementations, the candidate box optimization method for a three-dimensional detection object in the embodiment of the present disclosure further includes:
[0018] Based on the optimal candidate box of the target detection object, the target detection object in the target detection area is located.
[0019] The disclosed embodiment, based on removing redundant candidate frames of the target detection object to obtain the optimal candidate frame, can further not only quickly locate the target detection object in the target detection area, but also more accurately detect potential blurred areas even if the target detection object in the disclosed embodiment is affected by occlusion, sensor noise and other factors. Ultimately, the final positioning accuracy is improved.
[0020] In some optional implementations, optimizing multiple candidate boxes of target detection objects by segmentation processing includes:
[0021] Based on the position confidence of multiple candidate frames of the target detection object, a preset segmentation algorithm is used to perform a first segmentation on multiple candidate frames of the target detection object that are not connected to each other;
[0022] The interconnected target detection objects are obtained from the first segmentation results of multiple candidate boxes, and the interconnected target detection objects are segmented for the second time using the neighborhood maximum suppression method.
[0023] The disclosed embodiment processes the position confidences of multiple candidate boxes by segmentation processing in order to delete redundant candidate boxes of the target detection object, thereby achieving accurate positioning of the target detection object.
[0024] In some optional implementations, processing multiple candidate frames of the target detection object by normalization to obtain positioning certainty probability values of the multiple candidate frames includes:
[0025] Based on the position confidence of multiple candidate boxes of the target detection object, the total confidence sum of the target detection object is calculated;
[0026] Calculate the position confidence of the target candidate frame and divide it by the total confidence sum to obtain the positioning certainty probability value of the target candidate frame;
[0027] The positioning certainty probability values of the multiple candidate frames are calculated according to the calculation method of the positioning certainty probability value of the target candidate frame.
[0028] The disclosed embodiment processes the position confidences of multiple candidate boxes by normalization processing in order to delete redundant candidate boxes of the target detection object, thereby achieving accurate positioning of the target detection object.
[0029] In some optional implementations, suppressing the position confidences of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object includes:
[0030] The position confidences of multiple candidate boxes in the target detection area are sorted in descending order to obtain the candidate box corresponding to the maximum position confidence;
[0031] The candidate box corresponding to the maximum position confidence is taken as the optimal candidate box for the target detection object.
[0032] The disclosed embodiment uses the above method to achieve accurate removal of interfering candidate frames, thereby obtaining the optimal candidate frame of the target detection object, and ultimately achieving the purpose of accurately locating the target detection object.
[0033] According to a second aspect, an embodiment of the present disclosure provides a candidate box optimization device for three-dimensional detection objects, the device comprising:
[0034] A confidence prediction module, used to predict the position confidence of multiple candidate boxes of the target detection object using a target neural network based on the three-dimensional features of the target detection object, wherein the target neural network is pre-trained;
[0035] A confidence processing module is used to process the multiple candidate frames of the target detection object in sequence through a segmentation processing method and a normalization processing method based on the position confidence of the multiple candidate frames of the target detection object to obtain positioning certainty probability values of the multiple candidate frames;
[0036] A confidence determination module, used to obtain a target detection area of a target detection object based on the positioning certainty probability values of multiple candidate boxes, and obtain the position confidences of multiple candidate boxes in the target detection area;
[0037] The optimal box determination module is used to suppress the position confidence of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object.
[0038] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the candidate box optimization method for three-dimensional detection objects according to the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the candidate box optimization method for a three-dimensional detection object according to the first aspect or any corresponding embodiment thereof.
[0040] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the candidate box optimization method for three-dimensional detection objects according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1is a schematic flow chart of a candidate box optimization method for three-dimensional detection objects according to an embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of multiple candidate boxes of a target detection object before optimization according to an embodiment of the present invention;
[0044] Figure 3 is a schematic diagram of multiple candidate boxes of a target detection object after optimization according to an embodiment of the present invention;
[0045] Figure 4 is another schematic diagram of a flow chart of a candidate box optimization method for a three-dimensional detection object according to an embodiment of the present invention;
[0046] Figure 5 is a structural block diagram of a candidate box optimization device for three-dimensional detection objects according to an embodiment of the present invention;
[0047] Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0049] According to an embodiment of the present invention, an embodiment of a candidate box optimization method for a three-dimensional detected object is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0050] In this embodiment, a candidate box optimization method for three-dimensional detection objects is provided, which can be used in computer devices such as mobile phones, tablet computers, desktop computers, portable notebooks, servers, etc. The candidate box optimization method for three-dimensional detection objects in the disclosed embodiment can be specifically applied to target object detection in scenarios of autonomous driving or robot or drone navigation. In these scenarios, it helps the detection system to perceive and understand the surrounding environment in real time and make safe and accurate decision-making tasks, such as making decisions on multi-target tracking, vehicle operation trajectory, and other decision-making tasks. Figure 1 is a flowchart of a candidate box optimization method for a three-dimensional detection object according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0051] Step S101, based on the three-dimensional features of the target detection object, using the target neural network to predict the position confidence of multiple candidate boxes of the target detection object, wherein the target neural network is pre-trained.
[0052] Specifically, the target detection object can be any one of the multiple detection objects in the current detection environment. For example, the disclosed embodiment is applied to an autonomous driving scenario, and the target detection object can be a road sign or a bridge or a traffic light on the road where the vehicle is traveling. Through the three-dimensional features of the target detection object, the three-dimensional structure of the target detection object can be known, and then the spatial position, properties and dynamic changes of the target detection object can be known. The perception technology of three-dimensional features is not only widely used in the scenarios of autonomous driving or robot or drone navigation, but also in the development of immersive experiences such as augmented reality (AR) and virtual reality (VR), and plays an irreplaceable role in many fields such as medicine, construction, and entertainment.
[0053] In some optional implementations, the above step S101, obtaining the three-dimensional features of the target detection object, includes:
[0054] Step a1, obtaining grid sampling points from the target candidate frame of the target detection object.
[0055] Specifically, the target candidate box is any candidate box among the multiple candidate boxes, and the shape of any candidate box can be a rectangle. For example, grid-shaped sampling points are taken in the rectangular target candidate box, such as a 7x3x3 grid.
[0056] Step a2, using the internal and external parameters of the camera to project the grid sampling points into a two-dimensional space to obtain two-dimensional image data of the target detection object.
[0057] Specifically, the sampling points in the target candidate frame are projected onto the two-dimensional camera imaging plane using the camera internal and external parameters to obtain the two-dimensional image data of the target detection object.
[0058] Step a3, converting the two-dimensional image data of the target detection object into three-dimensional image data.
[0059] Step a4, extracting the three-dimensional features of the target detection object from the three-dimensional image data.
[0060] In a specific example, in step S101, based on the three-dimensional features of the target detection object, the target neural network is used to predict the position confidence of multiple candidate boxes of the target detection object, wherein the target neural network is pre-trained and includes:
[0061] The three-dimensional features of the target detection object are input into the detection module of the target neural network for prediction, and the position confidence of multiple candidate boxes of the target detection object output by the target neural network is obtained.
[0062] Specifically, the target neural network is a general neural network with a detection head. The target neural network is used to model the position confidence of the target detection object. In order to further locate the current position of the target detection object, the target neural network is pre-trained based on three-dimensional features. The detection module is a network of the detection head in the target neural network. The detection module is used to predict the position confidence of multiple candidate boxes of the target detection object. The position confidence is used to represent the relevant information of the position of the target detection object in the image; using the detection module of the target neural network to predict the position confidence of multiple candidate boxes of the target detection object can improve the robustness of the prediction result.
[0063] Step S102, based on the position confidence of multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in sequence by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames.
[0064] Specifically, in the process of 3D object detection, the 3D detection object is easily affected by occlusion or sensor noise, thus generating too many redundant candidate frames, which ultimately leads to inaccurate detection results of the 3D detection object. Therefore, it is necessary to further remove redundancy from multiple candidate frames of the target detection object to ensure the accuracy of the target detection object positioning.
[0065] Furthermore, the positioning certainty probability value indicates the degree of certainty that there is voxel data of the target detection object at the location of the target candidate frame, which is an indicator for measuring the quality of the positioning information of the target detection object. Generally, the larger the positioning certainty probability value, the greater the reliability of the positioning information of the target detection object. Both the segmentation processing method and the normalization processing method are to delete the redundant candidate frames of the target detection object, thereby achieving accurate positioning of the target detection object.
[0066] In some optional implementations, optimizing multiple candidate boxes of target detection objects by segmentation processing includes:
[0067] Step c1, based on the position confidence of multiple candidate frames of the target detection object, a preset segmentation algorithm is used to perform a first segmentation on multiple candidate frames of the target detection object that are not connected to each other.
[0068] Specifically, for example, a bird's-eye view is generated based on the position confidence of multiple candidate boxes of the target detection object, and the bird's-eye view represents an intensity distribution map of the position confidence of multiple candidate boxes in three-dimensional space. A segmentation algorithm such as a connected domain segmentation algorithm is applied to separate the confidence distribution areas of multiple candidate boxes that are not connected to each other.
[0069] Step c2, obtaining interconnected target detection objects from the first segmentation results of multiple candidate boxes, and performing a second segmentation on the interconnected target detection objects using a neighborhood maximum suppression method.
[0070] Specifically, for example, Figure 2 As shown, if the areas where the multiple blue candidate boxes are located are connected to the areas where the multiple purple candidate boxes are located, the blue area is separated from the purple area by using the neighborhood maximum suppression method. The maximum suppression method usually refers to the non-maximum suppression (NMS) algorithm. The working principle of the NMS algorithm is based on the judgment of local maximum values. For any candidate box, the NMS algorithm calculates the overlap between any candidate box and other candidate boxes (usually using the intersection-and-union ratio IOU as a metric). If the overlap between a candidate box and other candidate boxes exceeds the set threshold and its position confidence is not the highest, then the candidate box will be suppressed (i.e. filtered out), and those candidate boxes with the highest local position confidence and whose overlap with other candidate boxes does not exceed the threshold will be retained.
[0071] In some optional implementations, multiple candidate frames of the target detection object are processed by normalization to obtain positioning certainty probability values of the multiple candidate frames, including:
[0072] Step d1, based on the position confidences of multiple candidate boxes of the target detection object, calculate the total confidence sum of the target detection object.
[0073] Specifically, the position confidences of multiple candidate boxes of the target detection object are 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, and 0.9, respectively, and the sum of the confidences of the multiple candidate boxes is 3.
[0074] Step d2, calculate the position confidence of the target candidate box and divide it by the total confidence sum to obtain the positioning certainty probability value of the target candidate box.
[0075] Specifically, the target candidate box is any candidate box of multiple candidate boxes, then the position confidence of any of the above target candidate boxes divided by the total confidence sum is: 0.6 / 3=0.2; 0.5 / 3≈0.17, 0.4 / 3≈0.13, 0.3 / 3=0.1, 0.2 / 3≈0.07, 0.1 / 3≈0.03, 0.9 / 3=0.3.
[0076] Step d3, calculating the positioning certainty probability values of multiple candidate boxes according to the calculation method of the positioning certainty probability value of the target candidate box.
[0077] Specifically, according to the above calculation method, the positioning certainty probability values of multiple candidate boxes are obtained as 0.2, 0.17, 0.13, 0.1, 0.07, 0.03, and 0.3, respectively.
[0078] Step S103, based on the positioning certainty probability values of the multiple candidate boxes, obtain the target detection area of the target detection object, and obtain the position confidence of the multiple candidate boxes in the target detection area.
[0079] Specifically, for example, in Figure 2 In , for multiple red candidate boxes, the target detection area is the area where the multiple red candidate boxes are located, and the position confidences of the multiple red candidate boxes are further obtained.
[0080] In some optional implementations, the above step S103, based on the positioning certainty probability values of the plurality of candidate boxes, obtains the target detection area of the target detection object, including:
[0081] Step e1, determining whether the positioning certainty probability value of a target candidate frame of the target detection object is greater than a preset threshold, wherein the target candidate frame is any candidate frame among multiple candidate frames.
[0082] Specifically, for Figure 2 For multiple red candidate boxes, it is determined whether the positioning certainty probability value of the target candidate box (any red candidate box) is greater than a preset threshold. The preset threshold is a pre-selected reference value, which can be flexibly set in combination with the actual application scenario. The preset threshold can be set in combination with some parameters with smaller positioning certainty probability values. The preset threshold is used to filter out candidate boxes with smaller positioning certainty probability values.
[0083] Step e2: if the positioning certainty probability value of the target candidate box of the target detection object is greater than a preset threshold, it is determined that there are voxels in the area where the target candidate box is located.
[0084] Specifically, if the positioning certainty probability value of the target candidate box (any candidate box) of the target detection object is greater than a preset threshold, it is determined that there are voxels in the area where the target candidate box (any candidate box) is located; otherwise, there are no voxels in the area where the target candidate box (any candidate box) is located.
[0085] Step e3: If the positioning certainty probability value of the target candidate box of the target detection object is not greater than a preset threshold, the target candidate box is deleted.
[0086] Specifically, for example, in Figure 2 In the target detection, if the positioning certainty probability value of the target candidate box of the target detection object is not greater than the preset threshold, the corresponding candidate box is deleted.
[0087] The disclosed embodiment further determines whether the positioning certainty probability value of the target candidate frame of the target detection object is greater than a preset threshold based on the positioning certainty probability values of the multiple candidate frames obtained from the above-mentioned normalization processing results, in order to screen out candidate frames with better positioning certainty probability values, thereby facilitating the accurate removal of irrelevant redundant candidate frames.
[0088] Step e4, taking the entire area where the voxels exist as the target detection area of the target detection object.
[0089] Specifically, based on the above judgment results, the entire area where the voxels exist is used as the target detection area of the target detection object. Figure 2 In the above example, if there are voxels in the regions where multiple red candidate boxes are located, the regions where these multiple red candidate boxes are located are regarded as target detection regions of the same target detection object.
[0090] Step S104, suppressing the position confidence of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object.
[0091] In some optional implementations, the above step S104 suppresses the position confidences of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object, including:
[0092] Step d1, sorting the position confidences of multiple candidate boxes in the target detection area in descending order, and obtaining the candidate box corresponding to the maximum position confidence.
[0093] In step d2, the candidate box corresponding to the maximum position confidence is used as the optimal candidate box for the target detection object.
[0094] Specifically, for example, in Figure 2 For example, the position confidences of multiple red candidate boxes are 0.8, 0.75, 0.63, 0.61, and 0.55 in descending order, and the candidate box corresponding to the maximum confidence of 0.8 is selected as the optimal candidate box. Figure 2 In the example, multiple candidate boxes in the same color set represent the same target detection object. Figure 2 The multiple candidate boxes in the same color set in the same target detection object are optimized according to the above steps S101 to S104, and then the following is obtained: Figure 3 The optimization results are shown.
[0095] The candidate frame optimization method for three-dimensional detection objects in the embodiments of the present disclosure predicts the position confidences of multiple candidate frames of the target detection object by using a target neural network based on the three-dimensional features of the target detection object, wherein the target neural network is pre-trained; based on the position confidences of the multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in turn by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames; based on the positioning certainty probability values of the multiple candidate frames, a target detection area of the target detection object is obtained, and the position confidences of the multiple candidate frames in the target detection area are obtained; the position confidences of the multiple candidate frames in the target detection area are suppressed to obtain the optimal candidate frame of the target detection object. Finally, the embodiments of the present disclosure are not only conducive to accurately removing redundant candidate frames of the target detection object, but also conducive to accurately positioning the target detection object.
[0096] In this embodiment, a candidate box optimization method for three-dimensional detection objects is provided, which can be used in computer devices, such as mobile phones, tablet computers, desktop computers, portable notebooks, servers, etc. The candidate box optimization method for three-dimensional detection objects in the disclosed embodiment can be specifically applied to target object detection in scenarios of autonomous driving or robot or drone navigation. In these scenarios, it helps the detection system to perceive and understand the surrounding environment in real time and make safe and accurate decision-making tasks, such as multi-target tracking, vehicle operation trajectory and other decision-making tasks. Figure 4 is a flowchart of a candidate box optimization method for a three-dimensional detection object according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:
[0097] Step S101, based on the three-dimensional features of the target detection object, using the target neural network to predict the position confidence of multiple candidate boxes of the target detection object, wherein the target neural network is pre-trained.
[0098] Step S102, based on the position confidence of multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in sequence by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames.
[0099] Step S103, based on the positioning certainty probability values of the multiple candidate boxes, obtain the target detection area of the target detection object, and obtain the position confidence of the multiple candidate boxes in the target detection area.
[0100] Step S104, suppressing the position confidence of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object.
[0101] Specifically, the above steps S101 to S104 have been explained in the above process and will not be repeated here.
[0102] Step S105: locate the target detection object in the target detection area based on the optimal candidate box of the target detection object.
[0103] The disclosed embodiment locates the target detection object in the target detection area based on the optimal candidate frame of the target detection object, which is conducive to accurate tracking of the target detection object in different driving scenarios.
[0104] In this embodiment, a candidate box optimization device for a three-dimensional detection object is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated here. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0105] This embodiment provides a candidate box optimization device for three-dimensional detection objects, such as Figure 5 As shown, including:
[0106] A confidence prediction module 501 is used to predict the position confidence of multiple candidate boxes of the target detection object using a target neural network based on the three-dimensional features of the target detection object, wherein the target neural network is pre-trained;
[0107] A confidence processing module 502 is used to process the multiple candidate frames of the target detection object in sequence through a segmentation processing method and a normalization processing method based on the position confidence of the multiple candidate frames of the target detection object to obtain positioning certainty probability values of the multiple candidate frames;
[0108] A confidence determination module 503 is used to obtain a target detection area of the target detection object based on the positioning certainty probability values of the multiple candidate boxes, and obtain the position confidences of the multiple candidate boxes in the target detection area;
[0109] The optimal frame determination module 504 is used to suppress the position confidence of multiple candidate frames in the target detection area to obtain the optimal candidate frame of the target detection object.
[0110] In some optional implementations, the confidence determination module 503 includes:
[0111] A candidate box determination submodule is used to determine whether a positioning certainty probability value of a target candidate box of a target detection object is greater than a preset threshold, wherein the target candidate box is any candidate box among multiple candidate boxes;
[0112] A first determination submodule is used to determine that there are voxels in the area where the target candidate box is located if the positioning certainty probability value of the target candidate box of the target detection object is greater than a preset threshold;
[0113] A candidate box deletion submodule is used to delete the target candidate box if the positioning certainty probability value of the target candidate box of the target detection object is not greater than a preset threshold;
[0114] The second determination submodule is used to take the entire area where the voxels exist as the target detection area of the target detection object.
[0115] In some optional implementations, the candidate box optimization device for three-dimensional detection objects in the embodiment of the present disclosure, Figure 5 Also included:
[0116] The target positioning module 505 is used to position the target detection object in the target detection area based on the optimal candidate box of the target detection object.
[0117] In some optional implementations, the confidence processing module 502 includes:
[0118] A first segmentation submodule, configured to perform a first segmentation on multiple unconnected candidate frames of the target detection object using a preset segmentation algorithm based on the position confidence of multiple candidate frames of the target detection object;
[0119] The second segmentation submodule is used to obtain interconnected target detection objects from the first segmentation results of multiple candidate boxes, and perform a second segmentation on the interconnected target detection objects using a neighborhood maximum suppression method.
[0120] In some optional implementations, the confidence processing module 502 includes:
[0121] A first calculation submodule, configured to calculate a total confidence sum of the target detection object based on the position confidences of multiple candidate boxes of the target detection object;
[0122] The second calculation submodule is used to calculate the position confidence of the target candidate frame and divide it by the total confidence sum to obtain a positioning certainty probability value of the target candidate frame;
[0123] The third calculation submodule is used to calculate the positioning certainty probability values of multiple candidate boxes according to the calculation method of the positioning certainty probability value of the target candidate box.
[0124] In some optional implementations, the optimal frame determination module 504 includes:
[0125] The confidence ranking submodule is used to sort the position confidences of multiple candidate boxes in the target detection area in descending order to obtain the candidate box corresponding to the maximum position confidence;
[0126] The candidate box determination submodule is used to take the candidate box corresponding to the maximum position confidence as the optimal candidate box for the target detection object.
[0127] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0128] The candidate box optimization device for three-dimensional detection objects in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0129] An embodiment of the present invention further provides a computer device having the above-mentioned candidate box optimization device for three-dimensional detection objects.
[0130] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0131] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0132] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0133] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0134] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0135] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0136] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0137] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0138] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A candidate box optimization method for three-dimensional object detection, characterized in that: The method comprises: Based on the three-dimensional features of the target detection object, using a target neural network to predict the position confidence of multiple candidate boxes of the target detection object, wherein the target neural network is pre-trained; Based on the position confidences of the multiple candidate frames of the target detection object, the multiple candidate frames of the target detection object are processed in sequence by a segmentation processing method and a normalization processing method to obtain positioning certainty probability values of the multiple candidate frames; Based on the positioning certainty probability values of the multiple candidate boxes, a target detection area of the target detection object is obtained, and position confidences of the multiple candidate boxes in the target detection area are obtained; The position confidences of multiple candidate boxes in the target detection area are suppressed to obtain an optimal candidate box for the target detection object.
2. The method according to claim 1, characterized in that Based on the positioning certainty probability values of the plurality of candidate frames, obtaining the target detection area of the target detection object includes: Determine whether a positioning certainty probability value of a target candidate frame of the target detection object is greater than a preset threshold, wherein the target candidate frame is any candidate frame among the multiple candidate frames; If the positioning certainty probability value of the target candidate frame of the target detection object is greater than the preset threshold, it is determined that there are voxels in the area where the target candidate frame is located; If the positioning certainty probability value of the target candidate frame of the target detection object is not greater than the preset threshold, deleting the target candidate frame; The entire region where the voxels exist is regarded as the target detection region of the target detection object.
3. The method according to claim 1, characterized in that Also includes: Based on the optimal candidate frame of the target detection object, the target detection object in the target detection area is located.
4. The method according to claim 1, characterized in that Processing multiple candidate frames of the target detection object by segmentation processing includes: Based on the position confidences of the multiple candidate frames of the target detection object, a preset segmentation algorithm is used to perform a first segmentation on the multiple candidate frames of the target detection object that are not connected to each other; Interconnected target detection objects are obtained from the first segmentation results of multiple candidate boxes, and the interconnected target detection objects are segmented for a second time using a neighborhood maximum suppression method.
5. The method according to claim 1 or 4, characterized in that: Processing multiple candidate frames of the target detection object in a normalized manner to obtain positioning certainty probability values of the multiple candidate frames includes: Calculating a total confidence sum of the target detection object based on the position confidences of multiple candidate frames of the target detection object; Calculate the position confidence of the target candidate frame and divide it by the total confidence sum to obtain a positioning certainty probability value of the target candidate frame; The positioning certainty probability values of the multiple candidate frames are calculated according to the calculation method of the positioning certainty probability value of the target candidate frame.
6. The method according to claim 1, characterized in that Suppressing the position confidences of multiple candidate boxes in the target detection area to obtain the optimal candidate box of the target detection object includes: The position confidences of the multiple candidate boxes in the target detection area are sorted in descending order to obtain a candidate box corresponding to the maximum position confidence; The candidate box corresponding to the maximum position confidence is used as the optimal candidate box of the target detection object.
7. A candidate box optimization device for three-dimensional object detection, characterized in that: The device comprises: A confidence prediction module, used to predict the position confidence of multiple candidate boxes of the target detection object using a target neural network based on the three-dimensional features of the target detection object, wherein the target neural network is pre-trained; A confidence processing module, used to process the multiple candidate frames of the target detection object in turn by a segmentation processing method and a normalization processing method based on the position confidence of the multiple candidate frames of the target detection object, so as to obtain positioning certainty probability values of the multiple candidate frames; A confidence determination module, used to obtain a target detection area of the target detection object based on the positioning certainty probability values of multiple candidate boxes, and obtain the position confidences of multiple candidate boxes in the target detection area; The optimal frame determination module is used to suppress the position confidence of multiple candidate frames in the target detection area to obtain the optimal candidate frame of the target detection object.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the candidate box optimization method for three-dimensional detection objects according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the candidate box optimization method for three-dimensional detection objects according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the candidate box optimization method for a three-dimensional detection object according to any one of claims 1 to 6.
Citation Information
Patent Citations
Object recognition method and device, terminal equipment and storage medium
CN112348778A
Target detection positioning confidence determination method and device, electronic equipment and storage medium
CN112668573A
Non-maximum suppression method, system and device based on attention mechanism and medium
CN114723939A
Target object detection method and apparatus, and electronic device and storage medium
WO2022198786A1
Cited By
3D multi-target tracking method and system based on motion state prediction
CN120260017A