Target detection method, device, electronic device and storage medium

By using multiple object detection models and confidence correction models in object detection, combined with soft non-maximum suppression algorithm, the object detection results are screened and corrected, which solves the problem of low detection accuracy of a single model and improves the accuracy and reliability of object detection.

CN119295715BActive Publication Date: 2025-05-13北京观微科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411825834.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-13
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

In the prior art, the single object detection model has low accuracy in object detection, resulting in unreliable detection results.

Method used

The final object detection box and type are obtained by inputting the image to be detected into at least two object detection models and filtering and correcting the detection results using the confidence correction model and soft non-maximum suppression algorithm.

Benefits of technology

The accuracy of target detection is improved, the advantages of multiple target detection models are fully utilized, false detection and missed detection are reduced, and the reliability of detection results is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119295715B_ABST
    Figure CN119295715B_ABST
Patent Text Reader

Abstract

The present invention provides a target detection method, device, electronic device and storage medium, which relate to the field of detection technology, wherein the method comprises: inputting an image to be detected into at least two target detection models, obtaining target detection results output by each target detection model, the target detection results including detection frames and confidences corresponding to each target object; inputting each confidence under a type into a corresponding confidence correction model, obtaining each corrected confidence; screening each detection frame based on each corrected confidence under a type by a soft non-maximum suppression algorithm; determining the type corresponding to the screened target detection frame as the type of the target object. It can be seen that the present invention can screen a target detection frame from all detection frames by confidence correction and a soft non-maximum suppression algorithm, and determine the type corresponding to the target detection frame as the type of the corresponding target object, giving full play to the advantages of each target detection model, thereby improving the accuracy of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection technology, and in particular to a target detection method, device, electronic equipment and storage medium. Background Art

[0002] Object detection is a key computer vision technique that aims to simultaneously identify the class of objects and determine their locations in images or videos.

[0003] In the related art, an image to be detected including a target object is usually input into a target detection model to obtain the position and type of the target object output by the target detection model.

[0004] However, in the above-mentioned related technologies, the target is detected based on only a single target detection model, which reduces the accuracy of target detection. Summary of the invention

[0005] The present invention provides a target detection method, device, electronic device and storage medium, which are used to solve the defect of reducing the accuracy of target detection in the prior art.

[0006] The present invention provides a target detection method, comprising the following steps.

[0007] Inputting an image to be detected including at least one target object into at least two target detection models, and obtaining target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes;

[0008] For each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type by using a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0009] The detection frames are screened based on the corrected confidences of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0010] The type corresponding to the target detection frame is determined as the type of the target object included in the target detection frame.

[0011] According to a target detection method provided by the present invention, the confidence correction model is trained based on the following method:

[0012] Inputting a sample image including at least one sample object into each of the target detection models to obtain a prediction detection result output by each of the target detection models, wherein the prediction detection result includes a prediction detection box corresponding to each of the sample objects and a prediction confidence representing a prediction type of the sample object included in each of the prediction detection boxes;

[0013] For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type;

[0014] Based on the confidence correction sub-models of the target detection model under all the prediction types, a confidence correction model corresponding to the target detection model is determined.

[0015] According to a target detection method provided by the present invention, for each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type, including:

[0016] For each predicted detection box under each prediction type output by the target detection model, determine a first intersection-over-union ratio between the predicted detection box and the corresponding true detection box;

[0017] Determining a loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence;

[0018] Based on the loss function, the model parameters of the initial confidence correction model under the prediction type are iteratively adjusted to obtain the confidence correction sub-model of the target detection model under the prediction type.

[0019] According to a target detection method provided by the present invention, determining a loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence includes:

[0020] Determine the loss function under the prediction type based on the following formula (1);

[0021] (1)

[0022] in, Indicates the prediction type, Indicates the prediction type The next The first intersection-over-union ratio between the predicted detection box and the corresponding real detection box, Indicates the prediction type The corresponding prediction confidence under Indicates the prediction type The number of predicted detection boxes, satisfy , represents a monotonically increasing constraint.

[0023] According to a target detection method provided by the present invention, the method of screening each detection frame based on each corrected confidence of each target detection model under the type by a soft non-maximum suppression algorithm to obtain a target detection frame includes:

[0024] For a confidence set corresponding to all the corrected confidences under the type, determine a maximum corrected confidence under the type in the confidence set;

[0025] Determine a second intersection-over-union ratio between a detection frame corresponding to a maximum corrected confidence under the type and a detection frame corresponding to other corrected confidences under the type; the other corrected confidences are corrected confidences in the confidence set except the maximum corrected confidence;

[0026] For each of the second I / O ratios, updating other corrected confidences of the type corresponding to the second I / O ratio based on the second I / O ratio and an attenuation parameter to obtain an updated confidence of the type corresponding to the second I / O ratio;

[0027] Deleting the detection boxes corresponding to the updated confidences that are less than the preset confidences, to obtain the remaining detection boxes of the type;

[0028] Determine an object detection frame of the type based on remaining detection frames of the type.

[0029] According to a target detection method provided by the present invention, the updating of other corrected confidences corresponding to the second IoU ratio based on the second IoU ratio and the attenuation parameter to obtain an updated confidence corresponding to the second IoU ratio includes:

[0030] Determine an updated confidence level corresponding to the second intersection-over-union ratio based on the following formula (2);

[0031] (2)

[0032] in, Indicates the type, Representation Type Next Other corrected confidence levels, Indicates the type Next The detection boxes corresponding to other corrected confidences, Represents the confidence set The corrected confidence is the maximum corrected confidence, Indicates the type The detection box corresponding to the maximum corrected confidence level, represents the attenuation parameter, Indicates the type Next The updated confidence corresponding to the other corrected confidences.

[0033] The present invention also provides a target detection device, comprising:

[0034] A detection unit, configured to input an image to be detected including at least one target object into at least two target detection models, and obtain target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes;

[0035] A correction unit, for each type output by each target detection model, inputting each confidence under the type into a confidence correction model corresponding to the target detection model, and correcting each confidence under the type by a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0036] A screening unit, configured to screen each of the detection frames based on each corrected confidence of each of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0037] A determination unit is used to determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame.

[0038] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any of the target detection methods described above is implemented.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the target detection method as described above is implemented.

[0040] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the target detection methods described above.

[0041] The target detection method, device, electronic device and storage medium provided by the present invention input an image to be detected including at least one target object into at least two target detection models, obtain the detection frame and type confidence corresponding to each target object output by each target detection model, input each confidence under the type into the confidence correction model corresponding to the target detection model for each type output by each target detection model, correct each confidence under the type by the confidence correction submodel corresponding to the type in the confidence correction model, obtain each corrected confidence under the type, screen each detection frame based on each corrected confidence under the type of each target detection model by the soft non-maximum suppression algorithm, obtain the target detection frame, and then determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame. It can be seen that the present invention can screen the target detection frame from all detection frames through the confidence correction and soft non-maximum suppression algorithms, determine the type corresponding to the target detection frame as the type of the corresponding target object, give full play to the advantages of each target detection model, and thus improve the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 This is one of the flowcharts of the target detection method provided by the embodiment of the present invention.

[0044] Figure 2 This is one of the flow charts of the training method of the confidence correction model provided in an embodiment of the present invention.

[0045] Figure 3 This is the second flow chart of the training method of the confidence correction model provided in the embodiment of the present invention.

[0046] Figure 4 This is the second flowchart of the target detection method provided by the embodiment of the present invention.

[0047] Figure 5 It is a schematic diagram of the overall process of the target detection method provided by an embodiment of the present invention.

[0048] Figure 6It is a comparative schematic diagram of target detection of a single target detection model and multiple target detection models provided in an embodiment of the present invention.

[0049] Figure 7 It is a schematic diagram of the structure of a target detection device provided by an embodiment of the present invention.

[0050] Figure 8 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Object detection technology usually extracts features through convolutional neural networks (CNNs) to distinguish objects from backgrounds, and further generates bounding boxes and classification results of objects through regression methods. The accuracy and efficiency of object detection have been significantly improved in recent years, especially object detection based on deep learning algorithms, such as Faster Region Convolutional Neural Network (FasterR-CNN), YOLO or Single Shot MultiBox Detector (SSD). These algorithms use multi-layer convolutional networks to extract features from images, and combine with region proposal networks or directly predict the location and category of objects to achieve fast and accurate detection.

[0053] As research deepens, target detection technology will play an important role in more practical scenarios, such as autonomous driving, drone monitoring, smart security, etc. During the application process, a single target detection model will reduce the accuracy of target detection due to the characteristics of training or its own network structure.

[0054] Based on this, the present invention proposes a new target detection method, which can filter out a target detection frame from all detection frames through confidence correction and soft non-maximum suppression algorithm, and determine the type corresponding to the target detection frame as the type of the corresponding target object, giving full play to the advantages of each target detection model, thereby improving the accuracy of target detection.

[0055] Combine the following Figure 1-Figure 6The target detection method of the present invention is described. The execution subject of the target detection method can be an electronic device such as a camera device, a terminal or a computer, or a target detection device set in the electronic device, and the target detection device can be implemented by software, hardware or a combination of the two.

[0056] Figure 1 is one of the flow charts of the target detection method provided by the embodiment of the present invention, such as Figure 1 As shown, the target detection method includes the following steps:

[0057] Step 101: input an image to be detected including at least one target object into at least two target detection models to obtain target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes.

[0058] The image to be detected may be an image captured in real time by a camera device or a previously stored image, and the image to be detected may include one or more target objects, such as vehicles, persons or other targets to be detected.

[0059] For example, when an image to be detected including at least one target object is obtained, the image to be detected including at least one target object is input into each target detection model, and each target detection model performs feature analysis on each target object in the input image to be detected to obtain a corresponding target detection result output by each target detection model. The target detection result includes a detection box output by the target detection model where each target object is located and the confidence of the type of target object included in each detection box. For example, if the image to be detected includes vehicle 1 and person 1, the target detection model outputs a detection box where vehicle 1 is located and the confidence that vehicle 1 belongs to the vehicle type, as well as a detection box where person 1 is located and the confidence that person 1 belongs to the person type.

[0060] It should be noted that each target detection model is a trained detection model. Specifically, the training process of the target detection model is as follows: obtain multiple sample images, input the multiple sample images into the initial detection model, obtain the predicted detection box where the sample object is located in each sample image output by the initial detection model, determine the loss function based on the predicted detection box where the sample object is located and the actual detection box of the sample object, adjust the model parameters of the initial detection model based on the loss function, and obtain the target detection model.

[0061] It should be noted that the model structures of all target detection models are different. The specific structure of the target detection model can be deep neural network (DNN), convolutional neural network (CNN), Yolov5 or Faster R-CNN, etc. For example, if the number of target detection models is 2, the training data set can be used to train the target detection model based on Yolov5. And the object detection model based on Faster R-CNN , the present invention does not limit this.

[0062] Step 102: for each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type through a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type.

[0063] The confidence correction model includes at least one confidence correction sub-model corresponding to a type.

[0064] For example, for each target detection model, when the image to be detected includes a target object, the target detection model outputs only one type of confidence and detection box. The confidence of this type is input into the confidence correction model corresponding to the target detection model. The confidence of this type is corrected by the confidence correction sub-model corresponding to the type in the confidence correction model to obtain the corrected confidence of this type.

[0065] When the image to be detected includes multiple target objects, the target detection model outputs multiple types of confidences and detection boxes. Each type of confidence and detection box is input into the confidence correction model corresponding to the target detection model, and each type of confidence is input into the corresponding confidence correction sub-model. The confidence of the input type is corrected by the corresponding confidence correction sub-model to obtain the corrected confidence corresponding to the input type.

[0066] Step 103: Screen the detection frames based on the corrected confidences of the target detection models under the type using a soft non-maximum suppression algorithm to obtain a target detection frame.

[0067] For example, when obtaining the corrected confidences of all types output by the confidence correction models corresponding to all target detection models, considering that there may be overlapping redundancy in the detection frames output by multiple target detection models, the present invention uses a soft non-maximum suppression algorithm (Soft Non Maximum Suppression, Soft-NMS) to screen and filter all detection frames based on the corrected confidences of each target detection model under each type to obtain the final target detection frame.

[0068] Step 104: Determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame.

[0069] For example, when filtering out the target detection frame, the type corresponding to the target detection frame is used as the type of the target object included in the target detection frame, thereby realizing the detection of the target object in the image to be detected. If the image to be detected includes multiple target objects, each target object has a corresponding target detection frame and a corresponding target object type.

[0070] The target detection method provided by the present invention inputs an image to be detected including at least one target object into at least two target detection models, obtains the detection frame and type confidence corresponding to each target object output by each target detection model, inputs each confidence under the type into the confidence correction model corresponding to the target detection model for each type output by each target detection model, corrects each confidence under the type by the confidence correction submodel corresponding to the type in the confidence correction model, obtains each corrected confidence under the type, and screens each detection frame based on each corrected confidence under the type of each target detection model by a soft non-maximum suppression algorithm to obtain a target detection frame, and then determines the type corresponding to the target detection frame as the type of the target object included in the target detection frame. It can be seen that the present invention can screen the target detection frame from all detection frames through the confidence correction and soft non-maximum suppression algorithms, and determine the type corresponding to the target detection frame as the type of the corresponding target object, giving full play to the advantages of each target detection model, thereby improving the accuracy of target detection. In addition, the present invention performs target detection under a trained target detection model without the need for additional model training, which can reduce the amount of calculation caused by repeated training of the target detection model, thereby reducing the occupation of electronic device graphics card resources during the training process.

[0071] In one embodiment, Figure 2 is one of the flow charts of the training method of the confidence correction model provided in the embodiment of the present invention, such as Figure 2 As shown, the confidence correction model is trained based on the following method:

[0072] Step 201: Input a sample image including at least one sample object into each of the target detection models to obtain a predicted detection result output by each of the target detection models, wherein the predicted detection result includes a predicted detection box corresponding to each of the sample objects and a predicted confidence level representing a predicted type of the sample object included in each of the predicted detection boxes.

[0073] For example, take the target detection model and target detection models For example, a number of sample images including at least one sample object and the real detection box of the sample object in each sample image are randomly selected from the validation data set of the training data set, and these sample images are respectively input into the target detection model and target detection models , through the target detection model and target detection models The non-maximum suppression algorithm is used to obtain the predicted detection results. and , target detection model Output prediction test results Including the predicted detection box corresponding to each sample object and the prediction confidence of the prediction type corresponding to each predicted detection box, the target detection model Output prediction test results It includes the predicted detection box corresponding to each sample object and the prediction confidence of the prediction type corresponding to each predicted detection box.

[0074] Step 202: For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and the real detection box of the sample image corresponding to the prediction type, the initial confidence correction model of the target detection model under the prediction type is trained to obtain the confidence correction sub-model of the target detection model under the prediction type.

[0075] For example, target detection models with similar accuracy evaluation indicators may have different recognition confidences. This difference mainly comes from factors such as the model structure, model output head, training hyperparameters, and training time of the target detection model. Simply merging the results of multiple target detection models is easily affected by the target detection model with high confidence, so it is necessary to use the real detection frame to calibrate the confidence of each target detection model.

[0076] Specifically, for each prediction type output by each target detection model, based on each prediction detection frame under the prediction type, each prediction confidence under the prediction type, and the real detection frame of the sample image corresponding to the prediction type, the loss function corresponding to the prediction type is determined, and the initial confidence correction model of the target detection model under the prediction type is fitted using a monotone regression algorithm. The model parameters of the initial confidence correction model are adjusted based on the loss function corresponding to the prediction type. In the same way, the initial confidence correction model of the target detection model under all prediction types can be obtained.

[0077] Step 203: Determine a confidence correction model corresponding to the target detection model based on the confidence correction sub-models of the target detection model under all the prediction types.

[0078] For example, when the initial confidence correction model of the target detection model under all prediction types is obtained, the initial confidence correction models under all prediction types are combined to obtain the confidence correction model corresponding to the target detection model. and target detection models As an example, we can get the target detection model Corresponding confidence correction model , and target detection models Corresponding confidence correction model .

[0079] In this embodiment, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and the real detection box of the sample image corresponding to the prediction type, the initial confidence correction model of the target detection model under the prediction type is trained to obtain the confidence correction sub-model of the target detection model under the prediction type, and based on each confidence correction sub-model of the target detection model under all prediction types, the confidence correction model corresponding to the target detection model is determined, thereby realizing the acquisition of confidence correction models corresponding to different target detection models, facilitating the correction of the confidence output of the corresponding target detection model based on the confidence correction model during the application process, and further improving the accuracy of target detection.

[0080] In one embodiment, Figure 3 FIG. 2 is a flow chart of a training method for a confidence correction model provided in an embodiment of the present invention. Figure 3As shown, the above step 202 trains the initial confidence correction model of the target detection model under the prediction type for each prediction type output by each target detection model based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and the real detection box of the sample image corresponding to the prediction type, to obtain the confidence correction sub-model of the target detection model under the prediction type, which can be specifically achieved by the following steps:

[0081] Step 2021: for each predicted detection box under each prediction type output by the target detection model, determine a first intersection-over-union ratio between the predicted detection box and the corresponding true detection box.

[0082] For example, the first intersection-over-union ratio between the predicted detection box and the corresponding true detection box may be determined based on the following formula (3):

[0083] (3)

[0084] in, Indicates Predicted detection boxes, Indicates The real detection box corresponding to the predicted detection box, the two boxes represent the The predicted detection box and The real detection box corresponding to the predicted detection box, (Intersection over Union) means intersection over union.

[0085] Step 2022: Determine a loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence.

[0086] Optionally, a loss function under the prediction type is determined based on the following formula (1);

[0087] (1)

[0088] in, Indicates the prediction type, Indicates the prediction type The next The first intersection-over-union ratio between the predicted detection box and the corresponding real detection box, Indicates the prediction type The corresponding prediction confidence under Indicates the prediction type The number of predicted detection boxes, satisfy , represents a monotonically increasing constraint.

[0089] For example, the confidence between different target detection models is difficult to compare, but the confidence of a single target detection model still considers that the reliability of high confidence is higher than that of low confidence. The monotonic regression model is a statistical method for fitting monotonic functions. The output of the monotonic regression model increases as the input increases. The goal of the monotonic regression model is to minimize the loss function shown in the above formula (1).

[0090] Step 2023: iteratively adjust the model parameters of the initial confidence correction model under the prediction type based on the loss function to obtain a confidence correction sub-model of the target detection model under the prediction type.

[0091] For example, when the loss function shown in the above formula (1) is obtained, the model parameters of the initial confidence correction model under the prediction type are iteratively adjusted based on the loss function, in order to make the corrected confidence output by the confidence correction sub-model closer to the corresponding intersection-and-union ratio. The goal is to reduce the deviation between the corrected confidence predicted by the confidence correction sub-model and the corresponding intersection-and-union ratio, maintain the consistency of the corrected confidence and the corresponding intersection-and-union ratio, and finally obtain the confidence correction sub-model of the target detection model under the prediction type. Through each confidence correction sub-model, each target detection model can have a unified benchmark (intersection-and-union ratio) to compare confidence. The confidence correction sub-model finally obtained can be expressed by the following formula (4):

[0092] (4)

[0093] in, , and All are preset confidence levels. Indicates the confidence level of a detection box of a certain type output by the target detection model. When Input through function In the confidence correction submodel represented by Representation function Represents the corrected confidence of the confidence correction sub-model output; When Input through function In the confidence correction submodel represented by Representation function Represents the corrected confidence of the confidence correction sub-model output; When Input through function In the confidence correction submodel represented by Representation function represents the corrected confidence output by the confidence correction sub-model, Representation function The number of

[0094] It should be noted that in the above formula (4) , … It is obtained based on the confidence of the known detection frame, the known intersection-over-union ratio of the known detection frame and the corresponding real detection frame, and the present invention will not be repeated here.

[0095] In this embodiment, based on each first intersection-and-union ratio and each corresponding prediction confidence under the prediction type, a loss function under the prediction type is determined, and based on the loss function, the model parameters of the initial confidence correction model under the prediction type are iteratively adjusted to finally obtain a confidence correction sub-model of the target detection model under the prediction type, so that the corrected confidence output by the obtained confidence correction sub-model is closer to the intersection-and-union ratio determined based on the detection box and the corresponding real detection box, so as to improve the accuracy of target detection.

[0096] In one embodiment, Figure 4 FIG. 2 is a flow chart of a target detection method provided by an embodiment of the present invention. Figure 4 As shown, the above step 103 uses a soft non-maximum suppression algorithm based on the corrected confidence of each target detection model under the type to screen each detection frame to obtain a target detection frame, which can be specifically implemented by the following steps:

[0097] Step 1031: for a confidence set corresponding to all the corrected confidences of the type, determine a maximum corrected confidence of the type in the confidence set.

[0098] For example, when all corrected confidences of the type output by each confidence correction model are obtained, all corrected confidences of the type are combined to obtain a confidence set, and the maximum corrected confidence of the type is selected from the confidence set.

[0099] Step 1032: determine a second intersection-over-union ratio between the detection box corresponding to the maximum corrected confidence under the type and the detection boxes corresponding to other corrected confidences under the type; the other corrected confidences are the corrected confidences in the confidence set except the maximum corrected confidence.

[0100] For example, the detection frame corresponding to the maximum corrected confidence is used as the reference frame. For each corrected confidence in the confidence set except the maximum corrected confidence, the second intersection-over-union ratio between the detection frame corresponding to the maximum corrected confidence and the detection frames corresponding to other corrected confidences is calculated based on the above formula (3). The value of the second intersection-over-union ratio is between 0 and 1. The higher the value of the second intersection-over-union ratio, the higher the overlap.

[0101] Step 1033: for each of the second I / O ratios, update other corrected confidences of the type corresponding to the second I / O ratio based on the second I / O ratio and an attenuation parameter to obtain an updated confidence of the type corresponding to the second I / O ratio.

[0102] Optionally, determining an updated confidence level corresponding to the second intersection-over-union ratio based on the following formula (2);

[0103] (2)

[0104] in, Indicates the type, Representation Type Next Other corrected confidence levels, Indicates the type Next The detection boxes corresponding to other corrected confidences, Represents the confidence set The corrected confidence is the maximum corrected confidence, Indicates the type The detection box corresponding to the maximum corrected confidence level, represents the attenuation parameter, which is a parameter used to control the attenuation rate. The value of can be 0.5, Indicates the type Next The updated confidence corresponding to the other corrected confidences.

[0105] For example, for each second IoU, in the non-maximum suppression algorithm, when the second IoU is greater than a set threshold (e.g., 0.5), the detection frame corresponding to the other corrected confidences corresponding to the second IoU will be deleted. However, in the soft non-maximum suppression algorithm, the detection frames will not be directly removed, but the other corrected confidences of these detection frames will be attenuated, that is, the Gaussian attenuation formula shown in the above formula (2) will be used for attenuation, and finally the updated confidence of this type corresponding to each second IoU is obtained, thereby realizing the update of each other corrected confidence in the confidence set, even if the detection frame with a large second IoU is not removed.

[0106] Step 1034: Delete the detection boxes corresponding to the updated confidences that are less than the preset confidences, and obtain the remaining detection boxes of the type.

[0107] For example, when each updated confidence is obtained, each updated confidence is compared with a preset confidence. For example, the preset confidence is 0.3, and the detection boxes corresponding to the updated confidences that are less than the preset confidences are screened out. Since the detection boxes corresponding to the updated confidences that are less than the preset confidences cannot better represent the corresponding target objects, the detection boxes corresponding to the updated confidences that are less than the preset confidences are deleted to reduce the subsequent computational burden. In the same way, all detection boxes corresponding to the updated confidences that are less than the preset confidences can be deleted, and finally the remaining detection boxes of this type are obtained. The remaining detection boxes are detection boxes that can accurately represent the target objects. The number of remaining detection boxes can be one or two, etc.

[0108] Step 1035: determine the target detection frame of the type based on the remaining detection frames of the type.

[0109] For example, when the remaining detection frames of this type are obtained, if there is only one remaining detection frame, the remaining detection frame is directly determined as the target detection frame of this type; if there are more than two remaining detection frames, the detection frame corresponding to the maximum confidence among all the remaining detection frames can be determined as the target detection frame, or all the remaining detection frames can be determined as target detection frames.

[0110] In this embodiment, considering that the detection frames output by multiple target detection models have overlapping redundancy, the detection frames output by multiple target detection models are filtered by the soft non-maximum suppression algorithm corresponding to the above steps 1031 to 1035, and the overlapping detection frames are removed, and finally the target detection frames corresponding to the image to be detected are obtained. The soft non-maximum suppression algorithm has a certain improvement effect compared with the non-maximum suppression algorithm. The core idea of ​​the soft non-maximum suppression algorithm is: as the intersection and union ratio of the detection frame corresponding to other corrected confidences increases with the reference frame, the confidence of the detection frame corresponding to other corrected confidences gradually decays, rather than directly removing it, so that the valid detection frames with a certain overlap can be better retained, the missed detection can be reduced, and the accuracy of target detection can be further improved.

[0111] Figure 5 is a schematic diagram of the overall process of the target detection method provided by an embodiment of the present invention, such as Figure 5As shown, an image to be detected including at least one target object is input into multiple target detection models, and detection frames corresponding to each target object output by each target detection model and confidences representing the types of target objects included in each detection frame are obtained. For each type output by each target detection model, each confidence under the type is input into a confidence correction model corresponding to the target detection model, and each confidence under the type is corrected by a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type. Each detection frame is screened based on the corrected confidences of each target detection model under the type by a soft non-maximum suppression algorithm, and finally the detection result of the target object in the image to be detected is obtained.

[0112] Figure 6 FIG. 1 is a schematic diagram comparing target detection of a single target detection model and multiple target detection models provided by an embodiment of the present invention. Figure 6 As shown, taking two trained target detection models as an example, the detection results (detection boxes) of a single target detection model (Yolov5, Faster R-CNN) are compared with the detection results (detection boxes) of two trained target detection models (Yolov5+Faster R-CNN). It can be found that the detection results based on multiple target detection models proposed in the present invention can give full play to the advantages of each single target detection model and improve the accuracy of target detection. The target detection method based on multiple target detection models provided by the present invention can be applied to target detection operations of images taken by ordinary imaging devices such as digital cameras, and the beneficial effects achieved are also similar.

[0113] The target detection device provided by the present invention is described below. The target detection device described below and the target detection method described above can be referenced to each other.

[0114] Figure 7 is a schematic diagram of the structure of a target detection device provided by an embodiment of the present invention, such as Figure 7 As shown, the target detection device 700 includes a detection unit 701, a correction unit 702, a screening unit 703 and a determination unit 704; wherein:

[0115] A detection unit 701 is used to input an image to be detected including at least one target object into at least two target detection models to obtain target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes;

[0116] A correction unit 702 is used for inputting each confidence under each type output by each target detection model into a confidence correction model corresponding to the target detection model, and correcting each confidence under the type by a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0117] A screening unit 703 is used to screen each of the detection frames based on each corrected confidence of each of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0118] The determining unit 704 is configured to determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame.

[0119] The target detection device provided by the present invention inputs an image to be detected including at least one target object into at least two target detection models, obtains the detection frame and type confidence corresponding to each target object output by each target detection model, inputs each confidence under the type into the confidence correction model corresponding to the target detection model for each type output by each target detection model, corrects each confidence under the type by the confidence correction submodel corresponding to the type in the confidence correction model, obtains each corrected confidence under the type, screens each detection frame based on each corrected confidence under the type of each target detection model by the soft non-maximum suppression algorithm, obtains the target detection frame, and then determines the type corresponding to the target detection frame as the type of the target object included in the target detection frame. It can be seen that the present invention can screen the target detection frame from all detection frames through the confidence correction and soft non-maximum suppression algorithms, and determine the type corresponding to the target detection frame as the type of the corresponding target object, giving full play to the advantages of each target detection model, thereby improving the accuracy of target detection.

[0120] Based on any of the above embodiments, the confidence correction model is trained based on the following method:

[0121] Inputting a sample image including at least one sample object into each of the target detection models to obtain a prediction detection result output by each of the target detection models, wherein the prediction detection result includes a prediction detection box corresponding to each of the sample objects and a prediction confidence representing a prediction type of the sample object included in each of the prediction detection boxes;

[0122] For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type;

[0123] Based on the confidence correction sub-models of the target detection model under all the prediction types, a confidence correction model corresponding to the target detection model is determined.

[0124] Based on any of the foregoing embodiments, for each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type, including:

[0125] For each predicted detection box under each prediction type output by the target detection model, determine a first intersection-over-union ratio between the predicted detection box and the corresponding true detection box;

[0126] Determining a loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence;

[0127] Based on the loss function, the model parameters of the initial confidence correction model under the prediction type are iteratively adjusted to obtain the confidence correction sub-model of the target detection model under the prediction type.

[0128] Based on any of the foregoing embodiments, determining the loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence includes:

[0129] Determine the loss function under the prediction type based on the following formula (1);

[0130] (1)

[0131] in, Indicates the prediction type, Indicates the prediction type The next The first intersection-over-union ratio between the predicted detection box and the corresponding real detection box, Indicates the prediction type The corresponding prediction confidence under Indicates the prediction type The number of predicted detection boxes, satisfy , represents a monotonically increasing constraint.

[0132] Based on any of the above embodiments, the screening unit 703 is specifically used for:

[0133] For a confidence set corresponding to all the corrected confidences under the type, determine a maximum corrected confidence under the type in the confidence set;

[0134] Determine a second intersection-over-union ratio between a detection frame corresponding to a maximum corrected confidence under the type and a detection frame corresponding to other corrected confidences under the type; the other corrected confidences are corrected confidences in the confidence set except the maximum corrected confidence;

[0135] For each of the second I / O ratios, updating other corrected confidences of the type corresponding to the second I / O ratio based on the second I / O ratio and an attenuation parameter to obtain an updated confidence of the type corresponding to the second I / O ratio;

[0136] Deleting the detection boxes corresponding to the updated confidences that are less than the preset confidences, to obtain the remaining detection boxes of the type;

[0137] Determine an object detection frame of the type based on remaining detection frames of the type.

[0138] Based on any of the above embodiments, the screening unit 703 is further specifically configured to:

[0139] Determine an updated confidence level corresponding to the second intersection-over-union ratio based on the following formula (2);

[0140] (2)

[0141] in, Indicates the type, Representation Type Next Other corrected confidence levels, Indicates the type Next The detection boxes corresponding to other corrected confidences, Represents the confidence set The corrected confidence is the maximum corrected confidence, Indicates the type The detection box corresponding to the maximum corrected confidence level, represents the attenuation parameter, Indicates the type Next The updated confidence corresponding to the other corrected confidences.

[0142] Figure 8 is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention, such as Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the target detection method, which includes: inputting an image to be detected including at least one target object into at least two target detection models, obtaining a target detection result output by each of the target detection models, wherein the target detection result includes a detection frame corresponding to each of the target objects and a confidence level characterizing the type of the target object included in each of the detection frames;

[0143] For each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type by using a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0144] The detection frames are screened based on the corrected confidences of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0145] The type corresponding to the target detection frame is determined as the type of the target object included in the target detection frame.

[0146] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0147] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the target detection method provided by the above methods, the method comprising: inputting an image to be detected including at least one target object into at least two target detection models, obtaining a target detection result output by each of the target detection models, the target detection result including a detection box corresponding to each of the target objects and a confidence level characterizing the type of the target object included in each of the detection boxes;

[0148] For each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type by using a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0149] The detection frames are screened based on the corrected confidences of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0150] The type corresponding to the target detection frame is determined as the type of the target object included in the target detection frame.

[0151] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the target detection method provided by the above methods, the method comprising: inputting an image to be detected including at least one target object into at least two target detection models, obtaining a target detection result output by each of the target detection models, the target detection result comprising a detection box corresponding to each of the target objects and a confidence level characterizing the type of the target object included in each of the detection boxes;

[0152] For each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type by using a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type;

[0153] The detection frames are screened based on the corrected confidences of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame;

[0154] The type corresponding to the target detection frame is determined as the type of the target object included in the target detection frame.

[0155] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0156] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target detection method, characterized in that: include: Inputting an image to be detected including at least one target object into at least two target detection models, and obtaining target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes; For each type output by each target detection model, input each confidence under the type into a confidence correction model corresponding to the target detection model, and correct each confidence under the type by using a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type; The detection frames are screened based on the corrected confidences of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame; Determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame; The confidence correction model is trained based on the following method: Inputting a sample image including at least one sample object into each of the target detection models to obtain a prediction detection result output by each of the target detection models, wherein the prediction detection result includes a prediction detection box corresponding to each of the sample objects and a prediction confidence representing a prediction type of the sample object included in each of the prediction detection boxes; For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type; Based on the confidence correction sub-models of the target detection model under all the prediction types, a confidence correction model corresponding to the target detection model is determined.

2. The target detection method according to claim 1, characterized in that: For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type, including: For each predicted detection box under each prediction type output by the target detection model, determine a first intersection-over-union ratio between the predicted detection box and the corresponding true detection box; Determining a loss function under the prediction type based on each of the first intersection-over-union ratios under the prediction type and each corresponding prediction confidence; Based on the loss function, the model parameters of the initial confidence correction model under the prediction type are iteratively adjusted to obtain the confidence correction sub-model of the target detection model under the prediction type.

3. The target detection method according to claim 2, characterized in that: The determining, based on each of the first intersection-over-union ratios and each corresponding prediction confidence under the prediction type, a loss function under the prediction type, comprises: Determine the loss function under the prediction type based on the following formula (1); (1); in, Indicates the prediction type, Indicates the prediction type The next The first intersection-over-union ratio between the predicted detection box and the corresponding real detection box, Indicates the prediction type The corresponding prediction confidence under Indicates the prediction type The number of predicted detection boxes, satisfy , represents a monotonically increasing constraint.

4. The target detection method according to claim 1, characterized in that: The method of screening each of the detection frames based on each corrected confidence of each of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame includes: For a confidence set corresponding to all the corrected confidences under the type, determine a maximum corrected confidence under the type in the confidence set; Determine a second intersection-over-union ratio between a detection frame corresponding to a maximum corrected confidence under the type and a detection frame corresponding to other corrected confidences under the type; the other corrected confidences are corrected confidences in the confidence set except the maximum corrected confidence; For each of the second I / O ratios, updating other corrected confidences of the type corresponding to the second I / O ratio based on the second I / O ratio and an attenuation parameter to obtain an updated confidence of the type corresponding to the second I / O ratio; Deleting the detection boxes corresponding to the updated confidences that are less than the preset confidences, to obtain the remaining detection boxes of the type; Determine an object detection frame of the type based on remaining detection frames of the type.

5. The target detection method according to claim 4, characterized in that: The updating of other corrected confidences corresponding to the second IoU ratio based on the second IoU ratio and the attenuation parameter to obtain an updated confidence corresponding to the second IoU ratio includes: Determine an updated confidence level corresponding to the second intersection-over-union ratio based on the following formula (2); (2); in, Indicates the type, Representation Type Next Other corrected confidence levels, Indicates the type Next The detection boxes corresponding to other corrected confidences, Represents the confidence set The corrected confidence is the maximum corrected confidence, Indicates the type The detection box corresponding to the maximum corrected confidence level is: represents the attenuation parameter, Indicates the type Next The updated confidence corresponding to the other corrected confidences.

6. A target detection device, characterized in that: include: A detection unit, configured to input an image to be detected including at least one target object into at least two target detection models, and obtain target detection results output by each of the target detection models, wherein the target detection results include detection boxes corresponding to each of the target objects and confidence levels representing the types of the target objects included in each of the detection boxes; A correction unit, for each type output by each target detection model, inputting each confidence under the type into a confidence correction model corresponding to the target detection model, and correcting each confidence under the type by a confidence correction sub-model corresponding to the type in the confidence correction model to obtain each corrected confidence under the type; A screening unit, configured to screen each of the detection frames based on each corrected confidence of each of the target detection models under the type by using a soft non-maximum suppression algorithm to obtain a target detection frame; a determining unit, configured to determine the type corresponding to the target detection frame as the type of the target object included in the target detection frame; The confidence correction model is trained based on the following method: Inputting a sample image including at least one sample object into each of the target detection models to obtain a prediction detection result output by each of the target detection models, wherein the prediction detection result includes a prediction detection box corresponding to each of the sample objects and a prediction confidence representing a prediction type of the sample object included in each of the prediction detection boxes; For each prediction type output by each target detection model, based on each prediction detection box under the prediction type, each prediction confidence under the prediction type, and a real detection box of a sample image corresponding to the prediction type, an initial confidence correction model of the target detection model under the prediction type is trained to obtain a confidence correction sub-model of the target detection model under the prediction type; Based on the confidence correction sub-models of the target detection model under all the prediction types, a confidence correction model corresponding to the target detection model is determined.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the target detection method according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method according to any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the target detection method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Detection model training method and device, computing equipment and storage medium

    CN116977765A

  • Target detection method and device, visual detection system and electronic equipment

    CN117788798A