Target detection model training method and device, target detection method and device and electronic equipment

By copying the weight parameters in the target detection model and obtaining the false detection sample information, the problem of high false detection rate of the target detection model in intelligent traffic monitoring is solved, and the effect of reducing the false detection rate and improving the generalization of the model is achieved.

CN119963967APending Publication Date: 2025-05-09JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311493938.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In intelligent traffic monitoring scenarios, the target detection model has a high false detection rate due to the unbalanced proportion of positive and negative samples. Especially in long-distance target detection, false detection of human targets, branches, guardrails, etc. is prone to occur.

Method used

By copying the weight parameters of the object detection model to be trained into a detection model with the same structure as its structure, object detection is performed on the sample image, false detection sample information is obtained, and training and update the weight parameters of the object detection model to be trained based on this information to suppress the prediction information of the false detection prediction box.

Benefits of technology

The false detection rate of the target detection model is reduced, the generalization and robustness of the model is improved, especially in intelligent traffic monitoring scenarios, which effectively reduces false detection of long-distance targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963967A_ABST
    Figure CN119963967A_ABST
Patent Text Reader

Abstract

The invention provides a target detection model training method and device, a target detection method and device and electronic equipment, and relates to the technical field of computer vision, and the method comprises the steps: copying a weight parameter of a to-be-trained target detection model to a detection model having the same structure as the to-be-trained target detection model; performing target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image, and determining false detection sample information of the sample image based on the first target prediction frame and a target annotation frame of the sample image; and training and updating the weight parameters of the to-be-trained target detection model based on the sample image and the false detection sample information to obtain a trained target detection model. According to the technical scheme provided by the invention, the false detection rate of the target detection model can be reduced, and the generalization of the target detection model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a target detection model training method, a target detection method and device, and an electronic device. Background Art

[0002] Object detection is an important application in the field of computer vision and is directly or indirectly used in various industries, such as detecting faces, vehicles, or buildings from images.

[0003] Intelligent traffic monitoring is to automatically analyze image sequences through computers without human intervention to detect, track and analyze the behavior of moving targets in dynamic scenes. In intelligent traffic monitoring, the application scenarios of multi-target detection such as pedestrians, non-motor vehicles, motor vehicles, faces, license plates, etc. are relatively common, which also puts higher requirements on the accuracy and efficiency of target detection.

[0004] With the development of computer vision and artificial intelligence technology, the target detection model based on neural network has become the mainstream framework for target detection. In related technologies, the sample image is input into the target detection model for target prediction, and the prediction result is compared with the label data to determine the loss function value. Then, based on the loss function value, the weight parameters of the target detection model are adjusted using the back propagation algorithm until the loss function value reaches the expected value, completing the training of the target detection model. However, in the intelligent traffic monitoring scenario, the ratio of positive and negative samples used to train the target detection model is not balanced. The number of negative samples in the sample image is much larger than the number of positive samples, which makes the target detection model prone to false detection in actual application scenarios. For example, in some long-distance target detection, human targets, tree branches, guardrails, etc. are prone to false detection. Summary of the invention

[0005] The present invention provides a target detection model training method, a target detection method and device, and an electronic device, which are used to solve the false detection problem of the target detection model in the prior art, reduce the false detection rate of the target detection model, and improve the generalization of the model.

[0006] The present invention provides a target detection model training method, comprising:

[0007] Obtaining weight parameters of a target detection model to be trained, and copying the weight parameters into a detection model; the detection model has the same structure as the target detection model to be trained;

[0008] Performing target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image, and determining misdetected sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image;

[0009] The weight parameters of the target detection model to be trained are trained and updated based on the sample image and the false detection sample information to obtain a trained target detection model.

[0010] According to a target detection model training method provided by the present invention, the weight parameters of the target detection model to be trained are trained and updated based on the sample image and the false detection sample information, including:

[0011] Inputting the sample image into the target detection model to be trained to obtain a second target prediction frame output by the target detection model to be trained;

[0012] For each second prediction box in the second target prediction box, when the second prediction box is determined to be a falsely detected prediction box based on the falsely detected sample information, determining a first loss value of a first loss function, and training and updating the weight parameters of the target detection model to be trained based on the first loss value;

[0013] In a case where it is determined that the second prediction box is a non-falsely detected prediction box based on the falsely detected sample information, determining a second loss value of a second loss function, and training and updating the weight parameters of the target detection model to be trained based on the second loss value;

[0014] Among them, the first loss function is used to suppress the prediction information of the second prediction box; the first loss function is obtained by performing step-by-step gradient weighting on the second loss function.

[0015] According to a target detection model training method provided by the present invention, the second loss function is determined based on a target confidence loss function, a classification loss function and a regression loss function; and performing step-by-step gradient weighting on the second loss function includes:

[0016] Obtaining a first maximum intersection-over-union ratio between the false detection prediction frame and each false detection sample frame in the false detection sample information;

[0017] Determine a first step weight based on the first maximum intersection-over-union ratio and a first preset intersection-over-union ratio threshold;

[0018] Determine a second hierarchical weight based on the first maximum intersection-over-union ratio, the first preset intersection-over-union ratio threshold, and the total number of target categories;

[0019] The first order weight is added to the target confidence loss function in the second loss function, and the second order weight is added to the classification loss function in the second loss function to obtain the first loss function.

[0020] According to a target detection model training method provided by the present invention, the determining, based on the false detection sample information, that the second prediction box is a false detection prediction box includes:

[0021] Determine a second maximum intersection-over-union ratio between the second prediction box and a non-falsely detected prediction box in the sample image;

[0022] When the second maximum intersection-over-union ratio is less than a third preset intersection-over-union ratio threshold, determining a first maximum intersection-over-union ratio between the second prediction frame and each falsely detected sample frame in the falsely detected sample information;

[0023] When the first maximum IoU is greater than a fourth preset IoU threshold, the second prediction frame is determined as the false detection prediction frame.

[0024] According to a target detection model training method provided by the present invention, target detection is performed on a sample image based on the detection model to obtain a first target prediction frame of the sample image, including:

[0025] Inputting the sample image into the detection model to obtain target prediction frames at at least two scales output by the detection model, and merging the target prediction frames at at least two scales to obtain a third target prediction frame;

[0026] Perform non-maximum suppression processing on the third target prediction box to obtain the first target prediction box.

[0027] According to a target detection model training method provided by the present invention, the determining of misdetected sample information of the sample image based on the first target prediction box and the target annotation box of the sample image includes:

[0028] For each first prediction box in the first target prediction box, determining a first intersection-over-union ratio between the first prediction box and each target annotation box of the sample image;

[0029] When the first intersection-over-union ratios are all smaller than a second preset intersection-over-union ratio threshold, the information of the first prediction frame is determined as misdetected sample information of the sample image.

[0030] The present invention also provides a target detection method, comprising:

[0031] Acquire the surveillance image to be detected;

[0032] The monitoring image to be detected is input into a target detection model to obtain a target detection result of the monitoring image to be detected output by the target detection model; wherein the target detection model is trained based on any of the target detection model training methods described above.

[0033] The present invention also provides a target detection model training device, comprising:

[0034] A parameter copying module, used for obtaining weight parameters of a target detection model to be trained, and copying the weight parameters to a detection model; the detection model has the same structure as the target detection model to be trained;

[0035] A first target detection module, used to perform target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image;

[0036] a false detection determination module, configured to determine false detection sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image;

[0037] The model training module is used to train and update the weight parameters of the target detection model to be trained based on the sample image and the false detection sample information to obtain a trained target detection model.

[0038] The present invention also provides a target detection device, comprising:

[0039] An image acquisition module is used to acquire the monitoring image to be detected;

[0040] The second target detection module is used to input the monitoring image to be detected into a target detection model to obtain the target detection result in the monitoring image to be detected output by the target detection model; wherein the target detection model is trained based on any of the target detection model training methods described above.

[0041] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a target detection model training method as described in any one of the above, or implements the target detection method as described in the above.

[0042] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the target detection model training methods described above, or implements the target detection method described above.

[0043] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the target detection model training methods described above, or implements the target detection method described above.

[0044] The target detection model training method, target detection method and device, and electronic device provided by the present invention can obtain a first target prediction frame of the sample image by copying the weight parameters of the target detection model to be trained into a detection model with the same structure as the target detection model to be trained, and performing target detection on the sample image based on the detection model, wherein the first target prediction frame is the target prediction frame that can be predicted by the target detection model to be trained when performing target detection on the sample image under the current weight parameters; then, based on the first target prediction frame and the target annotation frame of the sample image, the false detection sample information of the sample image is determined, thereby simulating the test process of the target detection model to be trained under the current weight parameters, and obtaining the false detection samples detected by the current target detection model to be trained; then, based on the sample image and the false detection sample information, the weight parameters of the target detection model to be trained are trained and updated to obtain a trained target detection model, so that targeted reinforcement learning can be performed on the false detection samples in the sample image that will be falsely detected by the current target detection model to be trained, and false detections are effectively suppressed by reinforcement learning of false detection samples during the training process, thereby reducing the false detection rate of the finally trained target detection model and improving the generalization of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0046] Figure 1 This is one of the flowcharts of the target detection model training method provided by an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of the principle of obtaining misdetected sample information in a sample image based on a detection model in an embodiment of the present invention;

[0048] Figure 3 is a flow chart of a method for performing reinforcement learning on a second target prediction frame based on misdetected sample information in an embodiment of the present invention;

[0049] Figure 4 This is a second flow chart of a method for training a target detection model provided by an embodiment of the present invention;

[0050] Figure 5 is a flow chart of a target detection method provided by an embodiment of the present invention;

[0051] Figure 6 is a schematic diagram of the structure of a target detection model training device provided by an embodiment of the present invention;

[0052] Figure 7 is a schematic diagram of the structure of a target detection device provided by an embodiment of the present invention;

[0053] Figure 8 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] It should be noted that the serial numbers assigned to the objects described in the present invention, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning.

[0056] Target detection is a technology that uses computer vision to find specific targets in image sequences and determine their locations, such as the YOLO series of target detection algorithms. At present, the target detection model in most scenarios uses a single-stage detection model, that is, when training the target detection model, the input sample image can be predicted in multiple scales and grids, and the prediction results can be compared with the label data to determine the loss value of the loss function, such as performing at least one of the target loss calculation, regression loss calculation, and category loss calculation. Based on the calculated loss value, the model parameters of the target detection model are adjusted through back propagation until the loss value reaches the set requirements and the training is terminated. At this point, a trained target detection module can be obtained. When using the target detection model to detect targets in the surveillance image to be detected, the prediction box output by the model can be restored and then filtered by non-maximum suppression (NMS) to obtain the final target detection result.

[0057] False detection is the focus and difficulty of target detection tasks. For some large-resolution surveillance images, the existence of false detection samples is difficult to avoid, and such false detection samples are associated with some difficult samples in the training set. The improvement of the loss function may indirectly affect the learning of difficult samples, resulting in poor results of model optimization training. Among them, difficult samples are negative samples in the sample image that are similar to the target. For example, in the intelligent traffic monitoring scenario, some similar objects such as distant branches and guardrails are easily misdetected as human targets. In the training process of the target detection model, the sample images are only labeled with positive samples such as human bodies, non-motor vehicles, motor vehicles, faces, and heads. The model network does not discriminate the false detection areas during training. If it is forcibly adjusted through the loss function, it may affect the learning of difficult samples, thereby affecting the detection effect of the model.

[0058] Since the biggest difference between target detection tasks and algorithms such as target classification and target segmentation is that the training and testing processes are not completely consistent, that is, there is no post-processing process in the training process of the target detection model, therefore, the analysis of the model output results during the model training process is not completely consistent with the testing process. Based on this, the embodiment of the present invention combines the situation that the target detection model is prone to false detection of difficult samples in the later stage of training iteration, introduces the current optimal target detection model as a pre-training model into the secondary forward reinforcement learning strategy for a new round of iterative optimization, and proposes a target detection model training method based on the secondary forward reinforcement learning mechanism. Starting from the multi-target detection task, it can effectively suppress false detection through reinforcement learning of false detection samples, thereby reducing the false detection rate of the target detection model and enhancing the robustness and generalization of the model.

[0059] Combine the following Figure 1-Figure 4 The target detection model training method of the present invention is described. The target detection model training method can be applied to electronic devices such as terminal devices or servers. Among them, the terminal device may include a mobile phone, a computer, a tablet computer, a wearable device, etc.; the server may include an independent server, a cluster server or a cloud server, etc. The target detection model training method can also be applied to a target detection model training device set in an electronic device such as a terminal device or a server, and the target detection model training device can be implemented by software, hardware or a combination of both.

[0060] Figure 1 One of the flow charts of the target detection model training method provided by an embodiment of the present invention is exemplarily shown, referring to Figure 1 As shown, the target detection model training method may include the following steps 110 to 130.

[0061] Step 110: Obtain weight parameters of the target detection model to be trained, and copy the weight parameters to the detection model.

[0062] Among them, the target detection model to be trained can be a model that already has target detection capabilities after multiple rounds of model iterative training, and the training of the target detection model to be trained can be understood as further optimization training of the target detection model to be trained that already has target detection capabilities.

[0063] The detection model has the same structure as the target detection model to be trained. After the weight parameters of the target detection model to be trained are copied to the detection model, the detection model has the same target detection capability as the current target detection model to be trained.

[0064] For example, the target detection model to be trained may include a target detection deep learning model of the YOLO series structure, a target detection model of the single-stage multi-box detection (Single Shot Multibox Detector, SSD) network structure, etc. For example, in an exemplary embodiment of the present invention, the CSP-Darknet53 network structure in the YOLO series model can be used as the backbone network basis, and the DropBlock layer is introduced in the deep residual part of the model to improve the generalization ability of the model. At the same time, the Neck layer adopts the SPP_PAN structure, and the last Head layer adopts the Head layer of YOLOv3. Each layer of the Head can set 3 anchors as reference frames when predicting the target detection frame, thereby forming the network structure of the target detection model to be trained.

[0065] Step 120: Perform target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image, and determine misdetected sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image.

[0066] The detection model has the same target detection capability as the current target detection model to be trained. The sample image can be input into the detection model, and the detection model can be used to perform target detection on the sample image to obtain the first target prediction box output by the detection model. The first target prediction box is the target prediction box that the target detection model to be trained can predict when performing target detection on the sample image under the current weight parameters.

[0067] It takes multiple epochs to complete the training of a model, and it will converge after multiple rounds of iterative training. Among them, one epoch is the process of training once using all the samples in the training set. In an exemplary embodiment of the present invention, a large monitoring image in an intelligent traffic monitoring scenario can be used as a sample image for training a target detection model, and the targets such as human bodies, non-motor vehicles, motor vehicles, faces, and heads in the sample images are annotated to obtain a text label corresponding to each sample image, and the text label includes label data for each annotated target, and the label data may include the classification category and coordinate position of the target, wherein the classification category may be marked with numbers, such as the five different categories of targets, such as human bodies, non-motor vehicles, motor vehicles, faces, and heads, can be marked and distinguished with 1, 2, 3, 4, and 5, respectively.

[0068] For example, the sample images may be enhanced to expand the diversity of the training samples, thereby enhancing the robustness of the model. The image enhancement process may include at least one of horizontal flipping, mixing (MixUp), mosaic (Mosaic) processing, random horizontal flipping, HSV random enhancement, etc. HSV represents color space, H represents hue, S represents saturation, and V represents brightness.

[0069] In one epoch training process, the training set data can be divided into multiple batches of data, and a batch of data is fed in for training each time. Specifically, a part of the samples in the training set is used to perform a back propagation update on the weight parameters of the model. This part of the samples is a batch of data, and the size of a batch of data is the batch size.

[0070] For example, the current batch of sample images can be input into the detection model, and for each sample image in the current batch of sample images, the detection model can perform target detection on the sample image, and predict and output the first target prediction box corresponding to the sample image. For example, the detection model can perform prediction outputs on the sample image based on the anchor at least two different scales, and then integrate the prediction outputs at at least two different scales to obtain the first target prediction box corresponding to the sample image. Accordingly, performing target detection on the sample image based on the detection model to obtain the first target prediction box of the sample image can include: inputting the sample image into the detection model to obtain the target prediction box output by the detection model at at least two scales; merging the target prediction boxes at at least two scales to obtain the first target prediction box of the sample image.

[0071] Alternatively, the integrated prediction frame can be further subjected to NMS processing to obtain a first target prediction frame corresponding to the sample image, and redundant prediction frames can be removed by NMS processing. Specifically, performing target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image can include: inputting the sample image into the detection model to obtain a target prediction frame output by the detection model at at least two scales, and merging the target prediction frames at at least two scales to obtain a third target prediction frame; performing non-maximum suppression processing on the third target prediction frame to obtain a first target prediction frame of the sample image.

[0072] By way of example, determining the falsely detected sample information of a sample image based on a first target prediction box and a target annotation box of the sample image may include: determining, for each first prediction box in the first target prediction box, a first intersection-and-union ratio between the first prediction box and each target annotation box of the sample image; and determining the information of the first prediction box as the falsely detected sample information of the sample image when the first intersection-and-union ratios are all less than a second preset intersection-and-union ratio threshold.

[0073] Alternatively, when the first intersection-over-union ratio is less than the second preset intersection-over-union ratio threshold, it can be determined whether the confidence of the first prediction box is less than the preset confidence threshold. If it is, the information of the first prediction box is determined as the false detection sample information of the sample image. The preset confidence threshold can take a larger value, such as 0.9 or 0.95. By increasing the judgment of confidence, it can be prevented that the positive sample in the sample image is missed and the positive sample is judged as a false detection sample.

[0074] Specifically, for each sample image in the current batch of sample images, after obtaining the first target prediction box corresponding to the sample image, for each first prediction box in the first target prediction box, the first prediction box can be compared and analyzed with each target annotation box corresponding to the sample image, such as performing intersection over union (IOU) calculation. If the IOU values ​​are all less than the set threshold, the first prediction box is considered to be a false positive box, and the information of the false positive box can be determined as the false positive sample information of the sample image. The false positive sample information may include the image path information corresponding to the false positive sample box, the falsely predicted target category, and the falsely predicted coordinate information. The image path information can be used to identify the sample image, and the falsely predicted coordinate information can characterize the position of the falsely detected sample box in the sample image.

[0075] In this way, the sample image is detected by the detection model, and the false detection samples are judged based on the target annotation box of the sample image, and the false detection sample information is obtained. The purpose of adding implicit annotations to the false detection samples can be achieved during the training of the target detection model, which is convenient for the subsequent reinforcement learning of the false detection samples.

[0076] Step 130: Based on the sample images and the false detection sample information, the weight parameters of the target detection model to be trained are updated to obtain a trained target detection model.

[0077] After obtaining the false detection sample information corresponding to each sample image in the current batch of sample images, the current batch of sample images can be input into the target detection model to be trained, and the target detection model to be trained performs target detection on each sample image in the current batch of sample images to obtain the second target prediction box corresponding to each sample image output by the target detection model to be trained. For example, the target detection model to be trained can perform prediction outputs on the sample images at at least two different scales based on the anchors to obtain the prediction outputs of the sample images at different scales, and these prediction outputs can be determined as the second target prediction box corresponding to the sample image.

[0078] After obtaining the second target prediction box, the false detection prediction box in the second target prediction box can be determined based on the false detection sample information, and the weight parameters of the target detection model to be trained are updated with the goal of suppressing the prediction information of the false detection prediction box. The secondary forward processing of the current batch of sample images is completed, and the training of the next batch of sample images is entered until the training end conditions are met to obtain a trained target detection model.

[0079] In this way, in the process of training the target detection model to be trained using the current batch of sample images, the false detection prediction box of the target detection model to be trained when performing target detection on the current batch of sample images can be obtained, and then the false detection prediction box can be reinforced learning to effectively suppress false detection and enhance the robustness and generalization of the model.

[0080] It is understandable that a false positive sample frame requires multiple rounds of iterative reinforcement learning to learn well. If the weight parameters of the detection model are constantly updated, the false positive sample frame will be in a state of change, that is, during the secondary forward processing, the false positive prediction frame that needs to be suppressed in the sample image will also be in a state of change, which will cause the optimization effect of the target detection model to be trained to be uncontrollable. In addition, consider extreme cases, for example, after two rounds of iterative training, some false positive sample frames temporarily disappear due to reduced confidence, while new false positive sample frames appear in other areas, that is, during the secondary forward processing, the target detection model to be trained will predict a new false positive prediction frame, and the target detection model to be trained will suppress the new false positive prediction frame, but in the next few rounds of training, the initial false positive prediction frame appears again, but the learning rate at this time is too small, and it may not be possible to suppress the learning of these false positive prediction frames, which ultimately leads to unsatisfactory optimization of the target detection model to be trained. Therefore, after the first forward detection model copies the weight parameters of the target detection model to be trained, the weight parameters are not updated during the entire training process of the target detection model to be trained, and the false detection sample frame to be suppressed in learning is also the result of actual testing of the current target detection model to be trained.

[0081] The target detection model training method provided by the embodiment of the present invention can obtain a first target prediction frame of the sample image by copying the weight parameters of the target detection model to be trained into a detection model with the same structure as the target detection model to be trained, and performing target detection on the sample image based on the detection model, wherein the first target prediction frame is the target prediction frame that can be predicted by the target detection model to be trained when performing target detection on the sample image under the current weight parameters; then, the false detection sample information of the sample image is determined based on the first target prediction frame and the target annotation frame of the sample image, thereby simulating the test process of the target detection model to be trained under the current weight parameters, and obtaining the false detection samples detected by the current target detection model to be trained; then, the weight parameters of the target detection model to be trained are trained and updated based on the sample image and the false detection sample information, and a trained target detection model is obtained. In this way, targeted reinforcement learning can be performed on the false detection samples in the sample image that will be falsely detected by the current target detection model to be trained, and false detections can be effectively suppressed by reinforcement learning of false detection samples during the training process, thereby reducing the false detection rate of the finally trained target detection model and improving the generalization of the model.

[0082] based on Figure 1 The target detection model training method of the corresponding embodiment, Figure 2 The schematic diagram of the principle of obtaining the false detection sample information in the sample image based on the detection model is exemplified, which can be considered as the process of extracting the false detection sample information during the first forward processing of the sample image. Figure 2As shown in the figure, taking the prediction of the detection model at three different scales as an example, each sample image in the current batch of sample images can be input into the detection model. For each sample image, the detection model can extract feature maps at three different scales to obtain feature maps at three different scales. For each scale, the feature map at the scale is confidence filtered and multiple prediction boxes are restored according to the anchor of the Head layer. Then, the prediction boxes at three different scales are merged, and the merged prediction boxes are processed by NMS to obtain the prediction output result of the sample image, that is, the prediction box. Afterwards, each prediction box predicted by each sample image in the current batch of sample images can be traversed, and the IOU calculation of the prediction box and the target annotation box of the corresponding sample image can be performed, and the calculated IOU value can be compared with the preset IOU threshold. If it is less than the preset IOU threshold, the prediction box can be determined to be a false positive sample box, and the image path information, incorrectly predicted target category and incorrectly predicted coordinate information corresponding to the false positive sample box can be determined as false positive sample information and saved; or, if the calculated IOU value is less than the preset IOU threshold, it can be further determined whether the confidence of the prediction box is less than the preset confidence threshold. If it is less than the preset confidence threshold, the prediction box can be determined to be a false positive sample box, and the false positive sample information corresponding to the false positive sample box is saved. If it is greater than or equal to the preset confidence threshold, the prediction box may be a non-false positive sample box, that is, a positive sample. In this way, when the IOU value is less than the preset IOU threshold, further confidence judgment can prevent the positive sample in the sample image from being missed and causing the positive sample to be judged as a false positive sample, thereby ensuring the accuracy of the false positive sample information obtained by the first forward processing. The preset confidence threshold is a high confidence threshold, which can be set to 0.8 to 0.95, for example.

[0083] If the calculated IOU value is less than or equal to the preset IOU threshold, and / or the confidence of the predicted box is greater than or equal to the preset confidence threshold, the predicted box can be determined to be a non-false positive sample box, that is, a positive sample.

[0084] For example, a global variable may be used to store misdetected sample information predicted for each sample image in the current batch of sample images.

[0085] At this point, the first forward processing process for the current batch of sample images is completed. Through the first forward processing, the false detection sample information in the current batch of sample images can be extracted.

[0086] based on Figure 1The target detection model training method of the corresponding embodiment, in an exemplary embodiment, after using the detection model to obtain the false detection sample information corresponding to each sample image in the current batch of sample images, the current batch of sample images can be input into the target detection model to be trained, and secondary forward learning training is performed based on the obtained false detection sample information. During the training process, the loss function of the false detection prediction box output by the target detection model to be trained is subjected to step-by-step gradient weighting to suppress the prediction information of the false detection prediction box, strengthen the learning of the false detection samples, and thereby reduce the false detection rate of the model.

[0087] Specifically, training and updating the weight parameters of the target detection model to be trained based on the sample image and the false detection sample information may include: inputting the sample image into the target detection model to be trained to obtain a second target prediction box output by the target detection model to be trained; for each second prediction box in the second target prediction box, when the second prediction box is determined to be a false detection prediction box based on the false detection sample information, determining a first loss value of a first loss function, and training and updating the weight parameters of the target detection model to be trained based on the first loss value; when the second prediction box is determined to be a non-false detection prediction box based on the false detection sample information, determining a second loss value of a second loss function, and training and updating the weight parameters of the target detection model to be trained based on the second loss value; wherein the first loss function is used to suppress the prediction information of the second prediction box; and the first loss function is obtained by performing step-by-step gradient weighting on the second loss function.

[0088] For example, determining that the second prediction box is a false detection prediction box based on the false detection sample information may include: determining a first maximum intersection-and-union ratio between the second prediction box and each false detection sample box in the false detection sample information; and determining the second prediction box as a false detection prediction box when the first maximum intersection-and-union ratio is greater than a fourth preset intersection-and-union ratio threshold.

[0089] For example, it can be determined first whether the overlapping area between the second prediction box and the non-falsely detected prediction box in the sample image reaches a preset ratio requirement, that is, it can be determined whether the overlapping area between the second prediction box and the prediction box of the positive sample in the sample image reaches the preset ratio requirement. If so, it can be considered that the second prediction box has a large overlapping area with the positive sample and is treated as a non-falsely detected prediction box. Otherwise, the intersection and union ratio of the second prediction box and each falsely detected sample box in the falsely detected sample information is used to further determine whether the second prediction box is a falsely detected prediction box.

[0090] Specifically, determining that the second prediction frame is a false detection prediction frame based on the false detection sample information may include: determining a second maximum intersection-and-union ratio between the second prediction frame and the non-false detection prediction frame in the sample image; when the second maximum intersection-and-union ratio is less than a third preset intersection-and-union ratio threshold, determining the first maximum intersection-and-union ratio between the second prediction frame and each false detection sample frame in the false detection sample information; when the first maximum intersection-and-union ratio is greater than a fourth preset intersection-and-union ratio threshold, determining the second prediction frame as a false detection prediction frame.

[0091] For example, the current batch of sample images can be input again into the target detection model to be trained with the same network structure as the detection model for forward reasoning, and prediction results at at least two different scales are obtained, that is, the second target prediction box is obtained. For each prediction result of each sample image at different scales, that is, each second prediction box in the second target prediction box, the second prediction box is subjected to IOU traversal calculation with the false detection sample box in the stored false detection sample information, and the maximum IOU value is selected from the calculated IOU values. If the maximum IOU value is greater than the fourth preset intersection-and-union ratio threshold (for example, 0.8), the second prediction box can be determined to be a false detection prediction box, that is, the second prediction box will be tested as a false detection prediction box area in the test phase. Then, the loss function of the false detection prediction box can be weighted by step gradient, the target predicted by the false detection prediction box can be strengthened, and the loss value of the loss function after step gradient weighting can be calculated. If the maximum IOU value is less than or equal to the fourth preset intersection-and-union ratio threshold, the second prediction box can be determined to be a non-false detection prediction box, and the loss value of the loss function of the non-false detection prediction box can be directly calculated. Finally, the gradient result of the second target prediction box output when the current batch of sample images is subjected to secondary forward reasoning can be calculated, and then the gradient backpropagation is performed to update the network weight parameters of the target detection model to be trained, and the training of the current batch of sample images is completed, and the stored false positive sample information is cleared. In this way, the secondary forward training reasoning of the current batch of sample images is completed, and the training of the next batch of sample images is entered until the target detection model to be trained reaches the training end condition, and a trained target detection model is obtained.

[0092] For a prediction box, the weight distribution of its feature information from the inside to the outside decreases in sequence, even if the center point of the prediction box is not on the target object. The closer to the center of the prediction box, the more important the feature information. Based on this, in an exemplary embodiment of the present invention, different hierarchical weights can be set for the target confidence loss function and classification loss function of the false detection prediction box based on the IOU threshold of the false detection prediction box and the false detection sample box, respectively, wherein the hierarchical weights are proportional to the IOU values ​​of the false detection prediction box and the false detection sample box. In this way, when the overlapping area between the false detection prediction box and the false detection sample box is larger, the corresponding suppression weight also increases rapidly, and finally a sample suppression learning range with gradually decreasing weights from the inside to the outside is formed around the false detection sample box acquired forward for the first time, so that the purpose of suppressing the prediction information of the false detection sample box and the surrounding area can be achieved by performing hierarchical weighting according to the false detection sample box.

[0093] Based on this, in an example embodiment, the second loss function can be determined based on the target confidence loss function, the classification loss function and the regression loss function; accordingly, performing step-by-step gradient weighting on the second loss function can include: obtaining the first maximum intersection-and-union ratio between the false detection prediction box and each false detection sample box in the false detection sample information; determining the first step-by-step weight based on the first maximum intersection-and-union ratio and the first preset intersection-and-union ratio threshold; determining the second step-by-step weight based on the first maximum intersection-and-union ratio, the first preset intersection-and-union ratio threshold and the total number of target classifications; adding the first step-by-step weight to the target confidence loss function in the second loss function, and adding the second step-by-step weight to the classification loss function in the second loss function to obtain the first loss function.

[0094] For example, the second loss function L total2 It can be expressed as the following formula (1):

[0095]

[0096] Where N represents the number of detection layers of the target detection model to be trained, i∈[0,N]; B represents the number of anchors set for each Head layer; S represents the number of grids into which the current scale is divided; L CIOU represents the regression loss function. For example, the CIOU loss function can be used to calculate the positioning loss of positive samples. object represents the target confidence loss function, such as the cross entropy loss function; L class Represents the classification loss function, such as the cross entropy loss function.

[0097] For each false positive prediction frame, the false positive prediction frame can be calculated with each false positive sample frame in the corresponding false positive sample information by IOU, and the largest IOU value can be selected to obtain the first maximum intersection-over-union ratio, which can be used as the intersection-over-union ratio of the false positive prediction frame and the false positive sample frame. For example, based on the intersection-over-union ratio, the following formula (2) can be used to determine the first order weight g j :

[0098]

[0099] Where iou1 represents the first maximum intersection-over-union ratio calculated; iou th represents the first preset intersection-over-union ratio threshold; W1 is a constant.

[0100] For example, based on the first maximum intersection-over-combination ratio, the second order weight g can be determined using the following formula (3): c :

[0101]

[0102] Where iou1 represents the first maximum intersection-over-union ratio calculated; iou th represents the first preset intersection-over-union ratio threshold; n is the total number of target categories, that is, the total number of target categories that need to be detected in the target detection task; W2 is a constant.

[0103] For example, the first-order weight g is obtained j and the second-order weight g c After that, the first-order weight g can be used j For the second loss function L shown in the above formula (1), total2 The target confidence loss function L in object Weighted, and use the second-order weight g c For the second loss function L total2 The classification loss function L in class Weighted, we get the first loss function L total1 , which can be expressed as the following formula (4):

[0104]

[0105] According to the target detection model training method provided by the embodiment of the present invention, Figure 3 The schematic diagram of the method flow for performing reinforcement learning on the second target prediction frame based on the false detection sample information is shown as an example. Figure 3 As shown, the following steps 301 to 309 may be included.

[0106] Step 301: Obtain a second prediction box at the current scale.

[0107] The current batch of sample images is input into the target detection model to be trained with the same network structure as the detection model for forward reasoning to obtain the prediction results corresponding to each sample image at at least two different scales, that is, after obtaining the second target prediction frame, the prediction results at each scale can be processed separately, and for each scale, the second prediction frame at the scale (current scale) in the second target prediction frame is obtained.

[0108] Step 302: Determine a second maximum intersection-over-union ratio between the second prediction box and the corresponding positive sample.

[0109] For the prediction results of each sample image at each scale, each second prediction box at the scale can be calculated with all positive samples in the sample image, and the maximum IOU value is selected from the calculation results to obtain the second maximum intersection-over-union ratio iou2. The positive sample is the non-falsely detected prediction box in the sample image.

[0110] Step 303: Determine whether the second maximum IoU is less than a third preset IoU threshold.

[0111] After obtaining iou2, it is determined whether iou2 is less than a third preset intersection-over-union ratio threshold, such as whether it is less than 0.01. If it is, step 304 is executed, otherwise step 308 is executed.

[0112] Step 304: Determine a first maximum intersection-over-union ratio between the second prediction frame and the false detection sample frame.

[0113] When iou2 is less than the third preset intersection-over-union ratio threshold, the IOU calculation can be performed between the second prediction frame and all the falsely detected sample frames in the sample image, and the maximum IOU value can be selected from the calculation results to obtain the first maximum intersection-over-union ratio iou1.

[0114] Step 305: Determine whether the first maximum IoU is greater than a fourth preset IoU threshold.

[0115] After obtaining the first maximum intersection-and-union ratio iou1, it is determined whether iou1 is greater than a fourth preset intersection-and-union ratio threshold, such as whether it is greater than 0.8. If so, step 306 is executed, otherwise step 308 is executed.

[0116] Step 306: Determine a first order weight and a second order weight based on the first maximum intersection-over-union ratio.

[0117] When iou1 is greater than the fourth preset intersection-and-union ratio threshold, the second prediction frame can be determined to be a false positive prediction frame. At this time, the hierarchical weight calculation of the target confidence loss function and the classification loss function can be performed based on the first maximum intersection-and-union ratio iou1. Specifically, the first hierarchical weight g can be determined using the above formula (2): j , using the above formula (3) the second-order weight g c .

[0118] Step 307: Perform first-order weighting on the target confidence loss function, and perform second-order weighting on the classification loss function to calculate the loss value of the first loss function.

[0119] Determine the first-order weight g j and the second-order weight g c Afterwards, the first loss function L can be calculated according to the above formula (4): total1 The loss value.

[0120] Step 308: Calculate the loss value of the second loss function.

[0121] When the second maximum intersection-and-union ratio iou2 is greater than or equal to the third preset intersection-and-union ratio threshold, or when the first maximum intersection-and-union ratio iou1 is less than or equal to the fourth preset intersection-and-union ratio threshold, the loss value of the second loss function can be calculated using the above formula (1).

[0122] Step 309: Gradient calculation.

[0123] After completing the calculation of the loss value corresponding to each prediction box of each sample image at each scale in the current batch of sample images, the gradient result of the second target prediction box output when the current batch of sample images is subjected to secondary forward reasoning can be obtained, and then the gradient backpropagation is performed to update the network weight parameters of the target detection model to be trained, and the training of the current batch of sample images is completed, and the stored false positive sample information is cleared at the same time. In this way, the secondary forward training reasoning of the current batch of sample images is completed, and the training of the next batch of sample images is entered until the target detection model to be trained reaches the training end condition, and a trained target detection model is obtained.

[0124] The training method of the target detection model provided by the embodiment of the present invention can exclude the prediction boxes that may overlap with the positive samples around the misdetected samples by first performing IOU analysis with the positive samples for the prediction boxes obtained by the secondary forward reasoning, and then perform a step-by-step weighted loss function calculation on the prediction boxes that do not overlap with the positive samples around the misdetected samples. This can form a sample suppression learning range with gradually decreasing weights from the inside to the outside around the misdetected sample boxes obtained in the first forward pass, thereby achieving the purpose of suppressing the misdetected samples and the prediction information around them, thereby reducing the false detection rate of the target detection model and improving the generalization and robustness of the target detection model.

[0125] Based on the training method of the target detection model in the above embodiments, Figure 4 The second flow chart of the training method of the target detection model provided by the embodiment of the present invention is exemplified, taking model 1 as the target detection model to be trained and model 2 as the detection model as an example, referring to Figure 4 As shown, the method may include the following steps 401 to 410.

[0126] Step 401: Copy the weight parameters of model 1 to model 2.

[0127] Before the training of the first batch of sample images begins, the weight parameters of model 1 are copied to model 2. At this time, model 2 has the same target detection capability as model 1. That is, in the entire training, the weight parameters of model 1 are only copied to model 2 before the start of the first batch. In the subsequent training process, only the weight parameters of model 1 are updated, and the weight parameters of model 2 used for misdetected sample annotation are not updated.

[0128] Step 402: Perform the first forward inference on the current batch of sample images based on model 2 to obtain target prediction boxes of each sample image at different scales.

[0129] Step 403: After merging the target prediction frames of each sample image at different scales, perform NMS processing to obtain the first target prediction frame of each sample image.

[0130] Step 404: Determine and save misdetected sample information of each sample image based on the IOU value between the first target prediction box and the target annotation box.

[0131] Step 405: Perform secondary forward inference on the current batch of sample images based on model 1 to obtain a second target prediction box of each sample image at a different scale.

[0132] Step 406: Determine the falsely detected prediction frame and the non-falsely detected prediction frame in the second target prediction frame based on the falsely detected sample information.

[0133] Step 407: Determine a first loss value of a first loss function corresponding to the falsely detected prediction box, and a second loss value of a second loss function corresponding to the non-falsely detected prediction box.

[0134] Step 408: Perform gradient back propagation based on the first loss value and the second loss value, update the weight parameters of model 1, and clear the current misdetected sample information.

[0135] Step 409: Determine whether the current training has ended. That is, determine whether the current epoch has ended. If not, execute step 410, otherwise, end the current training.

[0136] Step 410: Enter the next batch of sample image training, that is, obtain the next batch of sample images, and return to step 402 to re-process the next batch of sample images.

[0137] In this way, the target detection model training method using secondary forward reinforcement learning can obtain the false detection samples predicted by the current target detection model to be trained based on the post-processing of the detection model with the same detection effect as the current target detection model to be trained during the first forward reasoning process, and then conduct targeted training optimization on these false detection samples during the secondary forward reasoning process, so that the target detection model to be trained can strengthen the learning of these false detection samples and optimize the recognition ability of the target detection model to be trained for the false detection area, thereby reducing the false detection rate of the trained target detection model and improving the generalization of the target detection model. This solution can achieve good target detection results in multi-target detection tasks such as human body, non-motor vehicle, motor vehicle, human face, human head, etc. in intelligent traffic monitoring scenarios.

[0138] The embodiment of the present invention also provides a target detection method, which can be applied to a monitoring device, or to an electronic device such as a terminal device or a server that is connected to the monitoring device for communication. The terminal device may include a mobile phone, a computer, a tablet computer, a wearable device, etc.; the server may include an independent server, a cluster server, or a cloud server, etc. The target detection method can also be applied to a target detection device set in an electronic device such as a terminal device or a server, and the target detection device can be implemented by software, hardware, or a combination of both.

[0139] Figure 5 The schematic diagram of the target detection method is shown as an example. Figure 5 As shown, the target detection method may include the following steps 510 to 520.

[0140] Step 510: Acquire a surveillance image to be detected.

[0141] Step 520: Input the surveillance image to be detected into the target detection model to obtain the target detection result of the surveillance image to be detected output by the target detection model.

[0142] Among them, the target detection model is trained based on the target detection model training method provided in an embodiment of the present invention.

[0143] The target detection model training device provided by the present invention is described below. The target detection model training device described below and the target detection model training method described above can be referenced to each other.

[0144] Figure 6 The schematic diagram of the structure of the target detection model training device provided by the embodiment of the present invention is exemplified. Figure 6 As shown, the target detection model training device may include: a parameter copying module 610, used to obtain the weight parameters of the target detection model to be trained, and copy the weight parameters to the detection model, wherein the detection model has the same structure as the target detection model to be trained; a first target detection module 620, used to perform target detection on the sample image based on the detection model to obtain the first target prediction frame of the sample image; a false detection determination module 630, used to determine the false detection sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image; a model training module 640, used to train and update the weight parameters of the target detection model to be trained based on the sample image and the false detection sample information to obtain a trained target detection model.

[0145] In an exemplary embodiment, the model training module 640 may include: a target detection unit, which is used to input a sample image into a target detection model to be trained to obtain a second target prediction box output by the target detection model to be trained; a first update unit, which is used to determine a first loss value of a first loss function for each second prediction box in the second target prediction box, when the second prediction box is determined to be a false detection prediction box based on the false detection sample information, and to perform training updates on weight parameters of the target detection model to be trained based on the first loss value; a second update unit, which is used to determine a second loss value of a second loss function, when the second prediction box is determined to be a non-false detection prediction box based on the false detection sample information, and to perform training updates on weight parameters of the target detection model to be trained based on the second loss value; wherein the first loss function is used to suppress the prediction information of the second prediction box; and the first loss function is obtained by performing step-by-step gradient weighting on the second loss function.

[0146] In an exemplary embodiment, the second loss function is determined based on the target confidence loss function, the classification loss function and the regression loss function. Accordingly, performing step-by-step gradient weighting on the second loss function includes: obtaining the first maximum intersection-over-union ratio between the false detection prediction box and each false detection sample box in the false detection sample information; determining the first step-by-step weight based on the first maximum intersection-over-union ratio and the first preset intersection-over-union ratio threshold; determining the second step-by-step weight based on the first maximum intersection-over-union ratio, the first preset intersection-over-union ratio threshold and the total number of target classifications; adding the first step-by-step weight to the target confidence loss function in the second loss function, and adding the second step-by-step weight to the classification loss function in the second loss function, to obtain the first loss function.

[0147] In an exemplary embodiment, the first updating unit may include: a first determining subunit, used to determine a second maximum intersection-and-union ratio between the second prediction box and a non-falsely detected prediction box in the sample image; a second determining subunit, used to determine a first maximum intersection-and-union ratio between the second prediction box and each falsely detected sample box in the falsely detected sample information when the second maximum intersection-and-union ratio is less than a third preset intersection-and-union ratio threshold; and a third determining subunit, used to determine the second prediction box as a falsely detected prediction box when the first maximum intersection-and-union ratio is greater than a fourth preset intersection-and-union ratio threshold.

[0148] In an exemplary embodiment, the first target detection module 620 may include: a prediction unit, used to input a sample image into a detection model, obtain a target prediction box at at least two scales output by the detection model, and merge the target prediction boxes at at least two scales to obtain a third target prediction box; a suppression unit, used to perform non-maximum suppression processing on the third target prediction box to obtain a first target prediction box.

[0149] In an exemplary embodiment, the false detection determination module 630 may include: a first determination unit, used to determine, for each first prediction box in the first target prediction box, a first intersection-and-union ratio between the first prediction box and each target annotation box of the sample image; and a second determination unit, used to determine the information of the first prediction box as false detection sample information of the sample image when the first intersection-and-union ratios are all less than a second preset intersection-and-union ratio threshold.

[0150] The target detection device provided by the present invention is described below. The target detection device described below and the target detection method described above can be referenced to each other.

[0151] Figure 7 The schematic diagram of the structure of the target detection device provided by the embodiment of the present invention is exemplarily shown. Figure 7 As shown, the target detection device may include: an image acquisition module 710, used to acquire a monitoring image to be detected; a second target detection module 720, used to input the monitoring image to be detected into a target detection model, and obtain a target detection result in the monitoring image to be detected output by the target detection model, wherein the target detection model is trained based on the target detection model training method provided in an embodiment of the present invention.

[0152] Figure 8 An example of a structural diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the training method of the target detection model provided by any of the above method embodiments, or execute the target detection method provided by the above method embodiments.

[0153] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0154] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the target detection model provided by any of the above method embodiments, or execute the target detection method provided by the above method embodiments.

[0155] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the training method of the target detection model provided by any of the above method embodiments, or to execute the target detection method provided by the above method embodiments.

[0156] By way of example, computer readable storage media include non-transitory computer readable storage media.

[0157] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0158] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target detection model training method, characterized in that: include: Obtain weight parameters of the target detection model to be trained, and copy the weight parameters to the detection model; The detection model has the same structure as the target detection model to be trained; Performing target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image, and determining misdetected sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image; The weight parameters of the target detection model to be trained are trained and updated based on the sample image and the false detection sample information to obtain a trained target detection model.

2. The target detection model training method according to claim 1, characterized in that: The training and updating of the weight parameters of the target detection model to be trained based on the sample image and the false detection sample information includes: Inputting the sample image into the target detection model to be trained to obtain a second target prediction frame output by the target detection model to be trained; For each second prediction box in the second target prediction box, when the second prediction box is determined to be a falsely detected prediction box based on the falsely detected sample information, determining a first loss value of a first loss function, and training and updating the weight parameters of the target detection model to be trained based on the first loss value; In a case where it is determined that the second prediction box is a non-falsely detected prediction box based on the falsely detected sample information, determining a second loss value of a second loss function, and training and updating the weight parameters of the target detection model to be trained based on the second loss value; Among them, the first loss function is used to suppress the prediction information of the second prediction box; the first loss function is obtained by performing step-by-step gradient weighting on the second loss function.

3. The target detection model training method according to claim 2, characterized in that: The second loss function is determined based on a target confidence loss function, a classification loss function, and a regression loss function; and performing step-by-step gradient weighting on the second loss function includes: Obtaining a first maximum intersection-over-union ratio between the false detection prediction frame and each false detection sample frame in the false detection sample information; Determine a first step weight based on the first maximum intersection-over-union ratio and a first preset intersection-over-union ratio threshold; Determine a second hierarchical weight based on the first maximum intersection-over-union ratio, the first preset intersection-over-union ratio threshold, and the total number of target categories; The first order weight is added to the target confidence loss function in the second loss function, and the second order weight is added to the classification loss function in the second loss function to obtain the first loss function.

4. The target detection model training method according to claim 2, characterized in that: The determining, based on the falsely detected sample information, that the second prediction box is a falsely detected prediction box includes: Determine a second maximum intersection-over-union ratio between the second prediction box and a non-falsely detected prediction box in the sample image; When the second maximum intersection-over-union ratio is less than a third preset intersection-over-union ratio threshold, determining a first maximum intersection-over-union ratio between the second prediction frame and each falsely detected sample frame in the falsely detected sample information; When the first maximum IoU is greater than a fourth preset IoU threshold, the second prediction frame is determined as the false detection prediction frame.

5. The target detection model training method according to any one of claims 1 to 4, characterized in that: The performing target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image includes: Inputting the sample image into the detection model to obtain target prediction frames at at least two scales output by the detection model, and merging the target prediction frames at at least two scales to obtain a third target prediction frame; Perform non-maximum suppression processing on the third target prediction box to obtain the first target prediction box.

6. The target detection model training method according to any one of claims 1 to 4, characterized in that: The determining the misdetected sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image includes: For each first prediction box in the first target prediction box, determining a first intersection-over-union ratio between the first prediction box and each target annotation box of the sample image; When the first intersection-over-union ratios are all smaller than a second preset intersection-over-union ratio threshold, the information of the first prediction frame is determined as misdetected sample information of the sample image.

7. A target detection method, characterized in that: include: Acquire the surveillance image to be detected; The monitoring image to be detected is input into a target detection model to obtain a target detection result of the monitoring image to be detected output by the target detection model; wherein the target detection model is trained based on the target detection model training method according to any one of claims 1 to 6.

8. A target detection model training device, characterized in that: include: A parameter copying module is used to obtain weight parameters of the target detection model to be trained and copy the weight parameters to the detection model; The detection model has the same structure as the target detection model to be trained; A first target detection module, used to perform target detection on the sample image based on the detection model to obtain a first target prediction frame of the sample image; a false detection determination module, configured to determine false detection sample information of the sample image based on the first target prediction frame and the target annotation frame of the sample image; The model training module is used to train and update the weight parameters of the target detection model to be trained based on the sample image and the false detection sample information to obtain a trained target detection model.

9. A target detection device, characterized in that: include: An image acquisition module is used to acquire the monitoring image to be detected; The second target detection module is used to input the monitoring image to be detected into the target detection model to obtain the target detection result in the monitoring image to be detected output by the target detection model; wherein the target detection model is trained based on the target detection model training method according to any one of claims 1 to 6.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the target detection model training method as described in any one of claims 1 to 6, or implements the target detection method as described in claim 7.

Citation Information

Cited By

  • Guardrail detection method, device and equipment and storage medium

    CN120976709A