Method, apparatus, and computer-readable storage medium for adjusting image analysis model

By introducing a pseudo-label mechanism into the detection model, the false detection box is marked as an unknown category and fine-tuned for the category branches, the problem of high false detection rate of detection models in open scenarios is solved, and the model identification ability and adjustment efficiency are improved.

CN114359669BActive Publication Date: 2025-06-10GUANGZHOU YUNCONG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111683471.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-10
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The existing detection models have a high risk of false detection in open scenarios, and existing methods such as model training based on threshold filtering and false detection as backgrounds have problems of inefficiency.

Method used

By introducing a pseudo-label mechanism into the detection model, the false detection box is marked as unknown categories, and these pseudo-labels are used to adjust the detection head of the model, especially fine-tuning for category branches, to reduce the output of unknown categories.

Benefits of technology

It effectively reduces the false detection rate of the model in open scenarios, improves the recognition ability and adjustment efficiency of the model, and avoids the problem of affecting subsequent identification results due to false detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359669B_ABST
    Figure CN114359669B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer processing technologies, and specifically provides a method and apparatus for adjusting an image analysis model, as well as a computer-readable storage medium, aiming to solve the technical problem of quickly adjusting the model to reduce the false detection situation. For this purpose, the method of the present invention includes: inputting training images into the model, which carry annotation data indicating the categories of the first detection boxes, the model detection head includes a first branch and a second branch, calculating the confidence and category of the second detection boxes; when the confidence is higher than a preset level, determining whether the category of the second detection box is the same as that of the first detection box, and when they are different, setting a pseudo-label for the second detection box and recording its category as an unknown category; adjusting the second branch; after the adjustment is completed, prohibiting the model from outputting the result that the category of the detection box is an unknown category. The present invention only adjusts the branch that outputs the category of the detection box, which is beneficial to improving the efficiency of model adjustment, prohibits the model from outputting the detection result of an unknown type, and improves the recognition ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and specifically provides a method and apparatus for adjusting an image analysis model, as well as a computer-readable storage medium. Background Art

[0002] With the development of the field of artificial intelligence, computer vision technology has been widely applied in life. Well-known applications such as face swiping payment, intelligent monitoring, and autonomous driving all rely on a system based on computer vision for support. Most of the first-step work of these systems is target detection tasks, which find the targets (objects) of interest in the image, determine their categories and positions, and then hand them over to subsequent modules such as recognition and tracking in the system for further processing. Such detection models are usually only for specific detection categories, such as faces, vehicles, pedestrians, commodities, etc., and the models can be deployed in any environment in an open scenario.

[0003] However, the learning of detection models based on specific tasks all face a common problem. The training data scenarios in the preparation stage are single. Even a model that performs well in the training set will have many unexpected false detections when deployed in an open scenario. Therefore, during the actual deployment process of the model, the scene diversity and complexity increase sharply, and the risk of false detection increases significantly. The existence of false detections will greatly affect the subsequent recognition results of the system. Therefore, how to quickly iterate the model on a limited training set and reduce the false detections of the model in an open scenario is a very meaningful and necessary technical problem.

[0004] For the problem of false detections of detection models, the currently mainly used methods include the following two:

[0005] (1) Method based on threshold filtering: Increase the output confidence threshold of the model and only output the detection targets that the model is certain about with higher scores. However, this will inevitably reduce the recall rate of the model and decrease the effective output.

[0006] (2) Model training with false detections as the background: If there are always false detections of a certain type of target, some corresponding samples can be added to the training set to improve the discrimination ability of the model, and false detections can be reduced without reducing the recall. This method also has a direct problem. In order to enable the model to fully learn the generalization of false detection objects as the background, end-to-end learning is required. For each newly added type of false detection object, the model needs to be retrained end-to-end. When the training set data volume is huge, such an iteration speed is obviously unacceptable. Summary of the Invention

[0007] In order to overcome the above defects, the present invention is proposed to provide a method and apparatus for adjusting an image analysis model, as well as a computer-readable storage medium, which can solve or at least partially and quickly adjust the model to reduce the situation of false detections.

[0008] In a first aspect, the present invention provides a method for adjusting a picture analysis model, the method comprising:

[0009] Input a training picture into the model, the training picture carrying annotation data indicating the category of a first detection box of a target object in the training picture, the detection head of the model including a first branch and a second branch, and respectively calculating the confidence and category of a second detection box for indicating the target object;

[0010] When the confidence of the second detection box is higher than a preset level, determine whether the category of the second detection box is the same as the category of the first detection box. When they are different, set a pseudo-label for the second detection box, and record the category of the second detection box as an unknown category in the pseudo-label;

[0011] Use the second detection box and the pseudo-label as the training data of the second branch, input them into the second branch, and adjust the second branch according to the output result;

[0012] After the model is adjusted, prohibit the model from outputting a result with the detection box category being the unknown category.

[0013] In a technical solution of the above picture analysis model adjustment method, the step of "adjusting the second branch according to the output result" includes:

[0014] Calculate the loss value of the output result according to a preset loss function, and adjust the parameters of the second branch according to the loss value;

[0015] And / or,

[0016] Before the step of "after the model is adjusted, prohibit the model from outputting a result with the detection box category being the unknown category", it further includes:

[0017] After detecting that the loss value is less than a preset threshold, determine that the model adjustment is completed;

[0018] And / or,

[0019] Before the step of "inputting the training picture into the model", it further includes:

[0020] Obtain the training picture according to the detection box category in the historical error results output by the model;

[0021] And / or,

[0022] Before the step of "using the second detection box and the pseudo-label as the training data of the second branch", it further includes:

[0023] If the second detection box and other detection boxes are in the same connected domain, update the category of the second detection box according to the category of the other detection boxes;

[0024] and / or,

[0025] The annotation data also indicates the position of the first detection box, and the detection head further includes a third branch for calculating the position of the second detection box;

[0026] and / or,

[0027] The model further includes a feature extraction layer, which includes a backbone network and a multi-scale feature fusion network. The backbone network is used to extract multi-scale features from the training pictures, and the multi-scale feature fusion network is used to fuse the multi-scale features of the training pictures into the features of the training pictures for inputting into the detection head.

[0028] In a second aspect, there is provided an apparatus for adjusting a picture analysis model, the apparatus including:

[0029] A picture input module for inputting a training picture into the model. The training picture carries annotation data indicating the category of a first detection box of a target object in the training picture. The detection head of the model includes a first branch and a second branch, which respectively calculate the confidence and category of a second detection box for indicating the target object;

[0030] A category setting module for, when the confidence of the second detection box is higher than a preset level, determining whether the category of the second detection box is the same as the category of the first detection box. When they are different, setting a pseudo-label for the second detection box and recording the category of the second detection box as an unknown category in the pseudo-label;

[0031] A branch adjustment module for using the second detection box and the pseudo-label as training data of the second branch, inputting them into the second branch, and adjusting the second branch according to the output result;

[0032] An output control module for, after the model is adjusted, prohibiting the model from outputting a result with the detection box category being the unknown category.

[0033] In a third aspect, there is provided a control device, which includes a processor and a storage device. The storage device is adapted to store multiple program codes, and the program codes are adapted to be loaded and run by the processor to execute the picture analysis model adjustment method according to any one of the technical solutions in the technical solution of the above picture analysis model adjustment method.

[0034] In a fourth aspect, a computer-readable storage medium is provided, in which multiple program codes are stored, and the program codes are adapted to be loaded and run by a processor to execute the above-mentioned picture analysis model adjustment method according to any one of the technical solutions of the above-mentioned picture analysis model adjustment method.

[0035] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:

[0036] In an embodiment of the present invention, the picture analysis model adjustment method may include the following steps: inputting a training picture into the model, where the training picture carries annotation data indicating the category of the first detection box of the target object in the training picture, and the detection head of the model includes a first branch and a second branch, respectively calculating the confidence and category of the second detection box for indicating the target object; when the confidence of the second detection box is higher than a preset level, determining whether the category of the second detection box is the same as the category of the first detection box, and when they are different, setting a pseudo-label for the second detection box and recording the category of the second detection box as an unknown category in the pseudo-label; using the second detection box and the pseudo-label as the training data of the second branch, inputting them into the second branch, and adjusting the second branch according to the output result; after the model is adjusted, prohibiting the model from outputting the result with the detection box category being an unknown category. In the technical solution of the present invention, the detection head of the model is decoupled, and only the branch for outputting the detection box category is adjusted when the model is adjusted, which is beneficial to improving the efficiency of model adjustment. By introducing the misdetection situation of the unknown type annotation model, and by prohibiting the model from outputting the detection result of the unknown type, in fact, the misdetection result output by the model is reduced, effectively improving the recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Referring to the accompanying drawings, the disclosure of the present invention will become easier to understand. It is easy for those skilled in the art to understand that: these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. Among them:

[0038] Figure 1 is a schematic flowchart of the main steps of the picture analysis model adjustment method according to an embodiment of the present invention;

[0039] Figure 2 is a schematic flowchart of the main steps of the picture analysis model adjustment method according to an embodiment of the present invention;

[0040] Figure 3 is a structural diagram of the model used in the picture analysis model adjustment method according to an embodiment of the present invention;

[0041] Figure 4 is a working principle diagram of the picture analysis model adjustment method according to an embodiment of the present invention;

[0042] Figure 5 It is a flowchart of the operation used in the method for adjusting a picture analysis model according to an embodiment of the present invention;

[0043] Figure 6 It is a schematic diagram of the main structural block diagram of a device for adjusting a picture analysis model according to another embodiment of the present invention;

[0044] Figure 7 It is a schematic diagram of the main structural block diagram of a device for adjusting a picture analysis model according to another embodiment of the present invention. Detailed implementation manners

[0045] The following describes some embodiments of the present invention with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.

[0046] In the description of the present invention, "module" and "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, memories, and may also include a software part, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. The non-transitory computer-readable storage medium includes any suitable medium for storing program code, such as magnetic disks, hard disks, optical disks, flash memories, read-only memories, random access memories, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "this" may also include the plural form.

[0047] Refer to the attached Figure 1 , Figure 1 It is a schematic diagram of the main step flow of a method for adjusting a picture analysis model according to an embodiment of the present invention.

[0048] As Figure 1 shown, the method for adjusting a picture analysis model in the embodiment of the present invention mainly includes the following steps:

[0049] Step S110, input the training pictures into the model. The training pictures carry annotation data indicating the category of the first detection box of the target object in the training pictures. The detection head of the model includes a first branch and a second branch, and respectively calculates the confidence and category of the second detection box for indicating the target object.

[0050] Generally, for currently mainstream models, the position of the first detection box is also indicated in the labeled data, and the detection head further includes a third branch for calculating the position of the second detection box. The model further includes a feature extraction layer, which includes a backbone network and a multi-scale feature fusion network. The backbone network is used to extract multi-scale features from the training images, and the multi-scale feature fusion network is used to fuse the multi-scale features of the training images into the features of the training images for input to the detection head.

[0051] In this embodiment, the type of the target object is not restricted. For example, it can be a pedestrian, a vehicle, a commodity, etc. The type of the model is not restricted in this embodiment. For example, the YOLOV5 model can be adopted, and the Obj branch, Cls branch, and Box branch of its detection head are the first branch, the second branch, and the third branch respectively. In this embodiment, decoupling processing is performed on the Box branch, Obj branch, and Cls branch. At this time, the model has been trained, but false detection is likely to occur, and the technical solution of this embodiment is required to fine-tune the model.

[0052] Step S120, when the confidence of the second detection box is higher than a preset level, determine whether the category of the second detection box is the same as that of the first detection box. When they are different, set a pseudo-label for the second detection box, and record the category of the second detection box as an unknown category in the pseudo-label.

[0053] Here, the confidence output by the Obj branch is obtained. If the confidence of the second detection box is higher than a certain threshold or belongs to the highest several values, then it is detected whether the category calculated by the Cls branch matches the true category labeled by the first detection box in the training image. If they do not match, it indicates that false detection has occurred in the model. Therefore, in this embodiment, those detection boxes with relatively high confidence scores but not belonging to any true category (i.e., false detection occurs) are classified into the "unknown" category, that is, the unknown category.

[0054] Step S130, use the second detection box and the pseudo-label as the training data of the second branch, input them into the second branch, and adjust the second branch according to the output result.

[0055] At this time, the feature extraction layer and the Box and Obj branches of the model are fixed, and only the Cls branch is trained. Since only the Cls branch network is trained, the finetune efficiency is improved. Using the detection boxes of the "unknown" category to train the Cls branch can enable the model to learn for false detection situations, that is, for objects that are prone to false detection, the model will recognize them as the "unknown" category.

[0056] Step S140, after the model is adjusted, prohibit the model from outputting the result that the detection box category is an unknown category.

[0057] When fine-tuning reaches a preset level, model training is completed. At this time, the model is prohibited from outputting detection boxes of the "unknown" category. Since the "unknown" category belongs to the result of incorrect recognition before model fine-tuning, the technical solution of this embodiment prohibits the output of false detection results to the user.

[0058] Through the technical solution of this embodiment, the detection head of the model is decoupled, and only the branch of the output detection box category is adjusted when the model is adjusted, which is beneficial to improving the efficiency of model adjustment. By introducing the false detection situation of the unknown type labeling model and prohibiting the model from outputting unknown type detection results, it actually reduces the false detection results output by the model, and effectively improves the recognition ability of the model.

[0059] See attached Figure 2 , Figure 2 It is a flowchart of main steps of a method for adjusting a picture analysis model according to an embodiment of the present invention.

[0060] The technical solution of this embodiment is mainly used in the model fine-tuning stage. Before the fine-tuning stage, the model needs to be trained end-to-end with all positive sample data, which specifically includes the following process:

[0061] (I) Data preprocessing: Construct a base training set, retain only the data containing the target detection object, annotate the target object in the image, including the category and detection box coordinates, obtain image data, and construct a Yo Lo type data set as the all-positive sample training set.

[0062] (ii) Data enhancement: We use a wide range of online data enhancement methods to maximize the diversity of training set data. Enhancement methods include Mosaic, mixup, scaling, flipping, rotation, affine transformation, brightness, contrast, saturation, motion blur, image compression blur, and other enhancement methods.

[0063] (III) Image feature extraction and fusion. The open source model YOLOV5 is used as the basic detection network framework. The backbone part uses the CSP (cross-stage local network) structure backbone network to extract the multi-scale features of the image. For the multi-scale features obtained by the backbone, the FPN (target detection algorithm) structure transfers high-level strong semantic features from top to bottom through upsampling for feature fusion. Then, the multi-scale features output by FPN are further fused by conveying strong positioning features from bottom to top through the PAN (pixel aggregation network) structure, thereby obtaining 1 / 8, 1 / 16, and 1 / 32 scale features of the input image, completing feature extraction. The overall structure of the detection network is as follows: Figure 3 shown.

[0064] (4) Decoupled detection head. The detection head consists of three branches, namely the Box branch (4-dimensional output) responsible for predicting the position information of the detection box, the Obj branch (1-dimensional output) responsible for predicting the confidence of the detection box, and the Cls branch (number of categories + 1) responsible for predicting the category of the detection box. The multi-scale features of the image obtained in the previous step are respectively input into the corresponding detection heads to obtain the network output. After post-processing operations such as nms (non-maximum suppression), detection boxes are obtained. The network structure of the detection head is as Figure 4 shown.

[0065] (5) End-to-end training with all positive samples. On the all positive samples training set in a simple scenario, end-to-end training is performed on the model. The loss function part is kept consistent with the original yolov5. The CIOU_Loss (a loss function) is adopted for the Box branch, the BCE_Loss (a loss function) is adopted for the Obj branch, and the BCE_Loss is also adopted for the Cls branch according to the labeled true category (assuming a total of C categories). At this time, only the 1-C categories of the Cls branch are trained, and the 0 category is ignored. The detection network trained in this way can ensure the detection of the target object to the greatest extent. At the same time, it is obvious that due to the single training samples, the false detection rate of the model in an open scenario is also extremely high.

[0066] As Figure 2 shown, the method for adjusting the image analysis model in the embodiment of the present invention mainly includes the following steps:

[0067] Step S210, obtain training images according to the categories of the detection boxes in the historical error results output by the model.

[0068] Update the training set here. By testing or pilot deploying the model, the categories that the model is prone to false detection are found, and some negative samples that are prone to false detection are added to the training set specifically. The data can be collected specifically; or directly add the false detection images in the actual deployment process to the training set to expand the training set samples.

[0069] Step S220, input the training images into the model. The training images carry annotation data indicating the category of the first detection box of the target object in the training images. The detection head of the model includes a first branch and a second branch, and calculate the confidence and category of the second detection box for indicating the target object respectively.

[0070] Step S230, when the confidence of the second detection box is higher than a preset level, determine whether the category of the second detection box is the same as the category of the first detection box. When they are different, set a pseudo-label for the second detection box, and record the category of the second detection box as an unknown category in the pseudo-label.

[0071] As Figure 5As shown, the finetuning of the Cls branch of the detection head starts at this time. It is known that the Cls branch responsible for predicting the categories of detection boxes outputs C + 1 category scores. In this embodiment, the additional category is defined as the "unknown" category with a label of 0, indicating those detection boxes with relatively high confidence scores but not belonging to any training categories. Generation method of "unknown" category pseudo-labels: During the forward propagation process, sort the detection boxes according to the scores output by the Obj branch, and select the top k detection boxes with the highest scores for training the Cls branch. Among them, the detection boxes that overlap with the true category labels maintain the true category labels, and the remaining detection boxes are defined as the "unknown" category.

[0072] Step S240, if the second detection box and other detection boxes are in the same connected domain, update the category of other detection boxes to the category of the second detection box.

[0073] Considering the definition rule of positive samples, in order to avoid the detection boxes at the edges of the target object being labeled as the "unknown" category, perform the minimum connected domain calculation on the output of the Cls branch, and obtain the same Cls label within the same connected domain.

[0074] Step S250, use the second detection box and the pseudo-labels as the training data of the second branch, input them into the second branch, calculate the loss value of the output result according to the preset loss function, adjust the parameters of the second branch according to the loss value, and after detecting that the loss value is less than the preset threshold, determine that the model adjustment is completed.

[0075] In this embodiment, the type of the loss function is not restricted. During the finetuning stage, fix the feature extraction layer, Box, and Obj branches of the network, and only train the Cls branch. Since only the Cls branch network is trained, the finetuning process is much faster than the end-to-end training. The existence of the "unknown" category not only ensures that the model can quickly learn the misdetected objects in the training set, but also has a very high probability of identifying the possible misdetected objects in unknown scenarios as the "unknown" category.

[0076] Step S260, after the model adjustment is completed, prohibit the model from outputting the results with the detection box category being the unknown category.

[0077] During the actual deployment process, as long as the output of the "unknown" category detection boxes is deleted, a part of the misdetections can be effectively filtered, and the detection performance can be improved.

[0078] According to the technical solution of this embodiment, it helps the detection model to further filter misdetected objects. This embodiment proposes a model fine-tuning solution based on the first-order detection model YOLOV5, decouples the model detection head network, which is divided into a Box branch, an Obj branch, and a Cls branch. During training, first, end-to-end training of the network is performed on the collected all-positive sample data. Subsequently, after obtaining data with complex scenes that are prone to misdetection, the Cls branch is fine-tuned by introducing the classification of the "unknown" category, and the model can quickly filter misdetections while ensuring detection. For the open-scene object detection task, this embodiment proposes a model training scheme that can effectively filter open-scene misdetections. By means of second-order fine-tuning, the detection model is rapidly iterated, effectively improving the misdetection situation of the detection model in the open scene, and providing stable and accurate object detection targets for computer vision recognition systems for specific tasks.

[0079] See Appendix Figure 6 , Figure 6 It is a schematic diagram of the main structural block diagram of an image analysis model adjustment device according to an embodiment of the present invention.

[0080] As Figure 6 shown, the image analysis model adjustment device in the embodiment of the present invention mainly includes the following modules:

[0081] An image input module 610 inputs training images into the model. The training images carry annotation data indicating the category of the first detection box of the target object in the training images. The detection head of the model includes a first branch and a second branch, which respectively calculate the confidence and category of the second detection box for indicating the target object.

[0082] Generally, for current mainstream models, the annotation data also indicates the position of the first detection box, and the detection head also includes a third branch for calculating the position of the second detection box. The model also includes a feature extraction layer, which includes a backbone network and a multi-scale feature fusion network. The backbone network is used to extract multi-scale features from the training images, and the multi-scale feature fusion network is used to fuse the multi-scale features of the training images into the features of the training images for input to the detection head.

[0083] In this embodiment, the type of the target object is not restricted. For example, it can be a pedestrian, a vehicle, a commodity, etc. The type of the model is not restricted in this embodiment. For example, the YOLO (a model) V5 model can be used. The Obj branch, Cls branch, and Box branch of its detection head are the first branch, the second branch, and the third branch respectively. In this embodiment, the Box branch, Obj branch, and Cls branch are decoupled. At this time, the model has been trained, but misdetection is likely to occur, and the model needs to be fine-tuned through the technical solution of this embodiment.

[0084] The category setting module 620 determines whether the category of the second detection box is the same as that of the first detection box when the confidence level of the second detection box is higher than the preset level. When they are different, a pseudo-label is set for the second detection box, and the category of the second detection box is recorded as an unknown category in the pseudo-label.

[0085] Here, the confidence level output by the Obj branch is obtained. If the confidence level of the second detection box is higher than a certain threshold or belongs to the highest several values, it is detected whether the category calculated by the Cls branch matches the true category labeled by the first detection box in the training image. If they do not match, it indicates that the model has a misdetection situation. Therefore, in this embodiment, the detection boxes with high confidence scores but not belonging to any true category (i.e., misdetection) are classified as the "unknown" category, that is, the unknown category.

[0086] The branch adjustment module 630 uses the second detection box and the pseudo-label as the training data of the second branch, inputs them into the second branch, and adjusts the second branch according to the output result.

[0087] At this time, the feature extraction layer and the Box and Obj branches of the model are fixed, and only the Cls branch is trained. Since only the Cls branch network is trained, the finetune efficiency is improved. Using the detection boxes of the "unknown" category to train the Cls branch can enable the model to learn for misdetection situations, that is, for objects that are prone to misdetection, the model will identify them as the "unknown" category.

[0088] The output control module 640 prohibits the model from outputting the result that the category of the detection box is an unknown category after the model adjustment is completed.

[0089] When the fine-tuning reaches the preset level, the model training is completed. At this time, the model is prohibited from outputting the detection boxes of the "unknown" category. Since the "unknown" category belongs to the result of incorrect recognition before the model fine-tuning, the technical solution of this embodiment is to prohibit the misdetection result from being output to the user.

[0090] Through the technical solution of this embodiment, the detection head of the model is decoupled. When adjusting the model, only the branch that outputs the category of the detection box is adjusted, which is beneficial to improving the efficiency of model adjustment. By introducing the misdetection situation of the unknown type annotation model, and by prohibiting the model from outputting the detection results of the unknown type, in fact, the misdetection results output by the model are reduced, effectively improving the recognition ability of the model.

[0091] Refer to the appendix Figure 7 , Figure 7 which is a schematic diagram of the main structure block diagram of the picture analysis model adjustment device according to an embodiment of the present invention.

[0092] The technical solution of this embodiment is mainly used in the model fine-tuning stage. Before the fine-tuning stage, end-to-end training of the model with all positive sample data is required, which specifically includes the following processes:

[0093] (1) Data preprocessing. Construct a base training set, only retain the data containing the target detection object, and perform data annotation on the target object in the picture, including the category and the coordinates of the detection box, to obtain image data, and construct a YOLO-type data set as the all-positive sample training set.

[0094] (2) Data augmentation. Adopt rich online data augmentation to improve the diversity of the training set data as much as possible. The augmentation methods include mosaic, mixup, scaling, flipping, rotation, affine transformation, brightness, contrast, saturation, motion blur, picture compression blur and other augmentation methods.

[0095] (3) Picture feature extraction and fusion. Use the open-source model YOLOV5 as the basic detection network framework, and use the CSP (Cross-Stage Partial Network) structure backbone network in the backbone part to extract multi-scale features of the picture; for the multi-scale features obtained by the backbone, the FPN (Feature Pyramid Network) structure passes the high-level strong semantic features from top to bottom through upsampling for feature fusion; then, through the PAN (Path Aggregation Network) structure, the strong localization features are conveyed from bottom to top for further fusion of the multi-scale features output by the FPN, thereby obtaining the 1 / 8, 1 / 16, and 1 / 32 scale features of the input picture, completing feature extraction. The overall structure of the detection network is as Figure 3 shown.

[0096] (4) Decoupled detection head. The detection head consists of three branches, namely the Box branch (4-dimensional output) responsible for predicting the position information of the detection box, the Obj branch (1-dimensional output) responsible for predicting the confidence of the detection box, and the Cls branch (number of categories + 1) responsible for predicting the category of the detection box. Input the multi-scale features of the picture obtained in the previous step into the corresponding detection heads respectively to obtain the network output, and obtain the detection box through post-processing operations such as nms (Non-Maximum Suppression). The network structure of the detection head is as Figure 4 shown.

[0097] (5) End-to-end training with all positive samples. On the all positive sample training set in a simple scenario, perform end-to-end training on the model. The loss function part is consistent with the original YOLOv5. The Box branch uses CIOU_Loss (a loss function), the Obj branch uses BCE_Loss (a loss function), and the Cls branch also uses BCE_Loss according to the labeled true category (assuming a total of C categories). At this time, only the 1 - C classes in the Cls branch are trained, and the 0 class is ignored. The detection network trained in this way can ensure the detection of the target object to the greatest extent. At the same time, it is obvious that due to the singularity of the training samples, the false detection rate of the model in an open scenario is also extremely high.

[0098] As Figure 7 shown, the picture analysis model adjustment device in the embodiment of the present invention mainly includes the following modules:

[0099] The picture acquisition module 710 acquires training pictures according to the detection box categories in the historical error results output by the model.

[0100] Update the training set here. By testing or pilot deploying the model, find the categories that the model is prone to false detection, and add some negative samples that are prone to false detection to the training set specifically. The data can be collected specifically; or directly add the false detection pictures in the actual deployment process to the training set to expand the training set samples.

[0101] The picture input module 720 inputs the training pictures into the model. The training pictures carry annotation data indicating the category of the first detection box of the target object in the training pictures. The detection head of the model includes a first branch and a second branch, which respectively calculate the confidence and category for indicating the second detection box of the target object.

[0102] The category setting module 730 determines whether the category of the second detection box is the same as the category of the first detection box when the confidence of the second detection box is higher than the preset level. When they are different, set a pseudo-label for the second detection box and record the category of the second detection box as an unknown category in the pseudo-label.

[0103] As Figure 5 shown, start finetuning the Cls branch of the detection head at this time. It is known that the Cls branch responsible for predicting the detection box category outputs C + 1 category scores. In this embodiment, the additional category is defined as the "unknown" category, and the label is 0, indicating those detection boxes with relatively high confidence scores but not belonging to any training category. Generation method of the "unknown" category pseudo-label: During the forward propagation process, sort the detection boxes according to the score output by the Obj branch, and take the top k detection boxes with the highest scores for training the Cls branch. Among them, the detection boxes that overlap with the true category labels maintain the true category labels, and the remaining detection boxes are defined as the "unknown" category.

[0104] The category setting module 730 updates the category of other detection boxes to the category of the second detection box if the second detection box and other detection boxes are in the same connected component.

[0105] Considering the definition rule of positive samples, in order to avoid the detection boxes at the edge of the target object being labeled as the "unknown" class, the minimum connected component calculation is performed on the output of the Cls branch, and the same Cls label is obtained within the same connected component.

[0106] The branch adjustment module 740 uses the second detection box and the pseudo-label as the training data of the second branch, inputs them into the second branch, calculates the loss value of the output result according to the preset loss function, adjusts the parameters of the second branch according to the loss value, and determines that the model adjustment is completed after detecting that the loss value is less than the preset threshold.

[0107] In this embodiment, the type of the loss function is not restricted. In the finetune stage, the feature extraction layer of the network, the Box branch, and the Obj branch are fixed, and only the Cls branch is trained. Since only the Cls branch network is trained, the finetune process is much faster than the end-to-end training. The existence of the "unknown" class not only ensures that the model can quickly learn the misdetected objects in the training set, but also has a very high probability of identifying the misdetected objects that may exist in unknown scenarios as the "unknown" class.

[0108] The output control module 750 prohibits the model from outputting the results with the detection box category being the unknown class after the model adjustment is completed.

[0109] In the actual deployment process, as long as the output of the "unknown" category detection box is deleted, a part of the misdetections can be effectively filtered, and the detection performance can be improved.

[0110] According to the technical solution of this embodiment, it helps the detection model to further filter misdetected objects. This embodiment proposes a model fine-tuning scheme based on the first-order detection model YOLOV5, decouples the model detection head network into a Box branch, an Obj branch, and a Cls branch. During training, first, the network is trained end-to-end on the collected all-positive sample data; subsequently, after obtaining the data with complex scenarios that are prone to misdetections, the Cls branch is fine-tuned by introducing the classification of the "unknown" class, and the model can quickly filter misdetections while ensuring detection. This embodiment proposes a model training scheme for open-scene object detection tasks that can effectively filter open-scene misdetections, quickly iterates the detection model through the second-order fine-tuning method, effectively improves the misdetection situation of the detection model in open scenes, and provides stable and accurate object detection objects for computer vision recognition systems for specific tasks.

[0111] The above Figures 6 to 7The illustrated picture analysis model adjustment device is used to execute Figures 1 to 2 the illustrated embodiment of the picture analysis model adjustment method. The technical principles, the technical problems solved, and the technical effects produced by both are similar. Those skilled in the art of this technology can clearly understand that for the convenience and conciseness of description, the specific working process and related explanations of the picture analysis model adjustment device can refer to the content described in the embodiment of the picture analysis model adjustment method, and will not be elaborated here.

[0112] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0113] Furthermore, the present invention also provides a control device. In an embodiment of the control device according to the present invention, the control device includes a processor and a storage device. The storage device can be configured to store a program for executing the picture analysis model adjustment method of the above-mentioned method embodiment, and the processor can be configured to execute the program in the storage device. The program includes, but is not limited to, the program for executing the picture analysis model adjustment method of the above-mentioned method embodiment. For the convenience of description, only the part related to the embodiment of the present invention is shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The control device can be a control device formed by various electronic devices.

[0114] Further, the present invention also provides a computer-readable storage medium. In an embodiment of the computer-readable storage medium according to the present invention, the computer-readable storage medium may be configured to store a program for executing the picture analysis model adjustment method in the above method embodiment, and this program can be loaded and run by a processor to implement the above picture analysis model adjustment method. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The computer-readable storage medium may be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present invention is a non-transitory computer-readable storage medium.

[0115] Furthermore, it should be understood that since the setting of each module is only for illustrating the functional units of the device of the present invention, the physical devices corresponding to these modules may be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only illustrative.

[0116] Those skilled in the art can understand that the various modules in the device can be adaptively split or combined. Such splitting or combination of specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combination will all fall within the protection scope of the present invention.

[0117] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A method for adjusting a picture analysis model, characterized in that, the method includes: Input the training picture into the model. The training picture carries annotation data indicating the category of the first detection box of the target object in the training picture. The detection head of the model includes a first branch and a second branch, and calculates the confidence and category of the second detection box for indicating the target object respectively; When the confidence of the second detection box is higher than a preset level, determine whether the category of the second detection box is the same as that of the first detection box. When they are different, set a pseudo-label for the second detection box, and record the category of the second detection box as an unknown category in the pseudo-label; Use the second detection box and the pseudo-label as the training data of the second branch, input them into the second branch, and adjust the second branch according to the output result; After the model is adjusted, prohibit the model from outputting the result with the detection box category being the unknown category.

2. The method for adjusting a picture analysis model according to claim 1, characterized in that, the step of "adjusting the second branch according to the output result" includes: Calculate the loss value of the output result according to a preset loss function, and adjust the parameters of the second branch according to the loss value.

3. The method for adjusting a picture analysis model according to claim 2, characterized in that, before the step of "after the model is adjusted, prohibit the model from outputting the result with the detection box category being the unknown category", further includes: After detecting that the loss value is less than a preset threshold, determine that the model is adjusted.

4. The method for adjusting a picture analysis model according to claim 1, characterized in that, before the step of "inputting the training picture into the model", further includes: Obtain the training picture according to the detection box category in the historical error results output by the model.

5. The method for adjusting a picture analysis model according to claim 1, characterized in that, before the step of "using the second detection box and the pseudo-label as the training data of the second branch", further includes: If the second detection box and other detection boxes are in the same connected domain, update the category of the second detection box according to the category of the other detection boxes.

6. The method for adjusting a picture analysis model according to claim 1, characterized in that, the annotation data also indicates the position of the first detection box, and the detection head further includes a third branch for calculating the position of the second detection box.

7. The method for adjusting a picture analysis model according to claim 1, characterized in that, the model further includes a feature extraction layer, and the feature extraction layer includes a backbone network and a multi-scale feature fusion network. The backbone network is used to extract multi-scale features from the training picture, and the multi-scale feature fusion network is used to fuse the multi-scale features of the training picture into the features of the training picture for inputting into the detection head.

8. A device for adjusting a picture analysis model, characterized in that, the device includes: An image input module inputs training images into the model. The training images carry annotation data indicating the category of the first detection box of the target object in the training images. The detection head of the model includes a first branch and a second branch, which respectively calculate the confidence level and category of the second detection box for indicating the target object; A category setting module, when the confidence level of the second detection box is higher than a preset level, determines whether the category of the second detection box is the same as the category of the first detection box. When they are different, a pseudo-label is set for the second detection box, and the category of the second detection box is recorded as an unknown category in the pseudo-label; A branch adjustment module inputs the second detection box and the pseudo-label as the training data of the second branch into the second branch, and adjusts the second branch according to the output result; An output control module, after the model is adjusted, prohibits the model from outputting a result with the detection box category being the unknown category.

9. A control device, comprising a processor and a storage device, wherein the storage device is adapted to store multiple program codes, characterized in that, the program codes are adapted to be loaded and run by the processor to execute the image analysis model adjustment method according to any one of claims 1 to 7.

10. A computer-readable storage medium, in which multiple program codes are stored, characterized in that, the program codes are adapted to be loaded and run by a processor to execute the image analysis model adjustment method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Smoke and fire detection method and system based on deep learning convolutional neural network

    CN111680632A

  • X-ray image target detection method and device

    CN111860510A