A method for eliminating false alarms in pet detection
By adding head categories to the pet detection network and adjusting the yolov3 framework structure, the problem of false alarms of similar objects in pet detection is solved, and the false alarm elimination effect is achieved without affecting the recall rate and accuracy.
Patent Information
- Application Number
- CN202110312560.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-03-24
AI Technical Summary
The prior art is difficult to eliminate false alarms of similar objects in pet detection without affecting recall and accuracy, such as dog false alarms in cat detection, dog false alarms in cat face detection and head false alarms in big head photos.
By adding head categories to the pet detection network, adjusting network structure and training, filtering head categories in the detection results, using the yolov3 framework for target detection, and adding class_id judgment to eliminate false alarms.
Without affecting pet detection recall and accuracy, it is cleverly eliminated false alarms of similar objects that cannot be handled by conventional methods.
Smart Images

Figure CN115131644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent video processing, and particularly relates to a method for eliminating false alarms in pet detection. Background Art
[0002] With the development of computer technology and the wide application of computer vision principles, it has become increasingly popular to use computer image processing technology to detect and track targets in real time. Dynamic real-time tracking and positioning of targets are used in intelligent transportation systems, intelligent monitoring systems, and military target detection, and it has broad application value to position surgical instruments in medical navigation surgery. The task of target detection is to find all the targets of interest in an image and determine their positions and sizes, which is one of the core problems in the field of machine vision. Due to the different appearances, shapes, postures of various objects, and the interference of factors such as illumination and occlusion during imaging, target detection has always been one of the most challenging problems in the field of machine vision.
[0003] Nowadays, keeping cats and dogs as pets is becoming increasingly popular among the public, and the related demands have emerged one after another. At present, cat and dog detection technologies are widely used. Cat and dog detection can be used in pet families, such as cat and dog feeders. When a cat or dog approaches the cat and dog feeder and the feeder detects the cat or dog, the lid will automatically open to facilitate the cat or dog to eat, etc.
[0004] Most of the common methods for reducing false alarms in existing target detection are: increasing negative sample data. Increasing negative sample data can be divided into two categories. One is to directly put some diverse pictures without targets into the training data; the other is to use an existing model to detect pictures without targets, and there will be some pictures with detection errors. These detected pictures are used as negative samples and put into the training data. In addition to increasing negative sample data, false alarms can also be reduced by improving the quality of positive samples, and the training samples need to be re-screened to filter out bad samples. Of course, a deeper and more complex network can also be used, but this will change the network structure, and accordingly, the detection time will also increase.
[0005] Existing methods for eliminating false alarms, such as increasing negative sample data, can solve some generalized false alarms, but increasing negative samples will inevitably lead to a decrease in the recall rate; screening positive samples and improving the sample quality. For example, the misdetection of people in cats and dogs is because some of the positive samples of the framed cats and dogs contain people. Just adjusting the positive samples to avoid human annotation can eliminate the false alarms of people. This method will reduce the diversity of positive samples, resulting in a decrease in the recall rate; using a deeper network will lead to changes in the network, increase the detection time, and increase the size of the model; and these traditional methods cannot eliminate false alarms of similar objects, such as false alarms of dogs in cat detection, false alarms of human faces in cat face and dog face detection, false alarms of human heads in mugshots. Even adding negative samples of people cannot eliminate the misdetection. Summary of the Invention
[0006] This application solves the above problems. The purpose of this application is to add a category to accommodate false positives, and perfectly eliminate false positives of humans in cat and dog detection without affecting the recall rate.
[0007] Specifically, the present invention provides a method for eliminating false positives in pet detection, and the method includes the following steps:
[0008] S1, Add head annotation: Add the category of human head, mark the human heads in the training samples, and only mark the front;
[0009] S2, Modify the network to add misdetection target classification: The pet detection network is improved based on the object detection framework. Increase the value of the classes attribute in the cfg in the configuration file by one, increase the value of the classes attribute in the data file in the data configuration file by one, and add the name attribute of the misdetection class in the name file in the data configuration file;
[0010] S3, Train the model: After the above two steps, the data of adding human heads and adding classification are completed, and the data can be trained. Since a new category is added, the previous pet class model can be loaded during training;
[0011] S4, Detect the model and adjust: Test the trained model. If it affects the pet detection accuracy, the human head samples can be appropriately reduced; if the elimination effect of human head false positives is not good, increase the human head samples appropriately without affecting the pet detection accuracy to adjust the data until the false positives are eliminated without affecting the accuracy;
[0012] S5, Filter false positive classes in the detection results: Since the added human heads are added to eliminate false positives, the detection results should filter out this class, find the class_id of this class, and filter out this class by judgment.
[0013] In step S2, the pet detection network can adopt the yolov3 framework. The yolov3 network structure is composed of a series of 1x1 and 3x3 convolutional layers. After each convolutional layer, there will be a BN layer and a LeakyReLU. There are 53 convolutional layers in the network. The default strides of the convolution are (1, 1), and the default padding is same. When the strides are (2, 2), the padding is valid; This method filters misdetection targets according to the class_id in 3 predict layers.
[0014] The yolov3 framework is an open-source deep learning framework for multi-category object detection.
[0015] The 53 convolutional layers are 2+1*2+1+2*2+1+8*2+1+8*2+1+4*2+1 = 53. When counted in order, the last Connected is a fully connected layer and is also counted as a convolutional layer, making a total of 53 layers.
[0016] The 3 predict layers output the four position coordinates of the target, a score for one object target, and a classification score: Obtain the class_id.
[0017] For the pet detection, the human head class is added, increasing the mis-detected target of the human head class. Set class_id = 2. Before obtaining the result, add a judgment on whether class_id is 2. If it is 2, then filter it out. In this way, false alarms can be removed, that is, separate first and then filter.
[0018] The specific steps to add classification to the YOLOv3 framework are as follows:
[0019] S2.1, increment the values of the two classes attributes in the cfg of the configuration file by one. If the original was classes = 2, then it becomes classes = 3 now; and increment the filters attribute value by 3. Since three anchors are taken, for each additional class, 3 is added. The change formula for classes is classes = total number of classes, and filters = (4 + 1 + 1) * 3, that is, 4 coordinates plus one score plus one classification and then multiplied by the number of anchors taken in the result layer;
[0020] S2.2, increment the value of the classes attribute in the data file of the data configuration file by one. The change formula for classes is classes = total number of classes;
[0021] S2.3, add the name attribute of the mis-detected class in the name file of the data configuration file. Here, one class, the person class, is added.
[0022] Thus, the advantage of this application is that: without affecting the existing accuracy, it cleverly eliminates the similar false alarms that cannot be eliminated by conventional methods. Specifically, it combines the idea of classification and the elimination of false alarms, and cleverly uses the classification method to eliminate the false alarms of similar targets that cannot be eliminated by conventional methods without affecting the recall rate and accuracy of the original detection targets. Description of the Drawings
[0023] The drawings described here are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.
[0024] Figure 1 It is a schematic diagram of the detection process of cats and dogs in the prior art.
[0025] Figure 2 It is a schematic diagram of the cat and dog detection process in the present invention.
[0026] Figure 3 It is a flowchart of the method of the present invention.
[0027] Figure 4 It is a structural diagram of the yolov3 network used in the present invention.
[0028] Figure 5 is Figure 4 a further structural diagram of the Convolutional Set part in Specific Embodiments
[0029] In order to more clearly understand the technical content and advantages of the present invention, the present invention will now be further described in detail with reference to the accompanying drawings.
[0030] As Figure 1 shown, the steps of the cat and dog detection process in the prior art taking cats and dogs as examples include: starting, obtaining data, data augmentation, entering the network, obtaining results, and judging whether class_id == 0? If it is YES, then it is judged as a cat and the process ends; if it is NO, then it is judged as a dog and the process ends. In fact, class_id is the pet classification, which is defined during data annotation. When there is one class, it can be not written. When there are multiple classes, it needs to be defined, starting from 0. In this application, the setting of classid is that a cat is 0, a dog is 1, and a human head is 2. These are all artificially set for the purpose of distinguishing different types of targets.
[0031] As Figure 2 shown, it is a specific embodiment of the method of the present invention, still taking cats and dogs as examples. The cat and dog detection adds the human head class (class_id = 2). The idea of this application is reflected in adding the misdetection target human head class and adding a judgment on whether class_id is 2 before obtaining the results to filter it out, so as to remove false alarms, that is, separate first and then filter. The specific process steps include: starting, obtaining data, data augmentation, entering the network, and judging whether class_id == 2? If it is YES, then remove this data; if it is NO, then obtain the results, and then judge whether class_id == 0? If it is YES, then it is judged as a cat and the process ends; if it is NO, then it is judged as a dog and the process ends.
[0032] As Figure 3 shown, it is a method for eliminating false alarms in pet detection, and the method includes the following steps:
[0033] S1, adding human head annotations: adding the human head class, marking the human heads in the training samples, and only marking the front;
[0034] S2. Change the network to add misdetection target classification: The pet detection network is improved based on the object detection framework. Increment the value of the "classes" attribute in the cfg of the configuration file by one, increment the value of the "classes" attribute in the data file of the data configuration file by one, and add the name attribute of the misdetection class in the name file of the data configuration file;
[0035] S3. Train the model: After the above two steps, the data of adding human heads and adding classifications are completed, and the data can be trained. Since one more class is added, the previous pet class model can be loaded during training. When loading the model, remove the result layer of the pre-trained model, that is, do not load the result layer of the pre-trained model;
[0036] S4. Detect the model and adjust: Test the trained model. If it affects the pet detection accuracy, the human head samples can be appropriately reduced; if the elimination effect of human head false alarms is not good, increase the human head samples appropriately under the condition of ensuring that the pet detection accuracy is not affected, so as to adjust the data until the false alarms are eliminated without affecting the accuracy;
[0037] S5. Filter the misdetection class from the detection results: Since the added human heads are added to eliminate misdetections, the detection results should filter out this class. Find the class_id of this class and filter out this class by judgment.
[0038] Furthermore, the method steps include:
[0039] Step S1. Add human head annotations
[0040] In cat and dog detection, a large number of false alarms of frontal human heads occur. Since the probability of humans appearing is relatively high, this false alarm must be eliminated. Since the class of human heads is added, the human heads in the training samples need to be annotated, and only the frontals need to be annotated. Otherwise, they will be treated as negative samples. Note that the number of added human head samples should not be too large, because it is only to eliminate human false alarms, so high precision is not required, and normal human detections are sufficient. This can not only ensure the accuracy of cat and dog detection, but also eliminate human false alarms.
[0041] Step S2. Change the network to add misdetection target classification
[0042] The current cat and dog detection network is improved based on the yolov3 framework. Therefore, taking the yolov3 framework (yolov3 is an open-source deep learning framework for multi-class object detection, and its network structure diagram is as Figure 4 ) as an example, this method is not limited to the yolov3 framework, and all object detection frameworks can be used. The addition method varies according to different frameworks. The specific steps to add classifications in the yolov3 framework are as follows:
[0043] S2.1 Increment the values of the two "classes" attributes in the cfg of the configuration file by 1. The original "classes = 2" now becomes "classes = 3"; and increment the value of the "filters" attribute by 3. Since three anchors are taken, 3 is added for each additional class. If there are multiple mis-detected classes, they can all be regarded as one class or each class is considered as a separate category. The change formula for "classes" is "classes = total number of classes"; "filters = (4 + 1 + 1) * 3", that is, (4 coordinates + one score + one classification) multiplied by the number of anchors taken in the result layer. Here, cfg is the configuration file in the YOLOv3 framework, with the suffix ".cfg", such as yolov3.cfg, which defines the parameters related to the network.
[0044] S2.2 Increment the value of the "classes" attribute in the data file of the data configuration file by 1. The change formula for "classes" is "classes = total number of classes".
[0045] S2.3 Add the "name" attribute of the mis-detected class in the name file of the data configuration file. Add as many as there are classes. Here, one class, the "person" class, is added.
[0046] Step S3. Train the model
[0047] After the above two steps, the data for adding human heads and adding classifications are completed, and the data can be trained. Since there is one more class, the previous cat and dog models can be loaded during training. When loading the model, remove the result layer of the pre-trained model, that is, do not load the result layer of the pre-trained model.
[0048] Step S4. Detect the model and adjust
[0049] Test the trained model. If the accuracy of cats and dogs is affected, the human samples can be appropriately reduced; if the effect of eliminating human false alarms is not very good, the human samples can be appropriately increased on the premise of not affecting the detection accuracy of cats and dogs to adjust the data until the false alarms are eliminated without affecting the accuracy.
[0050] Step S5. Filter out the mis-detected classes in the detection results
[0051] Since the added human heads are added to eliminate human false alarms, the detection results should filter out this class. Just find the class id of this class and add a judgment to filter out this class.
[0052] For example Figure 4As shown, it is the YOLOv3 network structure diagram adopted in this embodiment. This network is mainly composed of a series of 1x1 and 3x3 convolutional layers (each convolutional layer is followed by a BN (Batch Normalization) layer and a LeakyReLU). There are 53 convolutional layers in the network, (2 + 1*2 + 1 + 2*2 + 1 + 8*2 + 1 + 8*2 + 1 + 4*2 + 1 = 53. Counting in order, the last Connected is a fully connected layer and is also considered a convolutional layer, a total of 53). The strides of the convolution are defaulted to (1, 1), and the padding is defaulted to same. When the strides are (2, 2), the padding is valid. This method filters out false positive targets according to the classid at 3 predict layers (the predict layer outputs the four position coordinates of the target, a score for an object target, and a classification score: obtain the class_id).
[0053] Among them, as Figure 5 shown, the Convolutional Set part is further as follows: from Convolutional 1×1 to Convolutional 3×3 in sequence, then to Convolutional 1×1, then to Convolutional 3×3, and then to Convolutional 1×1.
[0054] The method for eliminating false positives in this method originated from cat and dog detection and can be extended to the elimination of similar false positives in all object detections.
[0055] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for eliminating false positives in pet detection, characterized in that, The method includes the following steps: S1. Add head annotation: Add the category of human heads, mark the human heads in the training samples, and only mark the front view; S2. Modify the network to add misdetection target classification: The pet detection network is improved based on the object detection framework. Increment the value of the "classes" attribute in the cfg of the configuration file by one, increment the value of the "classes" attribute in the data file of the data configuration file by one, and add the "name" attribute of the misdetection class in the name file of the data configuration file; S3. Train the model: After the above two steps, the data for adding human heads and the classification are completed, and then the data can be trained. Since one more class is added, the previous pet class model is loaded during training; S4. Detect the model and make adjustments: Test the trained model. If the pet detection accuracy is affected, appropriately reduce the human head samples; if the elimination effect of human head false alarms is not good, appropriately increase the human head samples on the premise of not affecting the pet detection accuracy to adjust the data until the false alarms are eliminated without affecting the accuracy; S5. Filter the misdetection class from the detection results: Since the added human heads are added to eliminate misdetections, the detection results should filter out this class. Find the class_id of this class and filter it out by judgment.
2. The method for eliminating false positives in pet detection according to claim 1, wherein, In step S2, the pet detection network uses the yolov3 framework. The network structure of the yolov3 framework consists of a series of 1x1 and 3x3 convolutional layers. After each convolutional layer, there is a BN layer and a LeakyReLU. There are 53 convolutional layers in the network. The default value of the convolutional stride is (1, 1), and the default value of the padding is "same". When the stride is (2, 2), the padding is "valid". In this method, misdetection targets are filtered according to the class_id at 3 "predict" layers.
3. A method for eliminating false alarms in pet detection according to claim 2, characterized in that The yolov3 framework is an open-source deep learning framework for multi-class object detection.
4. A method for eliminating false alarms in pet detection according to claim 2, characterized in that The 53 convolutional layers are 2 + 1*2 + 1 + 2*2 + 1 + 8*2 + 1 + 8*2 + 1 + 4*2 + 1 = 53. Counting in order, the last "Connected" is a fully connected layer and is also counted as a convolutional layer, with a total of 53 layers.
5. A method for eliminating false alarms in pet detection according to claim 2, characterized in that The 3 "predict" layers output the four position coordinates of the target, an object target score, and a classification score: Obtain the class_id.
6. A method for eliminating false alarms in pet detection according to claim 2, characterized in that The pet detection includes human heads, which increases the misdetected target human heads. Set class_id = 2, and add a judgment on whether class_id is 2 before obtaining the result. If it is 2, then filter it out. In this way, false alarms can be removed, that is, separate first and then filter.
7. A method for eliminating false alarms in pet detection according to claim 2, characterized in that The specific steps for adding classification to the yolov3 framework are as follows: S2.1, increment the values of the two classes attributes in the cfg of the configuration file by one. If the original is classes = 2, then it becomes classes = 3; and increment the filters attribute value by 3. Since three anchors are taken, each additional class adds 3. The change formula for classes is classes = total number of classes, and filters = (4 + 1 + 1) * 3, that is, 4 position coordinates plus one object target score plus one classification score and then multiplied by the number of anchors taken in the result layer; S2.2, increment the value of the classes attribute in the data file of the data configuration file by one. The change formula for classes is classes = total number of classes; S2.3, add the name attribute of the misdetected class in the name file of the data configuration file. Here, one class, namely the person class, is added.
8. A method for eliminating false alarms in pet detection according to claim 7, characterized in that In step S2.1, if there are multiple misdetected classes, they are all regarded as one class or each class is regarded as a separate class.
9. A method for eliminating false alarms in pet detection according to claim 7, characterized in that In step S2.3, if multiple misdetected classes are added, then add the name attributes of the misdetected classes according to the number of classes.
10. A method for eliminating false alarms in pet detection according to claim 1, characterized in that, The pet detection is for cat and dog detection.
Citation Information
Patent Citations
Target detection method and device and storage medium
CN111814832A
Target detection method and device, terminal equipment and computer readable storage medium
CN111914863A