A method for finding a threshold in binary image classification
By calculating the image classification threshold based on the 3σ principle, the problem of unreasonable setting of thresholds in the prior art relying on manual experience is solved, and the accuracy of the model and the universality of the threshold are improved.
Patent Information
- Application Number
- CN202111370900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-18
AI Technical Summary
When setting image classification thresholds, the prior art often relies on manual experience, resulting in unreasonable threshold settings and affecting the accuracy of the model.
By preparing the test set data, statistics of the confidence scores, and using symmetric completion and normal distribution of negative samples, a reasonable threshold is calculated based on the 3σ principle.
The accuracy of the model is improved, making the set threshold more versatile and can distinguish image categories more scientifically and effectively.
Smart Images

Figure CN116152538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent video processing, and particularly relates to a method for finding a threshold value in binary image classification. Background Art
[0002] In the field of computer vision, it is often necessary to determine whether an object in an image belongs to a certain category. The common method is to use a neural network to output a score between 0 and 1, and the magnitude of the score is used as the confidence for judgment. Then, a classification threshold is set to distinguish the two categories.
[0003] The Sigmoid activation function is commonly used at the output end of the neural network, which can constrain the output of the neural network between 0 and 1 and be used as the confidence for image classification.
[0004] The 3σ principle. In a normal distribution, σ represents the standard deviation and μ represents the mean. The 3σ principle is that the probability of the numerical distribution in (μ - σ, μ + σ) is 0.6826; the probability of the numerical distribution in (μ - 2σ, μ + 2σ) is 0.9544; the probability of the numerical distribution in (μ - 3σ, μ + 3σ) is 0.9974. It can be considered that the numerical values are almost all concentrated in the interval (μ - 3σ, μ + 3σ), and the possibility of exceeding this range only accounts for less than 0.3%.
[0005] The existing methods for setting the image classification threshold often rely on manual experience and set the threshold to 0.5. If it is greater than 0.5, it is considered to be one category, and vice versa for the other category. Summary of the Invention
[0006] In order to solve the problems in the prior art, the purpose of the method of the present application is to provide a scientific and effective way to reasonably set the classification threshold.
[0007] Specifically, the present invention provides a method for finding a threshold value in binary image classification, and the method includes the following steps:
[0008] S1. Prepare the test set data:
[0009] Positive samples are pictures belonging to the target category, and negative samples are pictures not belonging to the target category; prepare test sets of positive samples and negative samples with the same quantity; and the collected positive samples and negative samples should come from as many scenes as possible to ensure the universality and generality of the data.
[0010] The so-called target category means that the binary classification model processes pictures to see if they belong to a certain category or if there is a certain object in the pictures. The category belonging to this class or containing this object is called the target category.
[0011] S2. Statistically analyze the confidence scores:
[0012] Using open-source neural networks, such as the lightweight and fast mobile-net, and the well-performing classification resnet as the classification model. Since neural networks can approximate any function, the classification model can be abstractly represented by a mathematical formula as:
[0013] y = f(x)
[0014] where x is the image, and y represents the eigenvalue output by the network, with a range of (-∞, +∞). Then, the sigmoid activation function is used to map the eigenvalue y to a value between 0 and 1, which is the confidence score. Ideally, the score for a positive sample is 1, and the score for a negative sample is 0. The mathematical expression of sigmoid is as follows:
[0015]
[0016] When each test image is input into the model, the classification confidence of this image will be obtained, and the confidence scores are tested and recorded on two test sets respectively;
[0017] S3. Calculate the threshold:
[0018] Using symmetric completion and the normal distribution of negative samples, the mean of the confidence scores of the augmented positive samples is 1, and the standard deviation is denoted as σ1; the mean of the confidence scores of negative samples is 0, and its standard deviation is denoted as σ2; according to the 3σ principle: the threshold is set to: threshold = (3σ2 + 1 - 3σ1) / 2;
[0019] It can be seen from the formula that when σ1 and σ2 are equal, that is, when the confidence score distributions of the model on positive and negative samples are the same, threshold = 0.5 is the most ideal threshold.
[0020] The same quantity mentioned in step S1 is not less than 1000.
[0021] In step S2, the two results follow a truncated normal distribution. The confidence distribution of negative samples is the right half solid curve of its normal distribution curve, that is, the leftmost distribution curve in the figure. The confidence distribution of positive samples is the left half solid curve of its normal distribution curve, that is, the rightmost curve in the figure.
[0022] Therefore, the advantage of this application is that the threshold set by this method is more general and can improve the accuracy of the model. Brief Description of the Drawings
[0023] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0024] Figure 1 It is a flowchart of the method of this application.
[0025] Figure 2 It is a schematic diagram of an embodiment in this application. Specific embodiments
[0026] In order to more clearly understand the technical content and advantages of the present invention, the present invention will now be further described in detail with reference to the accompanying drawings.
[0027] By statistically analyzing the confidence scores of the classification model on positive and negative samples, it can be known that the confidence distributions of the model on positive and negative samples are not the same. The method of subjectively setting the classification threshold to 0.5 is actually unreliable. This application proposes a scientific and effective method to reasonably set the classification threshold. The present invention belongs to the technical field of intelligent video processing and relates to binary classification in image recognition. As Figure 1 shown, this application proposes a method for finding a suitable threshold for binary image classification, and the main implementation steps of the method are as follows:
[0028] S1. Prepare test set data:
[0029] Positive samples are pictures belonging to the target category, and negative samples are pictures not belonging to the target category; prepare test sets of positive and negative samples with the same number; and the collected positive and negative samples should come from as many scenarios as possible to ensure the universality and generality of the data.
[0030] The so-called target category means that when the binary classification model processes pictures, it checks whether they belong to a certain category or whether there is a certain target in the picture. The category that belongs to this category or contains this target is called the target category.
[0031] S2. Statistically analyze the confidence scores:
[0032] Use open-source neural networks, such as the lightweight and fast mobile-net and the high-performance resnet with good classification effects, as the classification model. Since a neural network can fit any function, the classification model can be abstractly expressed by the mathematical formula:
[0033] y = f(x)
[0034] where x is the picture, and y represents the eigenvalue output by the network, with a range of (-∞, +∞). Then, the sigmoid activation function is used to map the eigenvalue y to a value between 0 and 1, which is the confidence score. Ideally, the score for positive samples is 1, and the score for negative samples is 0. The mathematical expression of sigmoid is as follows:
[0035]
[0036] When each test image is input into the model, the classification confidence of this image will be obtained, and the confidence scores are tested and recorded on two test sets respectively; the same quantity mentioned in step S1 is not less than 1000 images.
[0037] As Figure 2 shown, in step S2, the two results present a truncated normal distribution. The confidence distribution of negative samples is the right half solid curve of its normal distribution curve, that is, the left distribution curve in the figure. The confidence distribution of positive samples is the left half solid curve of its normal distribution curve, that is, the right curve in the figure.
[0038] S3. Calculate the threshold:
[0039] Using symmetric complementation and the normal distribution of negative samples, as Figure 2 shown, the mean of the confidence scores of the expanded positive samples is 1, and the calculated standard deviation is denoted as σ1; the mean of the confidence scores of negative samples is 0, and the calculated standard deviation is denoted as σ2; according to the 3σ principle:
[0040] Set the threshold to: threshold = (3σ2 + 1 - 3σ1) / 2
[0041] It can be seen from the formula that when σ1 and σ2 are equal, that is, when the confidence score distributions of the model on positive and negative samples are the same, threshold = 0.5 is the most ideal threshold.
[0042] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for finding a threshold in binary image classification, characterized in that, the method comprises the following steps: S1. Prepare test set data: Positive samples are pictures belonging to the target category, and negative samples are pictures not belonging to the target category; prepare test sets of positive and negative samples with the same number; and the collected positive and negative samples should come from multiple scenarios; S2. Statistically calculate confidence scores: Using an open-source neural network, since a neural network can fit any function, the classification model can be abstractly expressed by a mathematical formula as: y = f(x) where x is the picture and y represents the eigenvalue output by the network, with a range of (-∞, +∞); then use the sigmoid activation function to map the eigenvalue y to a value between 0 and 1, which is the confidence score; where the mathematical expression of sigmoid is as follows: Each test picture input to the model will obtain the classification confidence of this picture, and the confidence scores are tested and recorded on two test sets respectively; S3. Calculate the threshold: Using symmetric complementation and the normal distribution of negative samples, the mean of the confidence scores of the augmented positive samples is 1, and the calculated standard deviation is denoted as σ1; the mean of the confidence scores of negative samples is 0, and the calculated standard deviation is denoted as σ2; according to the 3σ principle: set the threshold to: threshold = (3σ2 + 1 - 3σ1) / 2.
2. The method for finding a threshold in binary image classification according to claim 1, characterized in that, the "same number" in step S1 is not less than 1000.
3. The method for finding a threshold in binary image classification according to claim 1, characterized in that, in step S2, ideally, the score of positive samples is 1 and the score of negative samples is 0.
4. The method for finding a threshold in binary image classification according to claim 1, characterized in that, in step S2, the two results show a truncated normal distribution, the confidence distribution of negative samples is the right half curve of its normal distribution curve, and the confidence distribution of positive samples is the left half curve of its normal distribution curve.
5. The method for finding a threshold in binary image classification according to claim 1, characterized in that, when σ1 and σ2 are equal, that is, when the confidence score distributions of the model on positive and negative samples are the same, threshold = 0.5 is the most ideal threshold.
Citation Information
Patent Citations
Light-adaptive facial recognition method and system
CN106295571A
General target detection method of adaptive attention guidance mechanism
CN111259930A