Image classification method in self-adaptive environment during open set test
Through the method of dual-pattern matching and pairing distribution difference loss, the problem of identifying samples of unknown categories by deep neural networks in open-set testing environment is solved, efficient image classification in complex scenarios is achieved, and the adaptability and accuracy of the model is improved.
Patent Information
- Application Number
- CN202510527104.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-05
AI Technical Summary
In open-set testing environments, deep neural network models are difficult to identify unknown categories of samples, resulting in a decrease in classification accuracy and reliability, especially in complex scenarios such as biodiversity monitoring and airport foreign object detection.
The open set sample identification strategy based on dual-pattern matching and the paired distribution difference loss are adopted. By accurately identifying open set samples and optimizing model characterization learning, the model's ability to distinguish open set samples is improved and misclassification is avoided.
Quickly and effectively identify closed and open set samples in streaming data scenarios, improve the accuracy of image classification, avoid model degradation, and is suitable for scenarios such as medical diagnosis and industrial defects.
Smart Images

Figure CN120431388A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image classification and test-time adaptation, and in particular relates to an image classification method in an open set test-time adaptive environment. Background Art
[0002] In the field of image classification, deep neural networks have achieved remarkable results thanks to their powerful feature extraction and pattern recognition capabilities. This technology is widely used in scenarios ranging from object classification in everyday photos to disease diagnosis in medical imaging to traffic sign recognition. However, in practical applications, deep neural networks have exposed significant vulnerabilities. Imagine an image classification scenario for security surveillance. If the camera is disturbed by factors such as lighting fluctuations and dust obstruction, the captured image will become noisy or blurry. The accuracy of the target classifier based on the deep neural network (for example, distinguishing pedestrians from vehicles) will be significantly reduced. In autonomous driving, if the image data distribution of road signs changes due to stains, wear, etc., the image classification model on the vehicle may misjudge traffic signs, posing a serious safety hazard. Therefore, enhancing the robustness of deep models to changes in the test data distribution is key to ensuring the reliable operation of image classification applications.
[0003] Test-time adaptation technology provides a new opportunity to solve this problem. When image classification models are deployed in real-world scenarios (such as intelligent monitoring systems and self-driving cars), it can help the model cope with changes in data distribution. Traditional test-time adaptation methods mainly address the problem of covariate shift. For example, in flower classification applications, the source domain is flower images taken in a greenhouse, and the target domain is flower images taken in the wild. The flower categories of the two are the same, but the image data distribution is different due to factors such as lighting and background. Traditional methods aim to transfer the classification knowledge learned on greenhouse images to field images to reduce this distribution difference.
[0004] However, in more complex image classification scenarios, the problem goes far beyond this. For example, in the image classification task of biodiversity monitoring, the source domain data is common animal images, and new species may appear in the target domain (categories not covered by the source domain). This not only causes covariate shift, but also semantic shift. In this open set test environment, traditional adaptive methods become powerless during testing. Taking the airport's foreign object detection image classification system as an example, if a new type of foreign object appears (a category not covered by the training data), the traditional method will not allow the model to effectively identify it, which may lead to model misjudgment and cause safety risks.
[0005] Adaptation during open set testing requires the model to have the ability to recognize and process unknown categories. Otherwise, in practical applications, as samples of unknown categories continue to appear, the model performance will continue to deteriorate, seriously affecting the accuracy and reliability of image classification. Summary of the Invention
[0006] The purpose of the present invention is to propose an image classification method in an adaptive environment during open set testing. The method includes an open set sample recognition strategy based on dual-mode matching and a paired distribution difference loss. Through accurate open set sample recognition and robust open set representation learning, the model's ability to distinguish open set samples is improved. When each test batch arrives, the open set samples are quickly identified through the dual-mode matching strategy. This precise recognition capability prevents the model from mistakenly classifying open set samples into known categories to avoid the occurrence of misclassification. By optimizing the paired distribution difference loss, the model has a stronger adaptability to those open set samples that are difficult to distinguish, which is conducive to greatly improving the accuracy of image classification.
[0007] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0008] A method for image classification in an adaptive environment during open set testing, comprising the following steps:
[0009] Step 1. Obtain the pre-trained image classification model f θ ; where θ represents the model parameters; get the test data set D represents the image x i The unlabeled dataset N test is the total number of samples in the test dataset;
[0010] Step 2. Initialize the image classification model f θ Hyperparameters of
[0011] Step 3. At each time step t, receive a batch of test data;
[0012] Step 4. Based on the pattern matching strategy of model parameter update, starting from the perspective of gradient update of model parameters, an open set score ods(x) is constructed according to the model's prediction confidence to reflect the degree of certainty of the model's image classification results. By analyzing the relationship between gradient and model parameter update, combined with the open set score constructed by prediction confidence, a preliminary judgment is made on the open set samples;
[0013] Step 5. Based on the pattern matching strategy of feature space change, the position of the test image in the feature space and the difference in feature distribution with known category images are analyzed to determine whether it is an open set sample;
[0014] Step 6. Based on the intersection of the pattern matching strategies proposed in Step 4 and Step 5, perform final identification of open set samples and closed set samples, and perform entropy minimization loss only on the identified closed set samples;
[0015] Step 7. Design a pairwise distribution difference loss based on visual cue words, aiming to utilize information from easy open-set samples to help the model learn more robust representations to identify difficult open-set samples and better distinguish between open-set and closed-set samples.
[0016] Step 8. Calculate the model's loss function and update the model parameters until all data are tested; input the images to be classified obtained in real time into the tested image classification model to achieve image classification.
[0017] The present invention has the following advantages:
[0018] As described above, the present invention describes an image classification method in an adaptive environment during open set testing, in which an open set sample recognition strategy based on dual-mode matching and a paired distribution difference loss are designed. Among them, the open set sample recognition strategy based on dual-mode matching can quickly and efficiently identify reliable closed set samples and potential open set samples in a non-iterative manner in a scenario where data arrives in a streaming manner. This precise recognition capability prevents the model from mistakenly classifying open set samples into known categories, thereby avoiding the occurrence of misclassification. The paired distribution difference loss based on visual cue words can improve the distribution difference between open set categories and closed set categories, solve the problem of identifying difficult open set samples, and make the model more adaptable to those open set samples that are difficult to distinguish, which is conducive to greatly improving the accuracy of image classification. The method of the present invention can quickly and effectively identify open set samples at the moment when each data batch arrives during testing, while improving the discrimination of subsequent batch models for open set samples, thereby avoiding model degradation caused by open set samples. Significant improvement is shown in multiple adaptive scenarios during open set testing such as medical diagnosis and industrial defects, proving its feasibility in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 4 is a flowchart of an image classification method in an adaptive environment during open set testing according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] Glossary:
[0021] At the core of deep learning are artificial neural networks, which consist of multiple layers. Each layer contains multiple neurons, and the neurons in each layer have connection weights. Through the input of large amounts of data and repeated iterative training, the neural network can automatically adjust these weights, thereby realizing pattern recognition and feature extraction from the data.
[0022] Open set detection is an important research area in deep learning. Traditional deep learning classification tasks typically assume that all categories in the test data have also appeared in the training dataset, a concept known as the closed set assumption. In practical applications, we often encounter situations where the training data does not cover all possible categories, i.e., open set data exists. Open set detection aims to identify these unknown samples that do not belong to the known categories in the training set, while accurately classifying samples of known categories.
[0023] Transfer learning is a deep learning method that applies knowledge or experience learned from one task to another related task to improve learning outcomes. Typically, transfer learning involves transferring knowledge learned in a source domain to a target domain through some means. Transfer learning primarily addresses issues such as insufficient data or a lack of labels in the target domain.
[0024] Test-time adaptation, a subset of transfer learning, involves adjusting or adapting a model to new test data to accommodate the characteristics or distribution of the new data. Specifically, test-time adaptation involves fine-tuning or adjusting parameters based on new test data after the model has been trained to improve its performance on the new data.
[0025] Test-time adaptation (TTA) is a method for adjusting models to address shifts in the test data distribution, aiming to improve the performance of models (such as image classification models) in real-world applications. Traditional TTA methods typically focus on addressing covariate shift while ignoring semantic shift, particularly when the target domain contains open-set data. Indiscriminately using this data for test-time adaptation often leads to model degradation.
[0026] This embodiment describes an image classification method in an adaptive environment during open set testing, which includes an open set sample recognition strategy based on dual-mode matching and a paired distribution difference loss. Its core idea is to improve the model's ability to distinguish open set samples through accurate open set sample recognition and robust open set representation learning. When each test batch arrives, the open set samples are quickly identified through the dual-mode matching strategy. This precise recognition capability prevents the model from mistakenly classifying open set samples into known categories, thereby avoiding the occurrence of misclassification. By optimizing the paired distribution difference loss, the model has a stronger adaptability to those open set samples that are difficult to distinguish, such as the imaging features of rare diseases. This means that when faced with complex and diverse medical images, the model can more accurately judge the difference between normal images (closed set samples) and abnormal images (open set samples) representing rare diseases, greatly improving the accuracy of disease diagnosis.
[0027] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0028] like Figure 1 As shown in Figure 1, the image classification method in the adaptive environment during open set testing includes the following steps:
[0029] Step 1. Obtain the pre-trained image classification model f θ ; where θ represents the model parameters; get the test data set D represents the image x i The unlabeled dataset N test is the total number of samples in the test dataset.
[0030] Step 2. Initialize the image classification model f θ The hyperparameters of the image classification model f θ The hyperparameters include the learning rate of the Adam optimizer, the feature mask rate ρ, the loss function balance coefficient λ, and the update momentum α.
[0031] Step 3. At each time step t, receive a batch of test data x∈R B×H×W ; Where B is the batch size, H and W are the length and width of the image.
[0032] Step 4. Based on the pattern matching strategy of model parameter update, starting from the perspective of gradient updating of model parameters, a new open set score ods(x) is constructed according to the model's prediction confidence, which can reflect the degree of certainty of the model on the image classification results. By analyzing the relationship between gradient and model parameter update, combined with the open set score constructed by prediction confidence, the open set samples can be preliminarily judged.
[0033] Since the model is updated by a whole batch of data, the entropy of the label prediction for a single sample x may not necessarily decrease. Therefore, from the perspective of gradient update of model parameters, a new open set score ods(x) is constructed based on the model's prediction confidence. The calculation formula is as follows:
[0034]
[0035] in, Represents the category probability distribution predicted by the model at time t The probability value of the i-th category; Represents the category probability distribution predicted by the model at time t+1 The probability value of the i-th category.
[0036] Indicates the probability value of the model predicting that the input sample x belongs to category c′ at time t; Indicates that at time t, find the model prediction probability from all categories C The largest category.
[0037] Is the indicator function. When i (the currently traversed category index) is exactly equal to the category index with the highest model prediction probability at time t, The value is 1 if yes, otherwise it is 0.
[0038] C represents the number of categories, and c′ is a category in the category set C.
[0039] An efficient calculation method is designed by using the distribution similarity between test batch data, and the loss is calculated indifferently by the current sample and the gradient of the model is updated. To determine whether the sample belongs to the open set sample, as shown below:
[0040] θ′ t =α*θ′ t-1 +(1-α)*θ t ;
[0041]
[0042] Among them, θ t represents the model parameters obtained by completing the gradient update at time t, α is the update momentum; θ′ t-1 Refers to the model parameters updated at time t-1; θ′ t It refers to the model parameters updated by the exponential moving average EMA at time t.
[0043] It refers to the updated parameter θ′ at time t t The model predicts the probability value of the sample belonging to the i-th category; Indicates finding the model prediction probability from all categories C at time t The largest category.
[0044] Is the indicator function. When i (the currently traversed category index) is exactly equal to the category index with the highest model prediction probability at time t, The value is 1 if yes, otherwise it is 0.
[0045] After each test batch of data arrives, the open set score is calculated directly; according to the previous inference, the ods score of the open set sample will be less than 0, while the ods score of the closed set sample will be greater than 0, constructing a preliminary open set detection formula:
[0046]
[0047] Among them, ods(x) represents the score of open set samples, which is a numerical indicator used to measure whether sample x belongs to the open set sample; X OOD Represents the open set sample space, which contains all samples that do not belong to the closed set distribution for which the model is trained, that is, samples of categories or distributions that the model has never seen.
[0048] p(x∈X OOD ) represents a probability representation used to determine whether the sample x belongs to the open set sample space X OOD The probability of; when ods(x)<0, then p(x∈X OOD )=1, indicating that the sample x is an open set sample; when ods(x)≥0, then p(x∈X OOD )=0, indicating that the sample x is a closed set sample.
[0049] Step 5. Based on the pattern matching strategy of feature space changes, the position of the test image in the feature space and the difference in feature distribution with known category images are analyzed to determine whether it is an open set sample.
[0050] For open set samples, since their characteristic patterns are inherently false, the mask has little effect on them. Therefore, by analyzing the amplitude of the characteristic pattern change, the open set and closed set samples can be effectively distinguished.
[0051]
[0052] Among them, m h,w It represents a mask matrix element that determines whether the pixel with coordinates (h, w) in the image is masked. ρ is the probability of the mask, and H and W are the width and height of the image, respectively.
[0053] Bernoulli(1-ρ) stands for Bernoulli distribution, a discrete probability distribution with only two possible outcomes (success or failure). Here, the probability of "success" is 1-ρ, where ρ is the probability of being masked, and 1-ρ is the probability of not being masked.
[0054] Represents the indicator function to determine m h,w The feature space mask is constructed by calculating whether it is 1 (not masked) or 0 (masked). (h,w) represents the coordinates of the pixel in the image, where h is the height coordinate and w is the width coordinate. The values range from 1 to the image height H and from 1 to the image width W, respectively.
[0055] Define a new closed set score ids(x), the calculation formula is as follows:
[0056]
[0057] in, Represents the category probability distribution predicted by the model The probability value of the i-th category, x⊙m, represents the element-wise multiplication of the image sample x and the mask matrix m, and the image x is processed by the mask m.
[0058] Indicates that at time t, based on parameter θ t The model predicts the probability value of the masked sample x⊙m belonging to the i-th category, Indicates that at time t, find the model prediction probability from all categories C The largest category.
[0059] Is the indicator function. When i (the currently traversed category index) is exactly equal to the category index with the highest model prediction probability at time t, The value is 1 if yes, otherwise it is 0.
[0060] C represents the number of categories, and c′ refers to a category in the category set C.
[0061] The median of the ids score of the entire batch data X is selected as the dynamic threshold, and the ids score of a single sample x is calculated as:
[0062]
[0063] Here, median(ids(X)) represents the median of the ids score of the entire batch data X, that is, the dynamic threshold; ids(x) represents the score of closed set samples, which is a numerical indicator used to measure whether sample x belongs to the closed set sample.
[0064] X ID represents the closed sample space. p(x∈X ID ) represents a probability representation used to determine whether the sample x belongs to the closed sample space X ID When ids(x)>median(ids(X)), it means that sample x belongs to the closed set sample; otherwise, when ids(x)≤median(ids(X)), it means that sample x belongs to the open set sample.
[0065] Step 6. Based on the intersection of the pattern matching strategies proposed in Steps 4 and 5, the final identification of open set samples and closed set samples is performed, and entropy minimization loss is performed only on the identified closed set samples. Entropy minimization loss can make the model's prediction distribution on closed set samples more concentrated, thereby improving the classification accuracy of known categories.
[0066] In order to combine the open set sample recognition strategies of the two pattern matching, using their intersection as a measure, for any data batch X arriving at time t t , the identified closed set sample set X′ ID and open sample set X′ OOD They are:
[0067] X′ ID ={x|ids(x)>median(ids(X t ))∩ods(x)≥0};
[0068] X′ OOD ={x|ids(x)≤median(ids(X t )∩ods(x)<0)}.
[0069] In order to achieve test-time adaptation, the entropy minimization loss L is performed only on the identified closed set samples. en , the calculation formula is as follows:
[0070]
[0071] Among them, x i ∈X′ ID Closed set of samples. x i refers to X′ ID The i-th sample in , N = |X′ ID |, refers to the closed set sample X′ ID The number of samples in . C′ is the number of closed set categories, f(x i ) represents the model's response to sample x i The prediction result is usually a category probability distribution vector, h(f(x i )) c It means the model predicts sample x i The probability value of belonging to category c.
[0072] Step 7. Design a paired distribution difference loss based on visual cue words, aiming to utilize the information in simple open-set samples to help the model learn more robust representations to identify difficult open-set samples, better distinguish between open-set samples and closed-set samples, and lay the foundation for subsequent image classification.
[0073] Define a visual cue word ∈, where ∈ is a learnable parameter matrix whose size is equal to the image x; for any image x, using a visual cue word is equivalent to adding the two on x, that is: x = x + ∈.
[0074] For difficult open set samples x∈X OOD , where X OODRepresents the open set sample space. The ultimate optimization goal of the prompt word design is to ID , X ID Denotes the closed set sample space. The definition of difficult open set samples is as follows:
[0075] 1 T |f(x+∈)-f(x′+∈)|>δ1;
[0076] 1 T |h(f(x+∈))-h(f(x′+∈))|>δ2.
[0077] Among them, f(x+∈) is the prediction result of the model for the closed set sample x, which is a category probability distribution vector; f(x′+∈) is the prediction result of the model for the open set sample x, which is a category probability distribution vector; h(f(x)) refers to the probability value of the model prediction test sample x. T It refers to the transpose of the all-ones vector 1 and is used for matrix multiplication to aggregate the differences between vectors. δ1 and δ2 are the thresholds for distinguishing between open and closed set samples.
[0078] In order to distinguish difficult open-set samples in images, the results of open-set sample recognition based on dual-pattern matching are utilized.
[0079] Assume that at each time t, the open set sample set in the current batch data is identified and is X ID The subset of X′ OOD , a closed sample set and is X OOD The subset of X′ ID ; It is believed here Among them, X ID is a closed set of samples, X OOD is an open set sample set.
[0080] For a batch of samples {x1,x2,...,x b}, first apply the visual cue word ∈ to obtain batch samples {x1+∈,x2+∈,...,x b +∈}; the category probability distribution of each sample is represented as a matrix is a matrix.
[0081] Where b is the number of batch samples (i.e. the number of samples), C is the number of categories. Each row of this matrix corresponds to the category probability distribution of a sample; the i-th row is the sample x i+∈ The category probability distribution of :
[0082] P=[p1;p2;...;p b ].
[0083] Among them, P is a matrix used to represent the category probability distribution of batch samples; b represents the number of batch samples. Each row p i (i=1,2,…,b) corresponds to a sample x i +∈ category probability distribution.
[0084] The entire matrix P integrates the category probability distribution information of b samples.
[0085] Measure the distribution difference between open set samples and closed set samples D(p(x i +∈),p(x z +∈)) is as follows:
[0086]
[0087] Among them, p(x i +∈) represents the sample x i +∈ category probability distribution vector [p(i,1),p(i,2),…,p(i,C)],p(x z +∈) represents the sample x z +∈ category probability distribution vector [p(z,1),p(z,2),…,p(z,C)], p (i,c) Represents sample x i The class probability, p (z,c) Represents sample x z The category probability of .
[0088] For all samples, a pairwise distribution difference loss is designed to achieve fast calculation;
[0089] First, let B = min(|X′ ID |,|X′ OOD |).
[0090] Where B represents the open set sample set X′ OOD and the closed sample set X′ ID The minimum of the two quantities. |X′ DOD | represents the number of samples in the open set sample set, |X′ ID | represents the number of samples in the closed sample set.
[0091] Assume {u 1 ,u 2 ,...,u B} is the predicted closed set sample set, {v 1 ,v 2 ,...,v B} is the predicted open set sample set.
[0092] Define a random mapping π: π∶{1,2,...,B}→{1,2,...,B}.
[0093] The corresponding random pairing relationship is:
[0094] Among them, u i (i=1,2,...,B) represents the i-th closed set sample in the set. z (i=1,2,...,B) is the jzth open set sample in the set. π(i) is the value obtained by applying random mapping π to i. It is the index of a sample in the open set sample set and is used to convert the closed set sample u i Establish a random pairing relationship with the open set samples. π(i) According to the random mapping π, and the closed set sample u i The open set sample of random pairing. The loss function based on paired samples is expressed as:
[0095]
[0096] Among them, L p Denotes the pair distribution difference loss. B denotes the open set sample set X′ OOD and the closed sample set X′ ID The minimum of the two quantities. N id The number of closed set samples represented by M ood Represents the total number of open set sample sets.
[0097] The advantage of the random pairing distribution difference loss designed in this invention is that it is relatively robust to the division of open set samples and closed set samples. Even if there are closed set samples that are incorrectly predicted, the probability of pairing them with samples of the same category is very small. The noise gradient brought is also very small, which enables closed set samples and open set samples to be effectively separated, thereby improving the classification effect of the image classification model.
[0098] Step 8. Calculate the model's loss function and update the model parameters until all data are tested; input the images to be classified obtained in real time into the tested image classification model to achieve image classification.
[0099] Define the loss function of the model as L, L = L en +λL P Among them, L en represents the entropy minimization loss, L P represents the pairwise distribution difference loss, and λ is a hyperparameter.
[0100] The present invention proposes an image classification method in an adaptive environment during open set testing, which includes an open set sample recognition strategy based on dual-mode matching and a paired distribution difference loss. Its core idea is to improve the model's ability to distinguish open set samples through accurate open set sample recognition and robust open set representation learning. When each test batch arrives, the open set samples are quickly identified through the dual-mode matching strategy. This precise recognition ability prevents the model from mistakenly classifying open set samples into known categories, thereby avoiding the occurrence of misclassification. By optimizing the paired distribution difference loss, the model has a stronger adaptability to those open set samples that are difficult to distinguish, such as the imaging features of rare diseases. This means that when faced with complex and diverse medical images, the model can more accurately judge the difference between normal images (closed set samples) and abnormal images (open set samples) representing rare diseases, greatly improving the accuracy of image classification.
[0101] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above-mentioned embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with this field under the guidance of this specification fall within the substantive scope of this specification and should be protected by the present invention.
Claims
1. A method for image classification in an adaptive environment during open set testing, characterized in that The steps include: Step 1. Obtain the pre-trained image classification model f θ ; where θ represents the model parameters; get the test data set D represents the image x i The unlabeled dataset N test is the total number of samples in the test dataset; Step 2. Initialize the image classification model f θ Hyperparameters of Step 3. At each time step t, receive a batch of test data; Step 4. Based on the pattern matching strategy of model parameter update, starting from the perspective of gradient updating model parameters, an open set score ods(x) is constructed according to the model's prediction confidence to reflect the degree of certainty of the model's image classification results. By analyzing the relationship between gradient and model parameter update, combined with the open set score constructed by prediction confidence, a preliminary judgment is made on the open set samples; Step 5. Based on the pattern matching strategy of feature space change, the position of the test image in the feature space and the difference in feature distribution with known category images are analyzed to determine whether it is an open set sample; Step 6. Based on the intersection of the pattern matching strategies proposed in Step 4 and Step 5, perform final identification of open set samples and closed set samples, and perform entropy minimization loss only on the identified closed set samples; Step 7. Design a pairwise distribution difference loss based on visual cue words to leverage information from easy open-set samples to help the model learn more robust representations to identify difficult open-set samples and better distinguish between open-set and closed-set samples. Step 8. Calculate the image classification model f θ The loss function is used to update the model parameters until all data are tested. The images to be classified obtained in real time are input into the tested image classification model to realize image classification.
2. The image classification method in an adaptive environment during open set testing according to claim 1, characterized in that In step 4, the calculation formula of the open set score ods(x) is as follows: in, Represents the category probability distribution predicted by the model at time t The probability value of the i-th category; Represents the category probability distribution predicted by the model at time t+1 The probability value of the i-th category; Represents the probability value of the model predicting that the input sample x belongs to category c′ at time t; Indicates finding the model prediction probability from all categories C at time t The largest category; Is the indicator function, when i is exactly equal to the category index with the highest model prediction probability at time t, The value of is 1, otherwise it is 0; C represents the number of categories, and c′ is a category in C.
3. The image classification method in an adaptive environment during open set testing according to claim 2, characterized in that An efficient calculation method is designed by using the distribution similarity between test batch data, and the loss is calculated indifferently by the current sample and the gradient of the model is updated. To determine whether the sample belongs to the open set sample, as shown below: θ′ t =α*θ′ t-1 +(1-a)*θ t ; Among them, θ t represents the model parameters obtained by completing the gradient update at time t, α is the update momentum; θ′ t-1 Refers to the model parameters updated at time t-1; θ′ t Refers to the model parameters updated by the exponential moving average EMA at time t; It refers to the updated parameter θ′ at time t t The model predicts the probability value of the sample belonging to the i-th category; Is the indicator function, when i is exactly equal to the category index with the highest model prediction probability at time t, The value of is 1, otherwise it is 0; After each test batch of data arrives, the open set score is directly calculated; a preliminary open set detection formula is constructed: Among them, X OOD represents the open set sample space; p(x∈X OOD ) represents a probability representation used to determine whether the sample x belongs to the open set sample space X OOD The probability of; when ods(x)<0, then p(x∈X OOD )=1, indicating that the sample x is an open set sample; when ods(x)≥0, then p(x∈X OOD )=0, indicating that the sample x is a closed set sample.
4. The image classification method in an adaptive environment during open set testing according to claim 1, characterized in that The step 5 is specifically as follows: For open set samples, since their characteristic patterns are inherently false, the mask has little effect on them. Therefore, by analyzing the amplitude of the characteristic pattern change, the open set and closed set samples can be effectively distinguished. Among them, m h,w It represents a mask matrix element, which is used to determine whether the pixel with coordinates (h, w) in the image is masked. ρ is the probability of the mask, and H and W are the width and height of the image respectively. Bernoulli(1-ρ) represents Bernoulli distribution, which is a discrete probability distribution; Represents the indicator function to determine m h,w Is it 1 or 0, thus constructing a feature space mask; (h,w) represents the coordinates of the pixel in the image, h is the coordinate in the height direction, w is the coordinate in the width direction, and the value ranges of h and w are from 1 to the height H of the image and from 1 to the width W of the image respectively; Define a new closed set score ids(x), the calculation formula is as follows: in, Represents the category probability distribution predicted by the model The probability value of the i-th category, x⊙m represents the element-wise multiplication of the image sample x and the mask matrix m, and the image x is processed by the mask m; Indicates that at time t based on parameter θ t The model predicts the probability value of the masked sample x⊙m belonging to the i-th category; Indicates that at time t, find the model prediction probability from all categories C The largest category; Is the indicator function, when i is equal to the category index with the highest model prediction probability at time t, The value of is 1, otherwise it is 0; C represents the number of categories, and c′ refers to a category in C; The median of the ids score of the entire batch data X is selected as the dynamic threshold, and the ids score of a single sample x is calculated as: Where median(ids(X)) represents the median of the ids score of the entire batch data X, that is, the dynamic threshold; ids(x) represents the score of the closed set sample, which is a numerical indicator used to measure whether the sample x belongs to the closed set sample; X ID represents the closed sample space; p(x∈X ID ) represents a probability representation used to determine whether the sample x belongs to the closed sample space X ID probability; when ids(x)>median(ids(X)), it means that sample x belongs to the closed set sample; otherwise, when ids(x)≤median(ids(X)), it means that sample x belongs to the open set sample.
5. The image classification method in an adaptive environment during open set testing according to claim 1, characterized in that The step 6 is specifically as follows: In order to combine the open set sample recognition strategies of the two pattern matching, using their intersection as a measure, for any data batch X arriving at time t t , the identified closed set sample set X′ ID and open sample set X′ OOD They are: X′ ID ={x|ids(x)>median(ids(X t ))∩ods(x)≥0}; X′ OOD ={x|ids(x)≤median(ids(X t )∩ods(x)<0)}; Among them, median(ids(X)) represents the median of the ids score of the entire batch data X, that is, the dynamic threshold; ids(x) represents the score of closed set samples; ods(x) is the open set score.
6. The image classification method in an adaptive environment during open set testing according to claim 5, characterized in that In order to achieve test-time adaptation, the entropy minimization loss L is performed only on the identified closed set samples. en , the calculation formula is as follows: Among them, x i ∈X′ ID Closed sample set, x i refers to X′ ID The i-th sample in , N = |X′ ID |, refers to the closed set sample X′ ID The number of samples in; C′ is the number of closed set categories; f(x i ) represents the model's response to sample x i The prediction result is a category probability distribution vector; h(f(x i )) c It means the model predicts sample x i The probability value of belonging to category c.
7. The image classification method in an adaptive environment during open set testing according to claim 1, characterized in that The step 7 is specifically as follows: Define a visual cue word ∈, where ∈ is a learnable parameter matrix whose size is equal to the image x; for any image x, using a visual cue word is equivalent to adding the two on x, i.e., x = x + ∈; For difficult open set samples x∈X OOD , where X OOD Represents the open set sample space. The ultimate optimization goal of the prompt word design is to ID , X ID represents the closed set sample space; the definition of difficult open set samples is as follows: 1 T |f(x+∈)-f(x′+∈)|>δ1; 1 T |h(f(x+∈))-h(f(x′+∈))|>δ2; Among them, f(x+∈) is the prediction result of the model for the closed set sample x; f(x′+∈) is the prediction result of the model for the open set sample x; h(f(x)) refers to the probability value of the model predicting the test sample x; 1 T It refers to the transpose of the all-one vector 1, which is used for matrix multiplication operations to aggregate the differences between vectors; δ1 and δ2 are the thresholds for distinguishing open set samples from closed set samples.
8. The image classification method in an adaptive environment during open set testing according to claim 7, characterized in that In order to distinguish difficult open-set samples in images, the results of open-set sample recognition based on dual-pattern matching are used; Assume that at each time t, the open set sample set in the current batch data is identified and is X iD The subset of X′ OOD , a closed sample set and is X OOD The subset of X′ ID ; It is believed here Among them, X ID is a closed set of samples, X OOD is an open set sample set; For a batch of samples {x1,x2,...,x b }, first apply the visual cue word ∈ to obtain batch samples {x1+∈,x2+∈,...,x b +∈}; the category probability distribution of each sample is represented as a matrix is a matrix; where b is the number of batch samples, i.e. the number of samples, and C is the number of categories; each row of this matrix corresponds to the category probability distribution of a sample; the i-th row is the sample x i+∈ The category probability distribution of : P=[p1;p z ...p b ]; Among them, P is a matrix used to represent the category probability distribution of batch samples; each row p i Corresponding to a sample x i +∈ category probability distribution, i = 1, 2, ..., b; the entire matrix P integrates the category probability distribution information of each of the b samples; Measure the distribution difference between open set samples and closed set samples D(p(x i +∈),p(x z +∈)) is as follows: Among them, p(x i +∈) represents the sample x i +∈ category probability distribution vector; p(x z +∈) represents the sample x z +∈ category probability distribution vector, p (i,c) Represents sample x i The class probability, p (z,c) Represents sample x z The category probability of .
9. The image classification method in an adaptive environment during open set testing according to claim 8, characterized in that For all samples, a pairwise distribution difference loss is designed to achieve fast calculation; First, let B = min(|X′ ID |,|X′ OOD |); Where B represents the open set sample set X′ OOD and the closed sample set X′ ID The minimum value of the two quantities; |X′ OOD | represents the number of samples in the open set sample set, |X′ ID | represents the number of samples in the closed set of samples; Assume {u 1 ,u 2 ,...,u B } is the predicted closed set sample set, {v 1 ,v 2 ,...,v B } is the predicted open set sample set; Define a random mapping π: π∶{1,2,...,B}→{1,2,...,B}; The corresponding random pairing relationship is: Among them, u i represents the i-th closed set sample in the set; v z is the zth open set sample in the set, i = 1, 2, ..., B, π(i) is the value obtained after applying the random mapping π to i, which is the index of a sample in the open set sample set, and is used to convert the closed set sample u i Establish a random pairing relationship with the open set samples; v π(i) According to the random mapping π, and the closed set sample u i The open set sample of random pairing; the loss function based on paired samples is expressed as: Among them, L p represents the pairing distribution difference loss; B represents the open set sample set X′ OOD and the closed sample set X′ ID The minimum of the two numbers, N id The number of closed set samples represented by M ood Represents the total number of open set sample sets.
10. The image classification method in an adaptive environment during open set testing according to claim 1, characterized in that In step 8, the loss function of the model is defined as L: L = L en +λL P Among them, L en represents the entropy minimization loss, L P represents the pairwise distribution difference loss, and λ is a hyperparameter.
Citation Information
Cited By
A method and device for content management and control based on mixed enhancement of large models
CN122530717A