Image classification method based on continuous adaptive knowledge space

By introducing a continuous adaptive knowledge space into the image classification method, processing unknown images in the dynamic data flow, the problem of the transformation of unknown categories into known categories is solved, the accuracy of image recognition and classification is improved, and the consumption of storage resources is reduced.

CN119445263BActive Publication Date: 2025-05-20SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510038580.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-20
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process unknown images in dynamic data streams, cannot convert unknown categories into known categories, and storing unknown image samples consumes a large amount of resources.

Method used

Using an image classification method based on continuous adaptive knowledge space, the representation vectors of the image embedded representation are obtained, and they are divided into sets of positive and negative representation vectors, and the centroid and hypersphere of known categories are constructed, the knowledge space is updated to adapt to the new data distribution, and the accuracy of the pseudo-label is judged by calculating the distance, which is converted into real tags.

Benefits of technology

Real-time processing and classification of unknown images is realized, the recognition accuracy of unknown images and classification accuracy of known images is improved, the dependence on storage resources is reduced, and the application scenarios of scarce samples is adapted to.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445263B_ABST
    Figure CN119445263B_ABST
Patent Text Reader

Abstract

The present invention discloses an image classification method based on a continuous adaptive knowledge space, which belongs to the field of computer vision and image classification. The method comprises: constructing a knowledge space of known samples according to the centroid of known categories and the radius of the centroid, and updating the knowledge space of known samples to an adaptive knowledge space; in the adaptive knowledge space, calculating the second distance from the centroid of each known category in the current image classification task to all hyperspheres in the knowledge space before updating; judging whether the centroid of the hypersphere falls in the pseudo-label hypersphere according to the second distance, and if so, converting the pseudo-label of the hypersphere into a real label. The adaptive knowledge space can continuously contain the category information of all known samples, and evaluate the accuracy of the pseudo-label by calculating the distance between the centroid of the known category and the hypersphere, thereby converting the pseudo-label with high confidence into a real label, which greatly improves the recognition accuracy of unknown images and the classification accuracy of known images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image classification, and in particular to an image classification method based on a continuously adaptive knowledge space. Background Art

[0002] In recent years, deep neural networks have made remarkable progress in static computer vision tasks. With the development of technology and the expansion of application scenarios, more and more challenges of dynamic data streams are faced. In these dynamic environments, traditional unknown image detection methods show their limitations. These methods usually stay at the stage of unknown image detection, that is, the existing methods only regard unknown images as a single category and do not further process and transform these images, namely: the existing methods only have the clustering ability of unknown images, but cannot further classify and process unknown images. Facing the dynamic data stream, the limitation of storage space becomes a key issue. It is obviously unrealistic to store unknown image samples for a long time because it will consume a large amount of resources. Therefore, a more efficient strategy is needed to process these unknown samples. At the same time, as the training samples of unknown categories gradually appear, the model needs to have the ability to transform the unknown categories detected in the past into known categories. How to instantaneously process unknown samples and transform unknown categories into known categories in real time during the process of continuous learning is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0003] The purpose of the present invention is to overcome the problems of the prior art and provide an image classification method based on a continuously adaptive knowledge space.

[0004] The purpose of the present invention is achieved through the following technical solutions: an image classification method based on a continuously adaptive knowledge space, the method comprising the following steps:

[0005] Obtain an embedding representation of an input image of a current image classification task;

[0006] Obtain a feature vector of the embedding representation of the input image;

[0007] Divide the feature vectors into a positive feature vector set and a negative feature vector set according to the true labels corresponding to the feature vectors, and construct the centroids of known categories according to the positive feature vector set;

[0008] Calculate the first distances from the negative feature vectors in the negative feature vector set to the centroids, and construct the radius of each centroid;

[0009] Construct a knowledge space of known samples according to the centroids and the radii of the centroids, and update the knowledge space of known samples to an adaptive knowledge space; the knowledge space includes the centroids and hyperspheres of known categories, and the hyperspheres determine the boundaries and ranges of each category;

[0010] In the adaptive knowledge space, calculate the second distance from the centroid of each known category in the current image classification task to all the hyperspheres in the knowledge space before the update; determine whether the centroid of the hypersphere falls within the pseudo-label hypersphere based on the second distance. If so, convert the pseudo-label of the hypersphere to the true label.

[0011] In one example, before obtaining the embedding representation of the input image for the current image classification task, it further includes:

[0012] Construct a prompt pool for the current image classification task. The prompts in the prompt pool are learnable image vectors, and each prompt is associated with a learnable key as a value to form a key-value pair.

[0013] In one example, before obtaining the feature vector of the image embedding representation, it further includes enhancing the embedding representation, specifically including:

[0014] Calculate the similarity distance between the embedding representation and the keys in the prompt pool, search for multiple similar keys in the prompt pool with similarity distances less than the distance threshold, match the corresponding prompt subset according to the multiple similar keys, and splice the matched prompt subset with the image embedding representation to obtain an enhanced image embedding representation.

[0015] In one example, after obtaining the feature vector of the image embedding representation, it further includes:

[0016] Perform classification processing on the feature vector based on the image classification model, calculate the classification loss function according to the difference between the classification result and the true label value, perform backpropagation processing on the image classification model according to the classification loss function, thereby obtaining the optimized prompt pool for the current image classification task, and update the prompt pool according to the optimized prompt pool.

[0017] In one example, construct the knowledge space of the known samples using the optimized centroid and optimized radius. Obtaining the optimized centroid and optimized radius includes:

[0018] Calculate the class average margin loss function according to the feature vector, centroid, and radius, perform optimization processing on the image classification model according to the class average margin loss function, and obtain the optimized prompt pool, optimized centroid, and optimized radius for the current image classification task.

[0019] In one example, obtaining the optimized centroid and optimized radius further includes:

[0020] Calculate the average distance loss between the embedding representation and multiple similar keys, perform classification processing on the feature vector based on the image classification model, calculate the classification loss according to the difference between the classification result of the feature vector and the true label value, and determine the loss function based on the data augmentation paradigm according to the average distance loss and classification loss;

[0021] Based on the loss function and the class-averaged margin loss function in the data augmentation paradigm, the final optimized objective function for the current image classification task is obtained. According to the optimized objective function, the image classification model is optimized to obtain the optimized hint pool, optimized centroid, and optimized radius for the current image classification task.

[0022] It should be further noted that the technical features corresponding to the above method examples can be combined or replaced with each other to form a new technical solution.

[0023] In one example, an image classification method based on a continuously adaptive knowledge space is used to test the image classification method formed by combining any one of the above examples or multiple examples, including the following steps:

[0024] Obtain the embedding representation of the image to be tested;

[0025] Perform enhancement processing on the embedding representation to obtain an enhanced image embedding representation;

[0026] Obtain the feature vector of the image embedding representation;

[0027] In the adaptive knowledge space, calculate the third distance between the feature vector and all known class centroids in the knowledge space before update, and determine the minimum third distance; obtain the most similar class closest to the feature vector, compare the minimum third distance with the radius of the most similar class. If the minimum third distance is less than the radius of the most similar class, the image to be tested corresponding to the feature vector is classified as the most similar class; otherwise, the image to be tested corresponding to the feature vector is recognized as an unknown sample;

[0028] Perform clustering processing on the unknown samples to obtain the centroid and radius of each clustering cluster, and assign pseudo-labels to each clustering cluster and update them to the adaptive knowledge space.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] 1. In one example, the adaptive knowledge space is continuously updated through the knowledge space of known samples. The adaptive knowledge space can continuously contain the class information of all known samples. By calculating the distance between the known class centroid and the hypersphere in the knowledge space, it is judged whether the centroid of the hypersphere falls into the pseudo-label hypersphere, thereby evaluating the accuracy of the pseudo-label, and then converting the pseudo-label with high confidence into a true label, that is, gradually converting unknown classes into known classes in the process of adapting to the new data distribution through the adaptive knowledge space, greatly improving the recognition accuracy of unknown images and the classification accuracy of known images; at the same time, this classification method does not need to rely on a large amount of labeled data and can adapt to application scenarios with scarce samples.

[0031] 2. In one example, optimizing the image classification model through the class-average margin loss function can increase the inter-class distance and decrease the intra-class distance. Continuously detecting unknown images based on the margin can more effectively reduce the open-set risk and improve the model's recognition and processing ability for unknown images.

[0032] 3. In one example, optimizing the model based on the loss function of the data augmentation paradigm can overcome the problems of memory overhead caused by traditional sample augmentation and the non-reusability of augmented samples based on the provided data augmentation paradigm, effectively solve the overfitting problem caused by small samples, and improve the classification accuracy of known images. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings provided herein are used to provide a further understanding of the present application and form a part of the present application. The same reference numerals are used to represent the same or similar parts in these drawings. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application.

[0034] Figure 1 It is a flowchart of the method provided in one example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] In the description of the present invention, ordinal numbers (e.g., "first and second", "first to fourth", etc.) are used to distinguish objects and are not limited to this order, and should not be construed as indicating or implying relative importance.

[0037] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, terms such as "belonging to", "connected", and "connected to" should be understood in a broad sense. For example, it can also be an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal connection of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0038] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0039] In one example, as Figure 1As shown, an image classification method based on a continuous adaptive knowledge space, which can be used in the training phase of the model and also in the detection and testing phases. The method includes the following steps:

[0040] S1: Obtain the embedding representation of the input image for the current image classification task.

[0041] In this step, the embedding representation of the input image for the current image classification task is obtained through a feature extraction model, such as an image processing model based on Transformer. In this example, a Vision Transformer (ViT) pre-trained on the ImageNet-1k dataset (an image dataset containing 1,000 categories) is used, and the ViT parameters are frozen for feature extraction. Specifically, ViT first splits the image into a series of image patches and flattens them, and then inputs them as a sequence into the Transformer model. With the help of the self-attention mechanism and multiple Transformer encoders, ViT can effectively extract image features. ViT includes an image patch embedding module, a position encoding module, a Transformer encoder, a global average pooling layer, a fully connected layer, and an activation function layer. Among them, the image patch embedding module divides the input image into image patches of a fixed size and flattens these patches into a vector form; the position encoding module is responsible for encoding the information of each position in the sequence; the Transformer encoder consists of multiple encoding layers. In each encoding layer, the input sequence will first pass through the self-attention layer for feature extraction, and then through the fully connected feed-forward neural network for feature mapping; the global average pooling layer obtains the overall feature representation of the image based on global average pooling processing; the fully connected layer maps the pooled feature vector to the number of class labels, and then outputs the final classification probability through the Softmax activation function layer. In this example, the embedding representation of the input image for the current image classification task can be obtained through the image patch embedding module of ViT.

[0042] S2: Obtain the representation vector of the embedding representation of the input image.

[0043] Specifically, the representation vector of the embedding representation of the input image is obtained through a feature extraction model, such as an image processing model based on Transformer. In this example, a Vision Transformer (ViT) pre-trained on the ImageNet-1k dataset is used, and the ViT parameters are frozen for feature extraction. In this example, the representation vector of the embedding representation of the input image can be obtained through the image patch embedding module of ViT, that is, the image patch embedding module flattens the image patch into a one-dimensional vector and converts this one-dimensional vector into an embedding vector of a fixed dimension to obtain the representation vector of the embedding representation of the input image.

[0044] S3: Divide the representation vectors into a positive representation vector set and a negative representation vector set according to the true labels corresponding to the representation vectors, and construct the centroids of known categories based on the positive representation vector set.

[0045] Among them, the positive representation vectors in the positive representation vector set are the sample vectors whose training labels match the current category, and the negative representation vectors in the negative representation vector set are the sample vectors whose training labels do not match the current category; the centroid of a category represents the central feature vector of a category.

[0046] S4: Calculate the first distance from the negative representation vectors in the negative representation vector set to the centroids, and construct the radius of each centroid.

[0047] In the process of image classification, the category boundary can be defined by negative samples (samples corresponding to negative representation vectors). By calculating the distance from the negative representation vectors to the centroids, the boundary of the category can be better determined, and then the radius of each centroid can be constructed.

[0048] S5: Construct the knowledge space of known samples based on the centroids and radii, and update the knowledge space of known samples to the adaptive knowledge space.

[0049] Among them, the knowledge space includes the centroids and hyperspheres of known categories. The hyperspheres determine the boundaries and ranges of each category, that is, all samples of a category are located inside the corresponding hypersphere. The adaptive knowledge space also includes the centroids and corresponding hyperspheres of known categories, and is continuously updated through the knowledge space of known samples, and can continuously contain the category information of all known samples, and thus can continuously adapt to the new data distribution.

[0050] S6: In the adaptive knowledge space, calculate the second distance from the centroid of each known category in the current image classification task to all hyperspheres in the knowledge space before updating; judge whether the centroid of the hypersphere falls into the pseudo-label hypersphere according to the second distance. If so, convert the pseudo-label of the hypersphere into the true label.

[0051] When the next new training task comes, repeat the above steps S1 - S6.

[0052] Among them, the true label is used to represent the true category to which the hypersphere belongs, and the pseudo-label is used to represent the possible category to which the hypersphere belongs. In this step, by comparing the second distance with the boundary of the pseudo-label hypersphere, it can be judged whether the centroid of the hypersphere is located in the pseudo-label hypersphere. If the centroid of the hypersphere falls into the pseudo-label hypersphere, it means that the pseudo-label of the pseudo-label hypersphere has sufficient confidence to be considered as the true label, and thus the unknown category is converted into a known category, greatly improving the recognition accuracy of unknown images and the classification accuracy of known images.

[0053] In one example, before obtaining the embedding representation of the input image of the current image classification task, it further includes:

[0054] Construct a prompt pool for the current image classification task. The prompts in the prompt pool are learnable image vectors, and each prompt is associated with a learnable key as a value to form a key-value pair.

[0055] Specifically, during the training process of the current image classification task a task-specific prompt set is constructed and initialized with a uniform distribution:

[0056] ;

[0057] ;

[0058] ;

[0059] where the prompt pool contains L key-value pairs, represents the key, represents the value; represents the task number; ~ represents the case of "following" a certain data distribution; represents a uniform distribution from -1 to 1; , both follow a uniform distribution from -1 to 1. The uniform distribution is a commonly used probability distribution, and its key feature is that each numerical value within the interval has an equal probability of occurrence, that is, all possible values have the same probability. In addition, other common parameter initialization methods such as zero initialization, normal distribution random initialization, Glorot initialization, He initialization, etc. can also be selected. Optionally, for the task the classification head of ViT is initialized with a uniform distribution:

[0060] ;

[0061] where the classification head of ViT includes a fully connected layer and a Softmax activation function layer. In the above formula, is the weight matrix to be initialized; represents the uniform distribution; represents the input dimension of the current layer, which is the input dimension of the fully connected layer in the present invention.

[0062] In one example, obtaining the embedding representation of the input image of the current image classification task specifically includes:

[0063] For the current task the training sample set where represents a specific input training sample (input image); represents the category; is the category The sample set of = , where is the current task The number of training samples in the present invention is set to 5. After passing through a convolutional layer with a convolutional kernel size of in a pre-trained model (ViT in this example) and a flattening layer with a scale of as the query function for encoding, the input sample set is projected onto a vector space with a dimension of Then, an additional encoding (classification token) with a size of is concatenated and added with the positional encoding of After passing through the dropout layer, the final embedded representation of the input sample set is obtained, that is : :

[0064] ;

[0065] where represents concatenation of the same dimension; are the embedding network parameters of the image patch embedding module in ViT; and are the frozen parameters of the ViT model, represents the class encoding, represents the positional encoding; is the dropout layer. Additionally, it should be noted that the dropout rate adopted by the dropout layer in the present invention is 0, that is, in the network of the present invention, actually no embedded representation at any position is deactivated. In actual situations, the dropout rate can be adjusted according to the environment.

[0066] In one example, before obtaining the characterization vector of the image embedded representation, it further includes enhancing the embedded representation, specifically including:

[0067] Calculating the similarity distance between the embedded representation and the keys in the prompt pool, searching for multiple similar keys in the prompt pool with similarity distances less than the distance threshold, matching the corresponding prompt subsets according to the multiple similar keys, and concatenating the matched prompt subsets with the image embedded representation to obtain the enhanced image embedded representation. By setting different distance thresholds, different numbers of similar keys can be filtered, and the distance threshold can be set based on historical experience or actual usage requirements.

[0068] Specifically, according to the embedded representation , using the cosine similarity distance metric method to calculate An embedded representation and a hint pool The distance of the keys in; Based on the similarity distance, use the Top-K algorithm (the top K maximum and minimum value algorithm) to search for the top K keys with the smallest distance from the hint pool of the current task and match the corresponding hints ; The matched hints are concatenated with the embedded representation of the image to obtain an image embedded representation enhanced by hints , that is , The specific calculation includes:

[0069] a. According to the embedded representation , use the cosine similarity distance metric method to calculate and the distance of the keys stored in the hint pool In the present invention, the image embedded representation and can directly calculate the cosine similarity distance:

[0070] ;

[0071] Among them, represents the cosine similarity calculation; represents the vector modulus operator, is the th embedded representation modulus; is the key modulus. Of course, other distance metric methods, such as Euclidean distance, Manhattan distance, etc., can also be used for similarity measurement.

[0072] b. Based on the similarity distance obtained in step a, calculate the frequency of selecting each hint in the hint pool to obtain a hint frequency list , is the frequency list quantity label, and use the Top-K algorithm to select from the hint frequency list to search for the top K similar keys with the smallest distance from the hint pool of the current task and , The Top-K algorithm is as follows:

[0073] ;

[0074] Among them, represents the independent variable parameter value when the function obtains the minimum value; , M are both key quantity labels, which are the upper and lower limits of the summation function respectively.

[0075] c. According to the top K similar keys searched match the corresponding hints , the obtained prompts are concatenated with the embedding representation in the same dimension to obtain an image embedding representation with enhanced prompts :

[0076] ;

[0077] ;

[0078] Among them, is the sequence number set of K keys.

[0079] In this example, enhancing the embedding representation of the image can obtain richer image feature information, thereby improving the accuracy of image classification.

[0080] In one example, the enhanced image embedding representation is obtained through the vision self-attention mechanism ViT :

[0081] ,

[0082] Among them, ; is the encoder (parameter frozen) in the ViT model, including the position encoding module, Transformer encoder, etc.

[0083] In one example, constructing the centroid of known categories specifically includes:

[0084] Based on the obtained prompts-enhanced representation vectors and the corresponding true labels , the centroid of the known category set can be constructed, is the number of categories. For each category , the input training sample set is divided into a positive sample set and a negative sample set , that is:

[0085] ;

[0086] ;

[0087] Among them, represents a specific input training sample, represents the true label (true category label) corresponding to the input training sample; represents the input training sample set; represents the a known class;

[0088] Further, the representation vectors enhanced based on the prompts are divided into a set of positive representation vectors and a set of negative representation vectors , and the formula for constructing the centroid is as follows:

[0089] ;

[0090] In the above formula, is the centroid of class ; is the number of positive representation vectors; represents a representation vector in the set of positive representation vectors .

[0091] In one example, constructing the radius of each centroid specifically includes:

[0092] Based on the calculated centroid set , calculate the Euclidean distance from the negative representation vector to the centroid . Since the number of negative representation vectors is much larger than that of positive representation vectors, the calculation expression for constructing the radius of each centroid is as follows:

[0093] ;

[0094] In the above formula, represents the radius of the centroid; is the quantile function; is the Euclidean distance; is a hyperparameter for controlling the robustness of the radius, is a hyperparameter for controlling the margin between classes.

[0095] In one example, after obtaining the representation vectors of the image embedding representation or constructing the radius of each centroid, it further includes:

[0096] Classify the representation vectors based on the image classification model (ViT), calculate the classification loss function according to the difference between the classification result and the true label value, perform backpropagation optimization on the image classification model according to the classification loss function, and then obtain the optimized prompt pool for the current image classification task, and update the prompt pool according to the optimized prompt pool.

[0097] In this example, a prompt pool including model learnable parameters is constructed, and the model is optimized through the classification loss function, and then the model parameters are updated to improve the classification performance of the model.

[0098] In one example, using the optimized centroid and optimized radius to construct the knowledge space of known samples, obtaining the optimized centroid and optimized radius includes:

[0099] Calculate the class-average margin loss function based on the representation vector, centroid, and radius, and optimize the image classification model according to the class-average margin loss function to obtain the optimized hint pool, optimized centroid, and optimized radius for the current image classification task.

[0100] Preferably, obtaining the optimized centroid and optimized radius further includes:

[0101] Calculate the average distance loss between the embedded representation and multiple similar keys, classify the representation vector based on the image classification model, calculate the classification loss according to the difference between the classification result of the representation vector and the true label value, and determine the loss function based on the data augmentation paradigm according to the average distance loss and classification loss;

[0102] Calculate the class-average margin loss function based on the representation vector, centroid, and radius, obtain the final optimized objective function for the current image classification task according to the loss function based on the data augmentation paradigm and the class-average margin loss function, and optimize the image classification model according to the optimized objective function to obtain the optimized hint pool, optimized centroid, and optimized radius for the current image classification task.

[0103] In this example, optimizing the image classification model according to the optimized objective function specifically includes:

[0104] A. Calculate the embedded representation and the hint pool the average distance loss of the similar keys , and the calculation expression is:

[0105] ;

[0106] Among them, represents the th similar key.

[0107] B. Output the classification result of the representation vector through the fully connected layer of ViT, calculate the difference between the classification result and the true label value , and use the cross-entropy loss function to calculate the classification loss , and the calculation expression is:

[0108] ;

[0109] In the above formula, represents the cross-entropy loss function; is the classification head; , is the true probability that the sample belongs to class C, such as When the category is C, take 1; otherwise, take 0. is a sample The predicted probability of belonging to category C.

[0110] C. According to the average distance loss and the classification loss , construct a loss function for the data augmentation paradigm based on prompts :

[0111] ;

[0112] where is the hyperparameter of the average distance loss. The loss function based on the data augmentation paradigm optimizes the image classification model. At this time, the optimization objective function is:

[0113] ;

[0114] where are the learnable parameters of the model.

[0115] D. According to the representation vector , the centroid , and the radius of the centroid , calculate the class average margin loss function for all classes in task :

[0116] ;

[0117] In the formula, N is the number of categories in the training set of task ; is the radius of category ; is the set of positive representation vectors of category ; is the set of negative representation vectors of category ; is the centroid of category ; is the margin between classes; is the hyperparameter that constrains the radius size, , are the balance ratio factors of the positive and negative representation vector sets.

[0118] E. Perform weighted processing on the loss function of the data augmentation paradigm and the class average margin loss function to obtain the final optimization objective function of task :

[0119] ,

[0120] Among them, is and 's balance factor. According to the optimized objective function, the present invention uses the Adam optimizer for backpropagation update. Of course, other optimization methods can also be used, such as stochastic gradient descent, etc.

[0121] F. According to the final optimized objective function, the optimized model obtains the task The final prompt pool and the optimized centroid and optimized radius of the known samples.

[0122] Optionally, according to the prompt pool of the optimized task Update , by continuous accumulation, an updated prompt pool can be obtained, and the update calculation formula is as follows:

[0123]

[0124] When t = 1, then is initially empty.

[0125] In an example, according to the centroid and the radius of the centroid, construct the knowledge space of the known samples , and update the knowledge space of the known samples to the adaptive knowledge space, specifically including:

[0126] First, construct the knowledge space of each category according to the optimized centroid and radius, that is, the knowledge space of the known samples ;

[0127] Then, update the knowledge space of the known samples to the adaptive knowledge space, that is:

[0128]

[0129] Among them, the left side of the equation is the updated adaptive knowledge space, and the right side of the equation is the knowledge space before update; represents the known category; represents the known category of the current image classification task; is the set of centroids in the knowledge space (before update); represents the set of centroids of the known categories in the current image classification task; represents the set of radii of the centroids; represents the set of radii of the centroids in the current image classification task.

[0130] In an example, according to the updated adaptive knowledge space, based on the assumption, there is a hypersphere with a pseudo-label of in the space , if a hypersphere centroid Located on pseudo label Super Ball , then the pseudo-label Convert to true label , specifically including:

[0131] First, based on the adaptive knowledge space , calculate the current task The centroid of each known sample in The distance (second distance) to all hyperspheres in the knowledge space (before updating), with hypersphere as an example ( ), namely:

[0132] ;

[0133] Among them, , for the super ball Centroid To The set of distances to all centroids in .

[0134] Then according to the calculated distance , assuming there is a pseudo label in the space Super Ball , if the super ball Center of mass in Landing in the Super Ball Medium, that is, distance Smaller than the radius , then the pseudo-label Convert to true label , that is:

[0135] .

[0136] Combining the above examples, a preferred example of the present invention is obtained, in which the method comprises the following steps:

[0137] S10: Construct a prompt pool for the current image classification task. The prompt pool includes multiple key-value pairs. The key-value pairs are learnable parameters, including the weights and biases of the image classification model ViT;

[0138] S20: Obtain the embedded representation of the input image of the current image classification task through ViT;

[0139] S30: Calculate the similarity distance between the embedded representation and the key in the prompt pool, search for multiple similar keys with similarity distances less than a distance threshold from the prompt pool, match corresponding prompt subsets according to the multiple similar keys, and concatenate the matched prompt subsets with the image embedded representation to obtain an enhanced image embedded representation;​​

[0140] S40: Obtain the feature vector of the image embedding representation through ViT;

[0141] S50: Divide the feature vectors into a positive feature vector set and a negative feature vector set according to the true labels corresponding to the feature vectors, and construct the centroids of known classes based on the positive feature vector set;

[0142] S60: Calculate the first distance from the negative feature vectors in the negative feature vector set to the centroids, and construct the radius of each centroid;

[0143] S70: Calculate the average distance loss between the embedding representation and multiple similar keys, where the similar keys are multiple keys whose similarity distance to the embedding representation is less than the distance threshold; perform classification processing on the feature vectors based on the image classification model, calculate the classification loss according to the difference between the classification result of the feature vectors and the true label value, and determine the loss function based on the data augmentation paradigm according to the average distance loss and the classification loss; calculate the class average margin loss function according to the feature vectors, centroids and radii, obtain the final optimization objective function of the current image classification task according to the loss function based on the data augmentation paradigm and the class average margin loss function, perform optimization processing on the image classification model according to the optimization objective function to obtain the optimization hint pool, optimized centroids and optimized radii of the current image classification task, and update the hint pool according to the optimization hint pool;

[0144] S80: Construct the knowledge space of known samples according to the optimized centroids and optimized radii, and update the knowledge space of known samples to the adaptive knowledge space;

[0145] S90: In the adaptive knowledge space, calculate the second distance from the centroids of each known class in the current image classification task to all hyperspheres in the knowledge space before update; judge whether the centroids of the hyperspheres fall into the pseudo-label hyperspheres according to the second distance, and if so, convert the pseudo-labels of the hyperspheres into true labels.

[0146] The present invention designs a transformation strategy from unknown classes to known classes, that is, classifying unknown samples based on the continuously updated adaptive knowledge space, aiming to process unknown samples in real time, transform unknown classes into known classes, ensure that the model can adapt to new data distributions in a timely manner during the continuous learning process, and improve the classification performance of the model. At the same time, the present invention conducts continuous unknown image detection based on the margin, which can more effectively reduce the open risk and improve the model's recognition and processing ability for unknown images. The present invention also designs a data augmentation paradigm based on hints, which overcomes the problems of memory overhead caused by traditional sample augmentation and the non-reusable augmented samples, can effectively solve the overfitting problem caused by small samples, and improve the classification performance of known images.

[0147] After training the image classification model according to the above preferred method, the trained model is used for unknown detection and classification. At this time, testing the image classification method includes the following steps:

[0148] S100: Obtain the image data to be tested through the trained ViT = Embedded representation , is the number of the image data to be tested;

[0149] S200: Perform enhancement processing on the embedded representation to obtain the enhanced image embedded representation ;

[0150] S300: Obtain the characterization vector of the image embedded representation through ViT ;

[0151] S400: In the adaptive knowledge space, calculate the distance between the characterization vector and all the centroids of the known categories in the knowledge space . According to the calculated minimum distance , obtain the category with the minimum distance and compare its size with its radius . There are two cases for the comparison result, that is, (1) if is less than , then the sample is classified as ; (2) if is greater than , then it is recognized as an unknown sample.

[0152] S500: Perform clustering processing on the unknown samples to obtain the centroid and radius of each clustering cluster, and assign pseudo-labels to each clustering cluster and update them to the adaptive knowledge space.

[0153] Furthermore, the specific implementation methods of each sub-step of step S400 are as follows:

[0154] 1) In the adaptive knowledge space, calculate the third distance between the characterization vector and all the centroids of the known categories in the knowledge space before update, and determine the minimum third distance (the minimum distance );

[0155] Specifically, according to the prompted enhanced image characterization vector , calculate the distance between and all the centroids of the known categories in the knowledge space (the third distance):

[0156] ,

[0157] Among them, is the set of centroids of known categories in the knowledge space, is the centroid to is the set of distances from all centroids.

[0158] 2) Obtain the most similar category with the closest distance representation vector, and compare the magnitude of the minimum distance with the radius of the most similar category. If the minimum distance is less than the radius of the most similar category, the sample corresponding to the representation vector is classified as the most similar category; otherwise, the sample corresponding to the representation vector is identified as an unknown sample;

[0159] Specifically, obtain the category with the minimum distance to the centroid (the most similar category), and compare the magnitude of the minimum distance with the radius of the category :

[0160] If is less than , then the sample is classified as :

[0161] ;

[0162] If is greater than , then the sample is identified as unknown:

[0163] .

[0164] 3) Perform clustering on the unknown samples to obtain the centroids and radii of each cluster, and assign a pseudo-label to each cluster, and update it to the adaptive knowledge space:

[0165]

[0166] Among them, represents the category defined by the pseudo-label; represents the set of centroids defined by the pseudo-label; represents the set of radii defined by the pseudo-label.

[0167] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An image classification method based on a continuous adaptive knowledge space, characterized in that: The following steps are involved: Construct a prompt pool for the current image classification task. The prompts in the prompt pool are learnable image vectors. Each prompt is associated with a learnable key as a value to form a key-value pair. Get the embedded representation of the input image for the current image classification task; Get the representation vector of the embedded representation of the input image; The representation vectors are divided into a positive representation vector set and a negative representation vector set according to the true labels corresponding to the representation vectors, and the centroid of the known categories is constructed according to the positive representation vector set; In the positive characterization vector set, the positive characterization vector is the sample vector whose training label matches the current category, and in the negative characterization vector set, the negative characterization vector is the sample vector whose training label does not match the current category; the centroid of a category represents the central feature vector of a category; Calculate the first distance from the negative characterization vector to the centroid in the negative characterization vector set, and construct the radius of each centroid; The knowledge space of known samples is constructed by optimizing the centroid and the radius, and the knowledge space of known samples is updated to the adaptive knowledge space; The knowledge space includes centroids of known categories and hyperspheres, the hyperspheres determining the boundaries and range of each category; In the adaptive knowledge space, the second distance from the centroid of each known category in the current image classification task to all hyperspheres in the knowledge space before updating is calculated; based on the second distance, it is determined whether the centroid of a hypersphere of a known category falls within the pseudo-label hypersphere. If so, the pseudo-label of the hypersphere is converted into a true label; Obtaining the optimized centroid and radius also includes: Calculate the average distance loss between the embedded representation and multiple similar keys, classify the representation vector based on the image classification model, calculate the classification loss based on the difference between the classification result of the representation vector and the true label value, and determine the loss function based on the data augmentation paradigm based on the average distance loss and classification loss; The class average margin loss function is calculated according to the representation vector, centroid and radius. The final optimization objective function of the current image classification task is obtained according to the loss function based on the data augmentation paradigm and the class average margin loss function. The image classification model is optimized according to the optimization objective function to obtain the optimized prompt pool, optimized centroid and optimized radius of the current image classification task.

2. The image classification method based on continuous adaptive knowledge space according to claim 1 is characterized in that: Before obtaining the representation vector of the image embedding representation, the embedding representation is also enhanced, including: The similarity distance between the embedding representation and the key in the prompt pool is calculated, and multiple similar keys with similarity distances less than a distance threshold are searched from the prompt pool. The corresponding prompt subsets are matched according to the multiple similar keys, and the matched prompt subsets are concatenated with the image embedding representation to obtain an enhanced image embedding representation.

3. The image classification method based on continuous adaptive knowledge space according to claim 1, characterized in that: After obtaining the representation vector of the image embedding representation, it also includes: The representation vector is classified based on the image classification model, and the classification loss function is calculated according to the difference between the classification result and the true label value. The image classification model is back-propagated according to the classification loss function to obtain the optimized prompt pool for the current image classification task, and the prompt pool is updated according to the optimized prompt pool.

4. An image classification method based on a continuous adaptive knowledge space, characterized in that: Used to test the method according to any one of claims 1 to 3, comprising the following steps: Get the embedded representation of the image to be tested; Performing enhancement processing on the embedded representation to obtain an enhanced image embedded representation; Get the representation vector of the image embedding representation; In the adaptive knowledge space, the third distance between the representation vector and the centroid of all known categories in the knowledge space before updating is calculated to determine the minimum third distance; the most similar category closest to the representation vector is obtained, and the minimum third distance is compared with the radius of the most similar category. If the minimum third distance is less than the radius of the most similar category, the image to be tested corresponding to the representation vector is classified as the most similar category; otherwise, the image to be tested corresponding to the representation vector is identified as an unknown sample; The unknown samples are clustered to obtain the centroid and radius of each cluster, and a pseudo label is assigned to each cluster and updated to the adaptive knowledge space.