Image Classification System Based on Category Attribute Modeling

By adopting the method of category attribute modeling in the image classification system, the problem of large memory consumption of existing continuous learning methods is solved, and more efficient image classification and continuous learning are achieved, reducing the risk of computing resource requirements and privacy leakage, while improving the interpretability of the model.

CN116563635BActive Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310550550.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-06-13
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

Existing playback-based continuous learning methods have great challenges in memory consumption and computing resources, and may lead to data privacy leakage.

Method used

An image classification system based on category attribute modeling is adopted to realize image classification and continuous learning through technical means such as image preprocessing, backbone network, continuous learning knowledge distillation branches, self-supervised learning branches, attribute attention modules and classification modules.

Benefits of technology

This reduces the performance loss of the model in new tasks, reduces memory consumption and computing resource requirements, avoids the risk of data privacy leakage, and improves the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563635B_ABST
    Figure CN116563635B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and proposes an image classification system based on category attribute modeling, comprising an image preprocessing module for preprocessing an original image to obtain a plurality of data augmented transformed images of the original image; a backbone network for extracting a feature vector x of the first data augmented transformed image; s ; The knowledge distillation branch of continuous learning is used to minimize the distillation loss function to make the feature vector x s With the eigenvector x t The probability distribution error of is within the first setting range; the self-supervised learning branch is used to minimize the contrast loss function so that the feature vector x s and the eigenvector x c The error of is within the second setting range; the attribute attention module is used to calculate the feature vector x using attribute labeling and cross attention mechanism s The attribute encoding e of each category is used to calculate the classification score of each category. The above technical solution solves the problem of high memory consumption of the playback-based continuous learning method in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically, to an image classification system based on class attribute modeling. Background Art

[0002] Humans have the ability to continuously reuse and expand knowledge from experience, that is, they can not only apply previously learned knowledge and skills to new environments, but also use them as the basis for future learning. Currently, artificial intelligence models based on deep learning have achieved good performance in tasks such as image classification, speech recognition, and machine translation. However, most of these models are trained on static and identically distributed datasets and cannot adapt or expand their behavior over time. Experiments have shown that after a trained AI model is trained on new data or new tasks, its generalization performance on the original task will be significantly reduced. This phenomenon is called catastrophic forgetting.

[0003] The purpose of continual learning is to enable the trained machine learning model to have the ability to solve multiple tasks at different time periods respectively. It faces the learning task of data with inconsistent distributions at multiple moments, and requires the model to be able to learn different specific tasks at different moments under a set of parameters (plasticity), while not forgetting the previous knowledge (stability). Among them, the replay-based continual learning method, when learning a new task, in order to reduce forgetting, replays the samples stored in the previous task during the training process. These samples / pseudo-samples can be used for joint training or for constraining the optimization of the new task loss to avoid interfering with the previous task and achieve continual learning. Extra computing resources and storage space are required to recall old knowledge. When the number of task types increases continuously, either the training cost will become high, or the representativeness of the representative samples will be weakened. At the same time, in the actual production environment, this method may also have the problem of data privacy leakage. Summary of the Invention

[0004] The present invention proposes an image classification system based on class attribute modeling, which solves the problem of large memory consumption of the replay-based continual learning method in the related art.

[0005] The technical solution of the present invention is as follows: It includes:

[0006] An image preprocessing module, used to preprocess the original image to obtain multiple data-augmented transformed images of the original image. The preprocessing includes random cropping, or random scaling, or random color, or normalization;

[0007] A backbone network, which serves as a student network and is used to extract the feature vectors of the first data-augmented transformed images ;

[0008] A continual learning knowledge distillation branch, used to minimize the distillation loss function to make the feature vectors has a probability distribution error within a first set range with the eigenvector ; the eigenvector is the eigenvector of the first data-augmented transformed image extracted by the teacher network, and the teacher network has the same network structure as the student network. In the current round of task training, the teacher network is the student network obtained in the previous task training;

[0009] The self-supervised learning branch minimizes the contrast loss function to make the error between the eigenvector and the eigenvector within a second set range; the eigenvector is the eigenvector of the second data-augmented transformed image extracted by the contrast network. The second data-augmented transformed image and the first data-augmented transformed image are images obtained by preprocessing the same original image through different methods, and the contrast network has the same network structure as the student network;

[0010] The attribute attention module is used to calculate the attribute encoding e of the eigenvector using the attribute label and the cross-attention mechanism;

[0011] The classification module includes a linear classifier group composed of multiple class-independent parallel classifiers. The attribute encoding e of the eigenvector is input into each linear classifier and activated using sigmoid to calculate the classification score for each class.

[0012] Furthermore, the backbone network is specifically a Vision Transformer (ViT) or a Convolutional Neural Network (CNN).

[0013] Furthermore, the calculation process of the distillation loss function is as follows:

[0014]

[0015]

[0016]

[0017] where is the temperature coefficient used to control the smoothness of the probability distribution, represents the i-th eigenvector , represents the i-th eigenvector .

[0018] Furthermore, it is characterized in that the calculation process of the contrast loss function is as follows:

[0019]

[0020] Among them, the function is the cosine similarity function.

[0021] Furthermore, calculating the feature vector using the attribute tag and the cross-attention mechanism

[0022] dynamically adds the attribute tag and uses it as the Query in the cross-attention layer, and the feature vector is used as the Key in the cross-attention, and calculates the output of the cross-attention layer:

[0023]

[0024] Among them, , , are parameters in the cross-attention layer; is the dimension of the feature, is the number of attention heads in the cross-attention layer; are all process variables;

[0025] The output of the cross-attention obtains the attribute encoding e of the feature vector after layer normalization with residual and a linear layer.

[0026] Among them, in any training, the attribute tag is divided into and two parts, where is the attribute tag vector obtained from the training in the previous tasks and is not updated during the training of task ; is added to the model at the start of the training of task , and its initial value is sampled from the standard normal distribution and updated through backpropagation during the training of task ; the number of attribute tags , p is the number of categories, r = 1, 2, … l.

[0027] Furthermore, it also includes an orthonormalized cross-attention layer for the attribute tag , and the orthonormalization process specifically includes:

[0028] activates the calculated in the cross-attention layer using and normalizes it in the dimension of the batch to obtain ;

[0029] Calculate ;

[0030] Substitute into the regular term approximation formula for calculation:

[0031]

[0032] The working principle and beneficial effects of the present invention are as follows:

[0033] In the present invention, the self-supervised learning branch is used to extract general feature representations irrelevant to the task, reduce overfitting in specific tasks, and thus reduce the performance loss of the model in new tasks; the knowledge distillation branch of continuous learning is used to maintain the memory of knowledge in old tasks; then, in the attribute attention module, learnable orthogonal attribute tags and cross-attention mechanisms are used to calculate the attribute encoding e of the input features, so as to model the category attributes. Finally, the attribute encoding e is input into each category-independent parallel classifier and activated using sigmoid to calculate the classification scores for each category.

[0034] The present invention only saves a feature extractor obtained from training in the previous task as the teacher network, which has a smaller memory consumption compared to saving some samples for replay, and there is also no possibility of privacy leakage; at the same time, by modeling the category attributes, it can be analyzed through the visualization of the feature space and the encoding space, and the model has relatively higher interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments.

[0036] Figure 1 is the structural block diagram of the image classification system based on category attribute modeling of the present invention;

[0037] Figure 2 is the structural block diagram of the attribute attention module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0039] As Figure 1 shown, the image classification system based on category attribute modeling in this embodiment includes:

[0040] (1) Image preprocessing module, which is used to preprocess the original image to obtain multiple data-augmented transformed images of the original image. The preprocessing includes random cropping, or random scaling, or random color or normalization.

[0041] (2) Backbone network, which serves as the student network and is used to extract the feature vector of the first data-augmented transformed image I ;

[0042] (3) Knowledge distillation branch for continuous learning, which is used to minimize the distillation loss function so that the probability distribution error between the feature vector and the feature vector is within the first set range; the feature vector is the feature vector of the first data-augmented transformed image extracted by the teacher network. The teacher network has the same network structure as the student network. In the training of the current round of tasks, the teacher network is the student network obtained in the previous task training.

[0043] Knowledge distillation is a learning paradigm that treats the model as a black box and knowledge as the mapping relationship from input to output. By training the student network to fit the output of the teacher network, it realizes the transfer of knowledge from the teacher network to the student network and is used to maintain the memory of knowledge in old tasks in the continuous learning branch. Denote the parameters of the teacher network as , and the parameters of the student network as . In the training of task , the student network is the model backbone network; the teacher network is the backbone network obtained in the training of task , and its parameters are fixed in the training of task . Given an input picture , the output feature distributions of the teacher network and the student network are respectively denoted as and and . The student network is expected to output a feature probability distribution similar to that of the teacher network, and the two probability distributions are made similar by minimizing the distillation loss. The distillation loss calculation formula is as follows:

[0044]

[0045]

[0046]

[0047] where is the temperature coefficient, which is used to control the smoothness of the probability distribution, represents the i-th feature vector , represents the i-th feature vector 。

[0048] (4) A self-supervised learning branch, which is used to minimize the contrast loss function to make the error between the feature vector and the feature vector within a second set range; the feature vector is the feature vector of the second data-augmented transformed image extracted by the contrast network. The second data-augmented transformed image and the first data-augmented transformed image are images obtained by preprocessing the same original image through different methods. The network structures of the contrast network and the student network are the same.

[0049] Self-supervised learning requires learning a representation learning model by automatically constructing similar instances and dissimilar instances. Through this model, similar instances are close in the projection space, while dissimilar instances are far apart in the projection space. Self-supervised learning enables the model to learn general feature representations that are independent of tasks and categories, reducing overfitting in specific tasks and thus reducing the performance loss of the model in new tasks. The contrast network is the same as the student network, and its model parameters do not participate in backpropagation. By weighted updating the parameters of the student network, the formula is , where is a weight between 0 and 1. The calculation formula of the contrast loss is as follows:

[0050]

[0051] where, The function is a similarity function, usually using cosine similarity .

[0052] (5) An attribute attention module, which is used to calculate the attribute encoding e of the feature vector using the attribute label and the cross-attention mechanism, so as to model the category attributes.

[0053] As Figure 2 shown, during the continuous learning process, the attribute label is dynamically added and used as the Query in the cross-attention. The attribute label is used as the parameter of the attribute attention module and is jointly trained and updated with other parameters in the model. The feature vector is used as the Key in the cross-attention. The overall calculation formula of the cross-attention layer is as follows:

[0054]

[0055] where, , , is a parameter in the cross-attention layer; is the dimension of the feature, is the number of attention heads in the cross-attention layer; are all process variables;

[0056] Subsequently, the output of the cross-attention layer obtains the attribute encoding e after layer normalization with residual and a linear layer.

[0057] During the process of continuous learning, to ensure the memory of the knowledge learned previously, the attribute tags trained in the old tasks do not participate in the update in the new tasks. That is, in the training of task the attribute tags are divided into and two parts, where is the attribute tag vector obtained from the training in the previous tasks and is not updated in the training of task ; is added to the model at the start of the training of task , and its initial value is sampled from the standard normal distribution and updated through backpropagation in the training of task . The number of attribute tags increases as the number of categories increases during the model training. According to the binomial formula, when the number of categories is , only attributes are needed to represent all categories. Considering the slow growth rate and ensuring that there are attribute tag increases in each task, the final formula for the number of labeled attributes is , r = 1, 2, … l.

[0058] (6) Classification module, including a linear classifier group composed of multiple category-independent parallel classifiers. The attribute encoding e of the feature vector is respectively input into each linear classifier and activated using sigmoid to calculate the classification scores for each category.

[0059] As Figure 1 shown, the category-independent parallel classifiers refer to a linear classifier group trained for each category in each task. The attribute encoding e is respectively input into each linear classifier and activated using sigmoid to calculate the classification scores for each category.

[0060] In this embodiment, the self-supervised learning branch is used to extract general feature representations that are irrelevant to the task, reduce overfitting in specific tasks, and thus reduce the performance loss of the model in new tasks; the knowledge distillation branch for continuous learning is used to maintain the memory of knowledge in old tasks; then, in the attribute attention module, learnable orthogonalized attribute tokens and the cross-attention mechanism are used to calculate the attribute encoding e of the input features, thereby modeling the category attributes. Finally, the attribute encoding e is respectively input into each category-independent parallel classifier and activated using sigmoid to calculate the classification scores for each category.

[0061] In this embodiment, only one feature extractor obtained from the training in the previous task is saved as the teacher network, which has a smaller memory consumption compared to saving some samples for replay and there is also no possibility of privacy leakage; at the same time, by modeling the category attributes, analysis can be carried out through the visualization of the feature space and the encoding space, and the model has relatively higher interpretability.

[0062] Furthermore, the backbone network is specifically a Vision Transformer (ViT) or a Convolutional Neural Network (CNN).

[0063] In this embodiment, ViT is composed of stacked self-attention modules, and the structures of the teacher network and the contrast network are the same as that of the student network. The teacher network and the contrast network are only used during training, and the teacher network and the contrast network are discarded during model inference, and only the backbone network as the student network is retained.

[0064] The backbone network can also adopt a Convolutional Neural Network (CNN), where the convolutional neural network here is a stacked residual block;

[0065] Furthermore, it further includes:

[0066] Attribute tokens After being orthogonalized, orthogonalized attribute tokens are obtained , which are input into the cross-attention layer. The orthogonalization process is achieved by adding a regularization term to the loss function. In this embodiment, a fast regularization term construction module based on local sensitive hashing is adopted. This regularization term is relatively softer and more inclined to achieve orthogonal penalty from the overall perspective of the angular distribution, reducing the model performance loss. Its specific calculation method is:

[0067] (1) Activate the calculated in the cross-attention layer using and normalize it in the batch dimension to obtain ;

[0068] (2) Calculate ;

[0069] (3) Take It is calculated by substituting the approximate formula of the regularization term derived from locality-sensitive hashing. The formula is as follows:

[0070]

[0071] In addition, the orthogonalization of the attribute markers can also use a conventional regularization term to replace, where is the attribute marker matrix, is the identity matrix, and the norm is the matrix 2-norm or the matrix Frobenius norm.

[0072] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image classification system based on category attribute modeling, characterized in that, comprising: An image preprocessing module for preprocessing the original image to obtain multiple data-augmented transformed images of the original image, and the preprocessing includes random cropping, or random scaling, or random color or normalization; Backbone network, which serves as the student network and is used to extract the feature vectors of the first data-augmented transformed images ; A knowledge distillation branch for continuous learning, which is used to minimize the distillation loss function so that the probability distribution error between the feature vector and the feature vector is within a first set range; the feature vector is the feature vector of the first data-augmented transformed image extracted by the teacher network, and the teacher network has the same network structure as the student network. In the current round of task training, the teacher network is the student network obtained in the previous task training; A self-supervised learning branch for minimizing the contrast loss function to make the error between the feature vector and the feature vector within a second set range; the feature vector is the feature vector of the second data-augmented transformed image extracted by the contrast network, and the second data-augmented transformed image and the first data-augmented transformed image are images obtained by preprocessing the same original image through different methods. The network structures of the contrast network and the student network are the same; An attribute attention module for calculating a feature vector using an attribute token and a cross-attention mechanism to obtain an attribute encoding e; The classification module includes a linear classifier group composed of multiple independently parallel classifiers for each category, and the feature vector attribute encoding e is input into each linear classifier respectively and sigmoid activation is used to calculate the classification scores for each category; Calculating the feature vector using the attribute marker and the cross-attention mechanism The attribute encoding e specifically includes: Dynamically add attribute tags and use it as the Query in the cross-attention layer, with the feature vector as the Key in the cross-attention, and calculate the output of the cross-attention layer: Among them, , , are parameters in the cross-attention layer; is the dimension of the feature, is the number of attention heads in the cross-attention layer; are all process variables; Output of cross-attention After layer normalization with residual and linear layer, the feature vector is obtained Attribute encoding e of; Among them, in any training, the attribute marker is divided into and two parts, where is the attribute marker vector obtained from training in the first tasks and is not updated during the training of tasks ; is added to the model at the start of the training of task , and its initial value is sampled from the standard normal distribution and updated through backpropagation during the training of task ; the number of attribute markers , p is the number of categories, r = 1, 2, … l; Calculating the feature vector using the attribute marker and the cross-attention mechanism The attribute encoding e also includes: the attribute marker After being orthogonalized, it is input into the cross-attention layer.

2. The image classification system based on category attribute modeling according to claim 1, characterized in that, The backbone network is specifically a Vision Transformer (ViT) or a Convolutional Neural Network (CNN).

3. The image classification system based on category attribute modeling according to claim 1, characterized in that, The calculation process of the distillation loss function is: where is the temperature coefficient, which is used to control the smoothness of the probability distribution, represents the i-th eigenvector , represents the i-th eigenvector .

4. The image classification system based on category attribute modeling according to claim 1, characterized in that, The calculation process of the contrastive loss function is: Among them, The function is the cosine similarity function.

5. The image classification system based on category attribute modeling according to claim 1, characterized in that, The result calculated in the cross-attention layer is used for activation and normalized in the batch dimension to obtain ; Calculation ; Substitute into the regular term approximation formula for calculation: 。