A Continuous Learning Method Based on Hypersphere Geometric Structure

By projecting visual features in the hyperspherical geometric space and designing corresponding loss functions and learning solutions, the problems of restricted feature evolution and catastrophic forgetting in actual scenarios in the existing technology are solved, and efficient adaptation to new categories and maintaining the classification performance of old categories are achieved, which improves the generalization and continuous learning effect of the model.

CN117114128BActive Publication Date: 2025-06-20FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310972905.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-06-20
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

The existing class-incremental learning methods based on hypersphere geometry are limited by specific sample sizes in actual scenarios, limiting the evolution of features, resulting in the model's classification performance of new data and prone to catastrophic forgetting.

Method used

A continuous learning method based on hyperspherical geometric structure is designed to achieve efficient adaptation to new categories and maintain classification performance of old categories by performing visual feature projection in hyperspherical geometric space, combining basic training schemes and incremental learning schemes. Specific steps include designing the tight loss function of the instance prototype and the inter-class prototype separation loss function, conducting prototype construction and adapting, and overcoming catastrophic forgetting through the instance prototype relational distillation scheme.

Benefits of technology

This method can improve the generalization of the model, efficiently adapt to new categories, and effectively overcome the catastrophic forgetting of old data, significantly improving the effectiveness of continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117114128B_ABST
    Figure CN117114128B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of machine learning, and specifically relates to a continuous learning method based on a hypersphere geometric structure. The method of the present invention includes: projecting visual features into the hypersphere space, constructing class prototypes in this lower-dimensional embedding space, and performing continuous learning; based on the hypersphere structure, a basic training scheme is designed, including an instance prototype compactness loss function to reduce the distance between classes, and an inter-class prototype separation loss function to maximize the separability between classes; an incremental learning scheme is designed, including prototype construction and adaptation strategies to effectively adapt to new classes; and an instance prototype relationship preservation distillation scheme to overcome the problem of catastrophic forgetting. The above methods have been experimentally verified on multiple image datasets, demonstrating the superiority of the methods. The present invention can help deep learning models have stronger adaptability to future data in incremental learning scenarios and contribute to overcoming the problem of catastrophic forgetting of old data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a continuous learning method based on a hypersphere geometric structure. Background Art

[0002] In the actual scenario of risk control, the model needs to classify risk points. As new risk points keep emerging, the model needs to be continuously updated to adapt to the new risk points. If the model is updated from scratch every time, frequent updates will bring huge time overhead and resource consumption. In addition, when directly using the new risk point data for model update, deep learning models are prone to catastrophic forgetting, that is, the model's recognition performance for old risk points is poor. As an important issue in the field of artificial intelligence, continuous learning aims to continuously learn new tasks without forgetting what has been learned in old tasks, so it aligns with our purpose. We mainly consider one important continuous learning paradigm for solution design, focusing on how to learn to classify new classes while maintaining the classification ability for old class data, which is generally called class-incremental learning (CIL). Currently, CIL methods are roughly divided into three categories: data-centric, model-centric, and algorithm-centric methods. Data-centric methods design sample selection and replay strategies to effectively utilize samples of old classes. Model-centric methods tend to expand some functional modules to protect old parameter modules as much as possible to retain the classification performance of the old model. However, this will introduce additional parameters during the training process. Algorithm-centric methods are exploration algorithms, such as distilling the knowledge of old models to maintain the classification performance for old classes or correcting the weight bias of classifiers. Although there are significant differences in specific implementation technologies, one common feature of most existing CIL methods is that they use Euclidean space as the output space. That is to say, they conduct continuous learning in Euclidean space. In recent years, some few-shot class-incremental learning (FSCIL) uses constructed class prototypes in hypersphere space. The core idea is to increase the separation angle of class prototypes to make the model more adaptable to new data. Compared with Euclidean space, the advantage of using hypersphere as the output space is that the classifier prototypes in hypersphere space have unit norms, which helps to overcome the catastrophic forgetting caused by classifier bias resulting from the imbalance between new and old classes. These FSCIL methods, considering that only a small number of new class samples are available for training, in order to avoid feature overfitting and catastrophic forgetting of the model, they adopt the strategy of freezing (not participating in subsequent parameter updates) the feature extractor, which will limit the evolution of features and thus limit the classification performance for new data. All in all, some existing solutions designed based on hypersphere geometry are limited by specific sample size scenarios and limit the possibility of obtaining evolutionary representations, thus greatly restricting their applicability in actual scenarios. Summary of the Invention

[0003] The objective of the present invention is to provide a continual learning method based on a hypersphere geometric structure, which can project visual features into a hypersphere geometric space where continual learning is constructed, can improve the generalization of the model, and can efficiently adapt to new classes and overcome the problem of catastrophic forgetting of old data during the incremental process.

[0004] The continual learning method based on the hypersphere geometric structure provided by the present invention uses a variety of new technical means, including the design of a training scheme for the basic hypersphere geometric structure, the design of a scheme for adapting to new classes and overcoming catastrophic forgetting of old classes during the incremental learning process; the specific steps are as follows:

[0005] (1) Design of the basic training scheme: By designing the basic training loss function of the model, on the one hand, the intra-class distance is reduced by designing the instance prototype compactness loss function, and on the other hand, the inter-class prototype separation loss function is designed to maximize the separability between classes;

[0006] (2) Design of the incremental learning scheme: Design a construction and adaptation strategy for prototypes to effectively adapt to new classes; design an instance prototype relationship distillation scheme to overcome the problem of catastrophic forgetting.

[0007] The specific process of the basic training scheme design described in step (1) is as follows:

[0008] Different from most existing continual learning schemes that use the Euclidean space as the output space, the present invention takes the output space as a (d - 1)-dimensional hypersphere, which is defined as follows:

[0009]

[0010] Project the visual feature x (original image) into this hypersphere space, which can be achieved through a feature extractor (for computer vision tasks, we select the residual network here) f Θ as the backbone network, and then project it into the hypersphere geometric space through a non-linear projection head g Φ (usually implemented by two fully connected layers combined with a ReLU non-linear layer) to obtain the final result of the normalized norm embedding z = g Φ (f Θ (x)). Θ and Φ respectively correspond to the model trainable parameters of the backbone network and the non-linear projection head. To describe the distribution of z, further consider using the von Mises-Fisher (vMF) distribution for modeling, which is defined as follows:

[0011] f d (z; μ, κ) = C d (κ) exp(κμ T z), (2)

[0012] where κ is the concentration parameter, μ is a vector with unit norm used to average directions, and C d (κ) is the normalization constant. Using the vMF distribution, it is convenient to model the joint distribution formed by embeddings of different classes. Assume the prototype matrix formed by the centers of each class embedding is M p : = [p1, p2, …, p C , where p c is the distribution center corresponding to class c, and C is the total number of classes. It has a normalized norm. Next, the embedding distribution of class c is described through the vMF distribution, and the probability distribution result is as follows:

[0013]

[0014] where κ c is the concentration parameter of class c. Then, for a given embedding z, its posterior probability of belonging to class c is:

[0015]

[0016] where M p is the prototype matrix. For simplicity, the same centering parameter is taken for different classes, so we have:

[0017]

[0018] Next, for a given n samples by maximizing the log-likelihood, we aim to make the distribution of each class concentrate around the corresponding class prototype:

[0019]

[0020] where c(i) represents the class to which the sample x i belongs, Θ, Φ correspond to the trainable parameters of the backbone network and the non-linear projection head mentioned above, C is the number of classes, and z i represents the embedding calculated for the i-th sample.

[0021] To optimize using backpropagation, after taking the logarithm, the instance prototype compactness loss function is as follows:

[0022]

[0023] In addition to maximizing compactness, we further design a method to maximize the separability between classes. The framework is as shown in the upper line of Figure 1 This can be achieved by maximizing the angles between class prototypes. For a given prototype matrix M p : = [p1, p2, …, p C, calculate the average angle between each pair to obtain the result:

[0024]

[0025] Furthermore, in order to unify the format with the instance prototype compactness loss function and facilitate optimization, Equation (8) is further transformed by the logarithmic exponent method to obtain the between-class prototype separation loss function:

[0026]

[0027] After each optimization process, the class prototypes need to be updated. Reusing all samples of each class to calculate the class prototypes requires a large computational overhead. Therefore, an exponential smoothing incremental update strategy is adopted here. Taking class c as an example, it is implemented by the following formula:

[0028]

[0029] Among them, is the embedding average of a batch of samples obtained currently, z c is the embedding of the samples of class c, |z c | is its corresponding quantity, (z c ) i is the embedding of the i-th sample of class c. The parameter γ is an adjustable value close to 1 to control the update speed.

[0030] By optimizing the objective functions of Equations (7) and (9) above and combining the class prototype update steps of Equation (10), the process of basic training can be completed, that is, an embedding that maximizes the within-class compactness and maximizes the between-class separation on the hypersphere geometry can be obtained.

[0031] The design of the incremental learning scheme described in step (2) has the following specific process:

[0032] Based on the above basic training framework, the present invention further designs a scheme for the incremental learning scenario, which needs to overcome catastrophic forgetting on the old data while adapting to the new class data. Specifically: for the newly arrived class, prototype construction is performed, which is achieved by calculating the class center of the new class in the current task. For simplicity, consider performing prototype construction for the new class c as the initialization of its prototype, corresponding to the result of the following formula:

[0033]

[0034] Among them, represents the number of sampled samples. After that, the prototype adaptation step process is performed.

[0035] In the adaptation step, fix the prototype positions of the old classes to prevent drastic changes in the prototypes of the old classes during the process of adapting to the new classes, which may lead to catastrophic forgetting. For the prototypes of the new classes, on the one hand, consider modifying Equation (9) to achieve the purpose of separating the prototypes of both the old and new classes, and obtain the following equation:

[0036]

[0037] where C o and C n represent the numbers of the old and new classes respectively; on the other hand, minimize the distance between the embeddings of the new classes and their corresponding prototypes through the following equation:

[0038]

[0039] where represents the total number of samples, denotes the new samples among them, and denotes the old samples from the Replay buffer among them.

[0040] In addition to the above prototype construction and adaptation processes, since only a limited number of old samples can be retained in the actual process, it is necessary to design a further instance-prototype relationship distillation scheme to overcome catastrophic forgetting. By using the model trained on the old task, its corresponding old model parameters are and to obtain the embeddings corresponding to the old samples as Meanwhile, obtain the embeddings of the old samples after passing through the new model (whose corresponding coefficients are Θ n and Φ n ) as After that, the instance-prototype relationship distillation loss function can be calculated through the following equation:

[0041]

[0042] where M po =, p o1 , p o2 , …, p oC represents the prototype matrix corresponding to the prototypes of the old samples. The operation <:,:> represents the vector obtained by taking the inner product of the embedding with each embedding in the matrix, and the outermost norm is generally taken as the L2 norm.

[0043] To further utilize the information output by the intermediate layer of the model, the distillation scheme additionally adds a distillation term:

[0044]

[0045] The overall incremental process steps are as Figure 2As shown, by combining the construction and adaptation strategies of the prototype, corresponding to equations (12) and (13), and the instance-prototype relationship distillation scheme, corresponding to equations (14) and (15), the final result can be obtained by optimizing the model through gradient descent.

[0046] The present invention at least includes the following beneficial effects:

[0047] (1) The designed basic training scheme can project visual features into a hyperspherical geometric space where continuous learning is constructed. By adding inductive biases of maximizing intra-class compactness and maximizing inter-class separability to the loss function, the generalization ability of the model can be improved, thus better enhancing the effect of continuous learning;

[0048] (2) The designed incremental learning scheme, the design of the prototype construction and adaptation scheme, can efficiently adapt to new categories. At the same time, the instance-prototype relationship distillation scheme can overcome the catastrophic forgetting of the model for old categories.

[0049] Other advantages, objectives and features of the present invention will be partially reflected by the following description and partially understood by those skilled in the art through the research and practice of the invention. Brief Description of the Drawings

[0050] Figure 1 It is a framework diagram of the present invention.

[0051] Figure 2 It shows a summary of the specific scheme of the incremental learning process.

[0052] Figure 3 It shows the experimental results on two datasets. Detailed Embodiment

[0053] The following further detailed description of the present invention is made in conjunction with the drawings, so that those skilled in the art can implement it according to the description in the specification.

[0054] It should be understood that the terms such as "having", "comprising" and "including" used herein do not exclude the existence or addition of one or more other elements or their combinations.

[0055] I. Design of Basic Training Scheme

[0056] As Figure 1 shown in the above part, the embodiment of the present invention provides a basic training method for obtaining embeddings that maximize intra-class compactness and maximize inter-class separation on hypersphere geometry. Compared with the existing continuous learning schemes that mostly use Euclidean space as the output space, the present invention takes the output space as a hypersphere of d - 1 dimensions, defined as follows

[0057]

[0058] Project the visual feature x (the original image) into this hypersphere space, which can be achieved by a feature extractor (for computer vision tasks, we choose a residual network here) f Θ As the backbone network, and then further project it into the hypersphere geometric space through a non-linear projection head g Φ (usually implemented by two fully connected layers combined with a ReLU non-linear layer), and finally the embedding z with a normalized norm is obtained, z = g Φ (f Θ (x)). Θ and Φ correspond to the trainable parameters of the backbone network and the non-linear projection head respectively. To describe the distribution of z, further consider using the von Mises-Fisher (vMF) distribution for modeling, and its definition is as follows:

[0059] f d (z; μ, κ) = C d (κ) exp(κμ T z), (2)

[0060] where κ is the concentration parameter, μ is a vector with unit norm used to average the direction, and C d (κ) is the normalization constant. Using the vMF distribution, it is convenient to model the joint distribution formed by embeddings of different classes. Assume that the prototype matrix formed by the centers of each class embedding is M p : = [p1, p2,..., p C , where p c is the distribution center corresponding to class c, and C is the total number of classes. It has a normalized norm. Next, describe the embedding distribution of class c through the vMF distribution, and the probability distribution result is as follows:

[0061]

[0062] Then for a given embedding z, the posterior probability that it belongs to class c is

[0063]

[0064] For simplicity, we take the same concentration parameter for different classes, then

[0065]

[0066] Next, for a given n samples By maximizing the log-likelihood,

[0067]

[0068] where c(i) represents the sample x iBelonging to a category to achieve the purpose of concentrating the distribution of each category around the corresponding category prototype. To optimize using backpropagation, we take the logarithm and obtain the instance prototype compactness loss function as follows

[0069]

[0070] In actual deployment, the distribution range of each category embedding is restricted by setting κ to 10. In addition to maximizing compactness, we further design a method to maximize the separability between classes. The framework is as shown in the upper row of Figure 1 This can be achieved by maximizing the angle between category prototypes. For a given prototype matrix M p := [p1, p2, …, p C , calculate the average angle between each pair of them to obtain the result

[0071]

[0072] Furthermore, in order to unify the format with the instance prototype compactness loss function and facilitate optimization, Equation (8) is further transformed by the logarithmic exponential method to obtain the inter-class prototype separation loss function

[0073]

[0074] After each optimization process, the category prototype needs to be updated. Recalculating the category prototype using all samples of each class requires a large computational overhead. Therefore, an exponential smoothing update strategy is adopted here to achieve this. Taking category c as an example, it is achieved through the following formula

[0075]

[0076] where is the embedding average of a batch of samples obtained currently, z c is the embedding of samples of category c, |z c | is its corresponding quantity, and (z c ) i is the embedding of the i-th sample of category c. During the implementation process, γ is selected by the incremental method. Initially, it is set to 0, and then it is increased at a speed of 0.05 until it reaches the set upper bound of 0.95. This implementation method can help maintain a relatively large update speed in the initial stage and gradually decay to a relatively small update speed later for better convergence

[0077] The specific optimization objective is achieved by combining Equation (7) and (9), and the result is as follows

[0078]

[0079] where λ cIt is used to balance two loss functions and is set to 2 in the actual implementation process. During the actual implementation process, after obtaining the samples, their corresponding positive samples are obtained by augmentation. The two terms in Equation (11) are calculated using these two sets of samples respectively. Then, by optimizing Equation (11) and combining the class prototype update step in Equation (10), the process of basic training is completed, that is, an embedding that maximizes the within-class compactness and maximizes the between-class separation is obtained on the hypersphere geometry.

[0080] II. Incremental Learning Scheme Design

[0081] Based on the above basic training framework, the present invention further designs a scheme for the incremental learning scenario, which needs to overcome catastrophic forgetting on the old data while adapting to the new class data. The overall process is as Figure 1 shown in the next part below, mainly including two parts: prototype construction and adaptation, and instance-prototype relationship distillation. The more specific implementation process is as Figure 2 shown.

[0082] Specifically: For the newly arrived class, prototype construction is performed, which is achieved by calculating the class center of the new class in the current task. For simplicity, consider performing prototype construction for the new class c, corresponding to the result of the following formula

[0083]

[0084] where represents the number of sampled samples. Then, the prototype adaptation process is performed.

[0085] In the adaptation step, the prototype positions of the old classes are fixed to prevent drastic changes in the prototypes of the old classes during the process of adapting to the new class and causing catastrophic forgetting. For the prototypes of the new class, on the one hand, consider modifying Equation (9) to achieve the purpose of separating the prototypes of the old and new classes, and obtain the following formula

[0086]

[0087] where C o and C n represent the numbers of the old classes and the new class respectively; on the other hand, the distance between the embedding of the new class and its corresponding prototype is minimized through the following formula

[0088]

[0089] where represents the total number of samples, represents the new samples among them, represents the old samples from the Replay buffer among them.

[0090] In addition to the above prototype construction and adaptation processes, since only a limited number of old samples can be retained in the actual process, a further instance-prototype relationship distillation scheme needs to be designed to overcome catastrophic forgetting. By using the model trained on the old task, whose corresponding old model coefficients are and the embeddings corresponding to the old samples are obtained as Meanwhile, the old samples are passed through the new model, whose corresponding coefficients are Θ n and Φ n , and the resulting embeddings are After that, the instance-prototype relationship distillation loss function can be calculated by the following formula

[0091]

[0092] where M po = [p o1 , p o2 , …, p oC represents the prototype matrix corresponding to the prototypes of the old samples, where the operation <:,:> represents the vector obtained by taking the inner product of the embedding with each embedding in the matrix, and the outermost norm is generally taken as the L2 norm.

[0093] To further utilize the information output by the middle layer of the model, the distillation scheme additionally adds a distillation term

[0094]

[0095] In the actual process, the combined parameters λ c = 2, λ rp = 1 and λ fd = 1, and the total training loss function is

[0096]

[0097] It is obtained by combining equations (12) and (13), as well as the instance-prototype relationship distillation scheme, corresponding to equations (14) and (15). The model can be optimized by using gradient descent. After the model training is completed, representative samples are selected to construct In the actual process, the nearest neighbor selection strategy is adopted, and the samples closest to (angular distance) the prototype are selected to retain the information of the old samples as much as possible.

[0098] This application also provides a verification experiment to further prove the technical effects of this application.

[0099] To verify the performance of this method in continual learning, a class-incremental (CIL) setting was adopted, tested on multiple image datasets, and different task partitioning methods were set. The evaluation task set was partitioned into multiple tasks by category (randomly). The first task contained half of the number of categories, and then new categories were gradually added in small increments (i.e., a smaller number of categories).

[0100] CIFAR100 dataset (extracted from "Alex Krizhevsky, Geoffrey Hinton, et al. 2009. Learning multiple layers of features from tiny images. (2009)."): It contains 60,000 images with a resolution of 32 * 32 (length * width) pixels, a total of 100 categories, with 600 images in each category. The training set and the test set are partitioned in a ratio of 5:1, that is, each category contains 500 training data and 100 test data.

[0101] Tiny-ImageNet dataset (extracted from "Ya Le and Xuan Yang. 2015. Tiny imagenet visual recognition challenge. CS231N 7,7 (2015), 3."): It contains 100,000 training images with a resolution of 64 * 64 (length * width) pixels, a total of 200 categories, with 500 images in each category. Each category in the test set contains 50 images.

[0102] To verify the superiority of this method, under the above incremental settings and two publicly available datasets, it was compared with the following existing incremental learning methods. Specifically, several different existing methods were compared, including the parameter regularization method EWC (extracted from "James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences 114, 13 (2017), 3521–3526."), the distillation method LwF (extracted from "Zhizhong Li and Derek Hoiem. 2017. Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence 40, 12 (2017), 2935–2947."), iCaRL (extracted from "Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. 2017. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2001–2010."), UCIR (extracted from "Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. 2019. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE / CVF conference on Computer Vision and Pattern Recognition.831–839.”), PODNet (extracted from “Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. 2020. Podnet: Pooled outputs distillation for small-tasks incremental learning. In European Conference on Computer Vision. Springer, 86–102.”), the model correction method BiC (extracted from “Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. 2019. Large scale incremental learning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 374–382”), WA (extracted from “Bowen Zhao, Xi Xiao, Guojun Gan, Bin Zhang, and Shu-Tao Xia. 2020. Maintaining discrimination and fairness in class incremental learning. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 13208–13217.”) and the parameter expansion method Foster (extracted from “Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. 2022. Foster: Feature boosting and compression for class-incremental learning. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXV. Springer, 398–414.”).

[0103] In addition, the present invention creates two baselines for comparison: a lower-bound baseline that does not use any incremental learning techniques, called Finetune. Another baseline that only uses the random sample replay technique, called Replay. In particular, iCaRL is further divided into two categories according to the classifier type, iCaRL-CNN represents the use of a linear classifier, and iCaRL-NCM represents the use of a nearest neighbor classifier.

[0104] In this embodiment, the final accuracy and the average accuracy are used to measure the performance of each algorithm in the continuous learning process. Among them, the final accuracy is calculated by the total classification accuracy (Top-1 accuracy) on each class; the average accuracy is obtained by averaging the classification accuracies of the models after each task, which reflects the average performance of the continuous learning method in the incremental process.

[0105] The experimental results are as Figure 3 shown. The method of the present invention (C-HPN) is superior to the existing methods in multiple stage tasks on the two datasets of TinyImigeNet and CIFAR100. Taking the experimental results on the TinyImageNets dataset as an example, it corresponds to Figure 3 the results from top to bottom on the right side (5, 10, 20 incremental steps). It can be seen that the results of the hollow circle solid line (the method of the present invention) are all higher than the results of the existing optimal methods, and the final result is at least two percentage points higher, which confirms that C-HPN can effectively solve catastrophic forgetting in CIL. In addition, note a trend in the initial few tasks that, compared with other methods, the accuracy of C-HPN drops significantly slower. For example, on the CIFAR100 dataset, in the case of 10 incremental steps and 20 incremental steps, the results obtained by the hollow circle solid line (the method of the present invention) for the first three points are all at least 2 points higher than other methods; on the TinyImageNet dataset, it is even more obvious that after the first three increments, the results of the hollow circle solid line (the method of the present invention) are more than 5 points higher than other methods, which indicates that the model of the present invention has better generalization ability.

[0106] The method of the present invention has obtained consistent results on CIFAR100 and TinyImageNet. The obtained average accuracy results are shown in Table 1. The method of the present invention (C-HPN) is superior to the current existing methods in terms of the average accuracy on the CIFAR100 and TinyImageNet datasets. In summary, the method of the present invention has good adaptability to new classes and can effectively overcome catastrophic forgetting, that is, it can perform well on new tasks and still maintain good classification results on old tasks.

[0107] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrated and described examples here.

[0108] Table 1, Average accuracy on CIFAR-100 and TinyImageNet

[0109]

Claims

1. A continuous learning method based on a hypersphere geometric structure, characterized in that, Design of a training scheme including a basic hypersphere geometric structure, design of a scheme for adapting to new classes and overcoming catastrophic forgetting of old classes during the incremental learning process; The specific steps are as follows: (I) Design of the basic training scheme: Design the basic training loss function of the model, including designing the instance prototype compactness loss function to reduce the intra-class distance, and designing the inter-class prototype separation loss function to maximize the separation between classes; (II) Design of the incremental learning scheme: Include designing the construction and adaptation strategies of prototypes to effectively adapt to new classes; Design the instance prototype relationship distillation scheme to overcome the problem of catastrophic forgetting; The specific process of the basic training scheme design described in step (I) is as follows: Take a hypersphere with a d-1 dimensional output space, defined as follows: Project the visual feature x, i.e., the original image, into this hypersphere space through a feature extractor f Θ As the backbone network, and then through a non-linear projection head g Φ Further project it into the hypersphere geometric space to achieve, where Θ and Φ are the learnable parameters of the model, and finally obtain the embedded normalized norm: z = g Φ (f Θ (x)); where f Θ Select the residual network, and g Φ Is implemented by combining two fully connected layers with a ReLU non-linear layer; To describe the distribution of z, use the von Mises-Fisher vMF distribution for modeling, which is defined as follows: f d f(z; μ, κ) = C d f(κ) exp(κμ T z), (2) where k is the concentration parameter, μ is a vector with unit norm used to average directions, and C d (κ) is the normalization constant; the vMF distribution is used to model the joint distribution formed by the embeddings of different classes; assume that the prototype matrix formed by the centers of each class embedding is M p : = [p1, p2, …, p C , where p c is the distribution center corresponding to class c, which has a normalized norm; the embedding distribution of class c is described by the vMF distribution to obtain the probability distribution result: where κ c represents the concentration parameter of the class prototype, and for a given embedding z, the posterior probability that it belongs to class c is: M p is the prototype matrix. For simplicity, the same centering parameter is used for different classes, so we have: Next, for a given sample By maximizing the log-likelihood, the distribution of each class is concentrated around the corresponding class prototype: where c(i) represents the class to which the sample x i belongs, Θ and Φ are learnable parameters of the model, C is the total number of classes, and z i represents the embedding calculated for the i-th sample; To use backpropagation for optimization, take the logarithm of it to obtain the instance prototype compactness loss function as follows: In addition to maximizing compactness, the separability between classes is further maximized, which is achieved by maximizing the angles between class prototypes. Specifically, for a given prototype matrix M p : = [p1, p2, …, p C , calculate the average angle between each pair of them to obtain the result: To unify the format with the instance prototype compactness loss function and facilitate optimization, further transform equation (8) through the logarithmic exponential method to obtain the inter-class prototype separation loss function: After each optimization process, update the class prototypes, specifically implemented using an exponential smoothing incremental update strategy; for class c, it is specifically achieved through the following formula: Among them, is the embedding average of a batch of samples obtained currently, where z c is the embedding corresponding to class c, (z c ) i corresponds to the i-th sample, |z c | represents the total number of samples; the hyperparameter γ is used to control the update speed and is set to a value close to 1; By optimizing the objective functions of equations (7) and (9) above, combined with the class prototype update step of equation (10), the process of basic training is completed, that is, an embedding with maximized intra-class compactness and maximized inter-class separation is obtained on the hypersphere geometry.

2. The continuous learning method based on a hypersphere geometric structure according to claim 1, characterized in that, The specific process of the incremental learning scheme design described in step (II) is as follows: For newly arrived classes, perform prototype construction, which is achieved by calculating the class center of the new class in the current task. For simplicity, consider constructing a prototype for the new class c, corresponding to the following result: Among them, represents the number of sampled samples; then, the prototype adaptation step process is executed; In the adaptation step, fix the prototype positions of the old classes to prevent drastic changes in the prototypes of the old classes during the process of adapting to new classes and causing catastrophic forgetting; for the prototypes of the new classes, on the one hand, modify equation (9) to achieve the purpose of separating the prototypes of both new and old classes, obtaining the following formula: where C o and C n represent the numbers of the old and new categories, respectively; on the other hand, the distance between the new category embedding and its corresponding prototype is minimized by the following formula: Among them, represents the total number of samples, represents the new samples among them, represents the old samples from the replay buffer among them; Furthermore, a design instance prototype relationship distillation scheme is proposed to overcome catastrophic forgetting; specifically, by using the model trained on the old task, the corresponding old model coefficients are and to obtain the embedding corresponding to the old sample as Meanwhile, the embedding obtained by passing the old sample through the new model is The corresponding coefficients of the new model are Θ n and Φ n ; the instance prototype relationship distillation loss function is calculated by the following formula: Among them, M po = [p o1 , p o2 , …, p oC represents the prototype matrix corresponding to the prototype of the old samples. The operation <:,:> represents the vector obtained by taking the inner product of the embedding with each embedding in the matrix, and the outermost norm is taken as the L2 norm; To further utilize the information output by the middle layer of the model, the distillation scheme adds an extra distillation term: Combining the prototype construction and adaptation strategies, corresponding to equations (12) and (13), and the instance prototype relationship distillation scheme, corresponding to equations (14) and (15), optimize the model through gradient descent to obtain the final result.

Citation Information

Patent Citations

  • Image recognition method for incremental learning based on network expansion and memory recall mechanism

    CN114118207A

  • Image classification method based on multilevel adaptive feature fusion class incremental learning

    CN114612721A