Language guidance-based long-tail incremental image recognition method and system, and medium

By introducing language prior information and pseudo-feature spatial distribution in long-tail category incremental learning, a two-stage training method is adopted to solve the problems of new data overfitting and old task forgetting, and the performance and generalization ability of the model are improved.

CN120088797APending Publication Date: 2025-06-03SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510009027.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the incremental learning of long-tail category, the existing technology is prone to degradation in model performance due to overfitting of new data due to few sample categories, and due to the lack of old task data, the model will experience severe catastrophic forgetting of old tasks.

Method used

Using a language-guided method, by introducing language prior information, the pseudo-feature spatial distribution is constructed, and two-stage training is carried out in the image-text feature alignment and balance training stages, the generalization ability of image representation is enhanced and overfitting and forgetting problems are alleviated.

Benefits of technology

It effectively solves the problems of overfitting new data and catastrophic forgetting of old tasks in incremental learning of long-tail categories, and improves the performance and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088797A_ABST
    Figure CN120088797A_ABST
Patent Text Reader

Abstract

The invention discloses a long-tail incremental image recognition method and system based on language guidance and a medium. The method comprises the following steps: constructing a long-tail incremental image recognition model; obtaining multi-batch long-tail distribution training data; a fixed language template is designed, and corresponding category text description is generated through a large language model according to category labels of new category images in each batch of long-tail distribution training data; inputting multiple batches of long-tail distribution training data and corresponding category text descriptions into a long-tail incremental image recognition model in batches, and training by adopting a two-stage training mode to obtain a trained long-tail incremental image recognition model; and inputting the to-be-recognized long-tail distribution data into the trained long-tail incremental image recognition model to obtain a recognition result. According to the method, by introducing the language prior information, the problems of new data few sample category overfitting and old task knowledge disastrous forgetting in a long-tail category incremental learning scene can be solved at the same time, the performance of the model is improved, and the accuracy of long-tail incremental image recognition is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of long-tail incremental learning, and particularly relates to a language-guided long-tail incremental image recognition method, system and medium. Background Art

[0002] Incremental Learning is a machine learning method aiming to enable a model to continuously learn new knowledge over time without forgetting the knowledge already learned. The core challenge of incremental learning is the problem of catastrophic forgetting, that is, the model often forgets the knowledge learned in previous tasks when learning new tasks. Common incremental learning methods include: (1) Rehearsal-based methods, that is, when the storage space permits, a small number of representative old task samples are retained and trained together with new data when training new tasks, so as to keep the model's memory of old tasks; (2) Knowledge distillation-based methods, that is, the old model and the new model are used as the teacher model and the student model respectively, and through knowledge distillation, the knowledge in the old model is transferred to the new model; (3) Regularization-based methods, that is, by adding a regularization term during the training process of new tasks, the parameters of the model are kept as unchanged as possible to retain the knowledge in old tasks; (4) Based on the structure of a dynamic network, that is, new network branches are added when a new task comes, and the previous knowledge is transferred to the new structure to avoid forgetting.

[0003] Long-Tailed Learning focuses on training models on datasets with imbalanced class distributions. The classes in long-tailed distribution datasets are usually divided into multi-sample classes and few-sample classes, where multi-sample classes contain a large number of samples while few-sample classes have scarce samples. The goal of long-tailed learning is to enable the model to have good generalization ability for few-sample classes under the long-tailed distribution of training data and avoid overfitting problems during training. Common long-tailed learning methods include: (1) Resampling methods, that is, oversampling to increase the number of samples in few-sample classes to reduce the imbalance between classes during training; (2) Reweighting methods, that is, assigning higher weights to few-sample classes in the training loss function to make the model pay more attention to the learning of minority classes. (3) Data augmentation methods: Using data augmentation techniques to increase the diversity of few-sample class samples.

[0004] In the field of long-tail class incremental learning, in the prior art, some constraints are added during training to alleviate the problems of scarce samples in few-shot classes of new tasks and catastrophic forgetting of old tasks, including: (1) a two-stage decoupled learning method, which is based on traditional incremental learning algorithms and adds a second-stage balanced training during the training of new tasks. By learning a weight mapping layer, the network output of unbalanced classes is fine-tuned; (2) a gradient balance-based method, which adjusts the model gradients during training to reduce the gradient bias of unbalanced classes during the network update process.

[0005] However, the above prior art also has drawbacks: on the one hand, although the existing class incremental learning methods alleviate catastrophic forgetting to a certain extent, when the dataset has an unbalanced long-tail distribution, it is prone to overfitting to the few-shot classes at the tail, resulting in a serious decline in model performance; this is because existing methods generally assume that the dataset has a balanced distribution. On the other hand, although the existing long-tail learning methods alleviate the problem of intra-task class imbalance training to a certain extent, due to the lack of old task data, the model will suffer serious catastrophic forgetting of old tasks. On yet another hand, although the existing methods designed for long-tail incremental learning scenarios can improve the model performance to a certain extent compared with directly applying traditional class incremental learning algorithms, when constrained on a limited number of long-tail image datasets, the supervision information and knowledge provided by the images are limited, which also affects the model performance. Summary of the Invention

[0006] Aiming at the problem that the above prior art is limited by the supervision of limited unbalanced image data, the present invention provides a language-guided long-tail incremental image recognition method, system and medium, which introduce language prior information and construct a pseudo-feature space distribution, and simultaneously solve the problems of overfitting of few-shot classes of new data and catastrophic forgetting of old task knowledge in the long-tail class incremental learning scenario, improve the model performance, and improve the accuracy of long-tail class incremental learning.

[0007] To achieve the above object, on the one hand, the present invention adopts the following technical solutions:

[0008] The language-guided long-tail incremental image recognition method includes the following steps:

[0009] Construct a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network, and a classification head; the image feature backbone network is used to extract the visual image representation of the input image; the pre-trained text encoder is used to extract the semantic text representation of the input image; the feature remapping network is used to learn an enhanced image representation; the classification head classifies the image based on the enhanced image representation;

[0010] Obtain multi-batch long-tailed distribution training data; each batch of long-tailed distribution training data contains a fixed number of new class images and their corresponding class labels;

[0011] Design a fixed language template and use a large language model to generate corresponding class text descriptions for the class labels of new class images in each batch of long-tailed distribution training data;

[0012] Input multi-batch long-tailed distribution training data and corresponding class text descriptions into the long-tailed incremental image recognition model batch by batch, and train it using a two-stage training method to obtain a trained long-tailed incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; in the image-text feature alignment stage, generate semantic text representations of new class images and perform feature alignment with the visual image representations of the images to obtain image-text features; in the balanced training stage, construct a pseudo-feature space distribution and enhance the image-text features;

[0013] Input the long-tailed distribution data to be recognized into the trained long-tailed incremental image recognition model to obtain the recognition result.

[0014] As a preferred technical solution, the training steps of the image-text feature alignment stage are as follows:

[0015] Input the class text descriptions corresponding to the new class images in each batch of long-tailed distribution training data into the pre-trained text encoder batch by batch to generate multiple semantic text representations corresponding to the new class images;

[0016] Input the new class images in the same batch of long-tailed distribution training data into the image feature backbone network to extract the visual image representations corresponding to the new class images;

[0017] Calculate the average value on each dimension of the multiple semantic text representations of the new class images to obtain the semantic text representation center of the new class images;

[0018] Use the contrastive learning method to perform feature alignment between the visual image representation of the new class images and the semantic text representation center of the new class images to obtain image-text features.

[0019] As a preferred technical solution, the use of the contrastive learning method to perform feature alignment between the visual image representation of the new class images and the semantic text representation center of the new class images is specifically as follows:

[0020] Calculate the feature cosine similarity between the visual image representation of the new class images and the semantic text representation center of the new class images as the text confidence of the new class images;

[0021] Perform feature alignment on the text confidence of the new class and the class label vector through the cross-entropy loss function to obtain image-text features.

[0022] As a preferred technical solution, the training steps in the balance training stage are as follows:

[0023] Freeze the parameters of the image feature backbone network, then extract the visual image representations of the new-class images in each batch of long-tailed distribution training data, and calculate the image representation center of the new-class images;

[0024] Calculate the feature covariance matrix of the new-class images based on the visual image representations and the image representation center of the new-class images;

[0025] Construct a pseudo-feature space distribution using the image representation center and the corresponding feature covariance matrix of the new-class images;

[0026] Sample an image pseudo-feature sample from the pseudo-feature space distribution of the new-class images, and at the same time sample a text feature sample from the semantic text representations of the new-class images, and randomly assign a modality weight, then perform linear weighted interpolation to obtain a mixed modality feature;

[0027] Pass the mixed modality feature through a feature remapping network and an image classification head to obtain the corresponding classification output;

[0028] Use the cross-entropy loss of the weighted modality to supervise the learning in the balance training stage.

[0029] As a preferred technical solution, calculating the feature covariance matrix of the new-class images based on the visual image representations of the new-class images is specifically as follows:

[0030] Obtain the image representation center of the new-class images:

[0031]

[0032] where F c is the image representation center of the new-class images with class label c, k is the number of visual image representations of the new-class images, is the i-th visual image representation of the new-class images with class label c;

[0033] Calculate the feature covariance matrix of the new-class images:

[0034]

[0035] where K c is the feature covariance matrix of the new-class images with class label c, and T is the matrix transpose.

[0036] As a preferred technical solution, the constructed pseudo-feature normal distribution satisfies the multivariate normal distribution.

[0037] As a preferred technical solution, the mixed modal feature obtained by linear weighted difference is expressed as:

[0038] F mix = αS t + (1 - α)S i ,

[0039] where F mix is the mixed modal feature, α is the modal weight, and S t is the text feature sample sampled from the semantic text representation of the new class image, and S i is the pseudo-feature sample sampled from the pseudo-feature space distribution of the new class image.

[0040] As a preferred technical solution, the calculation method of the cross-entropy loss of the weighted modality is:

[0041] Obtain the mixed modal label of the mixed modal feature:

[0042] y mix = αy t + (1 - α)y i ,

[0043] where y mix is the mixed modal label vector of the mixed modal feature, α is the modal weight, y t is the one-hot encoded vector of the class label of the sampled text feature sample, and y i is the one-hot encoded vector of the class label of the sampled image pseudo-feature sample;

[0044] Based on the number of new class images in each batch of long-tail distribution training data, calculate the long-tail distribution weight of the class label:

[0045]

[0046] where w c is the long-tail distribution weight of class label c, C is the number of class labels of new class images in each batch of long-tail distribution training data, k c is the number of new class images of class label c, and γ is a hyperparameter;

[0047] Calculate the cross-entropy loss of the weighted modality:

[0048]

[0049] where L mix is the cross-entropy loss of the weighted modality, w i is the long-tail distribution weight of class label i, y mix,i is the mixed modal label of class label i, and p mix,iIt is the output value of the classification head for the mixed-modal feature category label i.

[0050] On the other hand, the present invention provides a language-guided long-tail incremental image recognition system, which is applied to the above-mentioned language-guided long-tail incremental image recognition method, and includes a model construction module, a data acquisition module, a text description module, a model training module, and an incremental recognition module;

[0051] The model construction module is used to construct a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network, and a classification head; the image feature backbone network is used to extract the visual image representation of the input image; the pre-trained text encoder is used to extract the semantic text representation of the input image; the feature remapping network is used to learn the enhanced image representation; the classification head classifies the image based on the enhanced image representation;

[0052] The data acquisition module is used to acquire multi-batch long-tail distribution training data; each batch of long-tail distribution training data contains a fixed number of new category images and their corresponding category labels;

[0053] The text description module is used to design a fixed language template and generate corresponding category text descriptions for the new category images in each batch of long-tail distribution training data through a large language model according to the category labels;

[0054] The model training module is used to input the multi-batch long-tail distribution training data and the corresponding category text descriptions into the long-tail incremental image recognition model batch by batch, and train it in a two-stage training manner to obtain a trained long-tail incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; the image-text feature alignment stage generates the semantic text representation of the new category image and performs feature alignment with the visual image representation of the image to obtain image-text features; the balanced training stage constructs a pseudo-feature space distribution and enhances the image-text features;

[0055] The incremental recognition module is used to input the long-tail distribution data to be recognized into the trained long-tail incremental image recognition model to obtain the recognition result.

[0056] On yet another aspect, a computer-readable storage medium is provided, storing a program, which when executed by a processor, implements the above-mentioned language-guided long-tail incremental image recognition method.

[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0058] A long-tail incremental image recognition method based on language guidance proposed by the present invention can simultaneously solve the problems of overfitting of new data few-shot classes and catastrophic forgetting of old task knowledge in the long-tail class incremental learning scenario by introducing language prior information. On the one hand, aiming at the problem that existing class incremental learning methods are prone to overfitting to the tail few-shot classes on long-tail distribution datasets, resulting in a serious decline in model performance, the present invention alleviates the overfitting problem by obtaining rich class text semantic features and using cross-modal alignment and cross-modal enhancement strategies to enhance the representation generalization of few-shot image classes. On the other hand, aiming at the problem that existing long-tail learning methods will cause serious catastrophic forgetting of old tasks due to the lack of old task data, the present invention constructs the pseudo-feature space distribution of old tasks and obtains the text description of old tasks, so as to jointly learn to alleviate the forgetting problem. Finally, compared with existing methods, the present invention breaks through the single unbalanced image modality and uses text priors to improve the performance upper limit of the model, thereby improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0060] Figure 1 It is a flowchart of the long-tail incremental image recognition method based on language guidance according to the embodiment of the present invention.

[0061] Figure 2 It is a block diagram of the long-tail incremental image recognition system based on language guidance according to the embodiment of the present invention.

[0062] Figure 3 It is a structural diagram of an electronic device according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0064] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0065] In the image recognition task of long-tailed class incremental learning, the model faces the problem of overfitting new data and catastrophic forgetting of old data, that is, the model forgets historical experience knowledge when it is only exposed to new data, and is prone to overfitting to new data. Under the condition of unbalanced long-tail distribution of new data, this situation will be more serious, and the model will overfit to the tail few-sample category, thereby affecting the generalization ability and stability of the model. In view of the problem that the above-mentioned prior art is limited by limited unbalanced image data supervision, the present invention explores a language-guided long-tail incremental image recognition method. The motivation of the invention mainly comes from two key advantages of natural language models: (1) After the pre-trained visual language model, the specific category text description has rich category semantic features, which can enhance the tail few-sample category features of the long-tail data distribution, so that the model can learn richer generalization ability; (2) The text description of the previous task can serve as the supervision information of the old task, which can be used to reconstruct the previous feature space, thereby alleviating the problem of catastrophic forgetting. The present invention proposes a language-guided long-tail incremental image recognition method, which utilizes the prior knowledge of language modality to enhance image representation, and solves the problems of overfitting of new data and few-sample categories and catastrophic forgetting of old task knowledge encountered in long-tail category incremental learning.

[0066] See also Figure 1 In one embodiment of the present application, a language-guided long-tail incremental image recognition method is provided, comprising the following steps:

[0067] S1. Construct a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network and a classification head; wherein, the image feature backbone network is used to extract the visual image representation of the input image; the pre-trained text encoder is used to extract the semantic text representation of the input image; the feature remapping network is used to learn the enhanced image representation; the classification head classifies the image based on the enhanced image representation.

[0068] S2. Obtain multi-batch long-tail distribution training data; In the long-tail category incremental learning task, the model needs to sequentially learn a series of new data, where each batch of new data consists of images of different categories and presents an unbalanced long-tail distribution (for example, contains a large number of images of the cat category while very few images of the dog category); Each batch of long-tail distribution training data will contain a fixed number of new category images (including few-shot categories and multi-shot categories) and their corresponding category labels, such as "cat" and "dog".

[0069] S3. Design a fixed language template and use a large language model to generate corresponding category text descriptions according to the category labels of the new category images in each batch of long-tail distribution training data.

[0070] Before training, it is necessary to generate a fixed number of category text descriptions according to the category labels of the new category images. Specifically:

[0071] Use the publicly available CuPL algorithm on the Internet (https: / / github.com / sarahpratt / CuPL) to generate. The generation principle is: design a fixed language template and use a large language model (such as ChatGPT3) to obtain the category text description. For example, for the category with the label "cat", a text description generated by the large language model may be: "A cat is an animal with soft fur, bright eyes, and a flexible tail."

[0072] S4. Input the multi-batch long-tail distribution training data and the corresponding category text descriptions into the long-tail incremental image recognition model batch by batch, and use a two-stage training method for training to obtain a trained long-tail incremental image recognition model.

[0073] This application uses a two-stage training method to enable the model to learn each batch of long-tail distribution training data, including the image-text feature alignment stage and the balanced training stage; Among them, the image-text feature alignment stage generates semantic text representations of new category images and performs feature alignment with the visual image representations of the images to obtain image-text features; The balanced training stage constructs a pseudo-feature space distribution and enhances the image-text features.

[0074] Furthermore, for the image-text feature alignment stage, its training steps are:

[0075] S4.1.1. Input the category text descriptions corresponding to the new category images in each batch of long-tail distribution training data into the pre-trained text encoder (such as CLIP) batch by batch to generate multiple semantic text representations corresponding to the new category images.

[0076] S4.1.2. Input the new category images in the same batch of long-tail distribution training data into the image feature backbone network to extract the visual image representations corresponding to the new category images.

[0077] S4.1.3. Calculate the average value on each dimension of multiple semantic text representations of the new category image to obtain the semantic text representation center of the new category image. The semantic text representation center is the average value of multiple semantic text representations on the feature dimension, expressed as:

[0078]

[0079] where T c is the semantic text representation center of the category label c, is the i-th semantic text representation of the category label c, and n is the number of semantic text representations of the category label c.

[0080] S4.1.4. Use the contrastive learning method to align the visual image representation of the new category image and the semantic text representation center of the new category image category label to obtain the image-text feature.

[0081] Specifically, the principle of contrastive learning is that for the images and text representations belonging to the same category in the new data, the distance between the two is as close as possible, while the distance between the images and text representations of different categories is as far as possible, that is, it is hoped that the image can be aligned with the corresponding text to prepare for the subsequent balanced training stage. The feature alignment steps are as follows:

[0082] Calculate the feature cosine similarity between the visual image representation of the new category image and the semantic text representation center of the new category image as the text confidence of the new category image;

[0083] Align the text confidence of the new category and the category label vector through the cross-entropy loss function to obtain the image-text feature:

[0084]

[0085] where L align is the image-text feature, C is the number of new category images in each batch of long-tailed distribution training data, y i is the category label vector of the i-th new category image, and s i is the text confidence of the i-th new category image.

[0086] Furthermore, for the balanced training stage, its training steps are as follows:

[0087] S4.2.1. After the training in the previous stage is completed, freeze the parameters of the image feature backbone network; then extract the visual image representation of the new category image in each batch of long-tailed distribution training data, and use the same method as obtaining the semantic text representation center to obtain the image representation center of the new category image through statistical calculation.

[0088] S4.2.2. Meanwhile, calculate the feature covariance matrix of the new-class images based on the visual image representation and the image representation center of the new-class images.

[0089] Specifically, the process of calculating the feature covariance matrix is as follows:

[0090] First, obtain the image representation center of the new-class images:

[0091]

[0092] where F c is the image representation center of the new-class images with class label c, k is the number of visual image representations of the new-class images, is the i-th visual image representation of the new-class images with class label c;

[0093] Then, calculate the feature covariance matrix of the new-class images:

[0094]

[0095] where K c is the feature covariance matrix of the new-class images with class label c, and T is the matrix transpose.

[0096] S4.2.3. Construct a pseudo-feature space distribution using the image representation center and the corresponding feature covariance matrix of the new-class images. Specifically, the pseudo-feature normal distribution satisfies the multivariate normal distribution, expressed as

[0097] Due to the loss of the old task, by storing the old distribution information with a very small cache cost, the image representations of the pseudo-old tasks can be sampled from the old task distribution space, thus alleviating the catastrophic forgetting problem.

[0098] By sampling from the pseudo-feature space distribution, the image representations of the current data and historical data can be obtained. Combining with the text representations containing new and old task categories obtained previously, a cross-image-text modality feature mixing and enhancement strategy is designed by referring to the classic Mixup feature enhancement method. Specifically:

[0099] S4.2.4. Sample an image pseudo-feature sample S i from the pseudo-feature space distribution of the new-class images, and at the same time sample a text feature sample S t from the semantic text representation of the new-class images, and randomly assign a modality weight, and then perform linear weighted interpolation to obtain the mixed modality feature.

[0100] Specifically, the process of obtaining the mixed modality feature by performing linear weighted interpolation is described as:

[0101] F mix = αSt +(1 - α)S i ,

[0102] Among them, F mix is the mixed - modality feature, α is the modality weight, and S t is the text feature sample sampled from the semantic text representation of the new - class images, and S i is the image pseudo - feature sample sampled from the pseudo - feature space distribution of the new - class images.

[0103] S4.2.5. Then, the mixed - modality feature is passed through the feature remapping network and the image classification head to obtain the corresponding classification output p mix .

[0104] S4.2.6. Use the cross - entropy loss of the weighted modality to supervise and balance the learning in the training stage, transfer the rich text semantic information to the few - sample image classes in the long - tail data, thereby enhancing the few - sample class representation and alleviating the over - fitting problem of the imbalanced new classes. In addition, due to the pseudo - image representation space distribution and text description of the old tasks, this strategy can also further alleviate catastrophic forgetting. Thus, the present invention can use the prior knowledge of the language modality to enhance the image representation, while solving the problems of over - fitting of the few - sample classes in the new data and catastrophic forgetting of the old - task knowledge encountered in the incremental learning of long - tail classes.

[0105] Specifically, the calculation method of the cross - entropy loss of the weighted modality is as follows:

[0106] First, obtain the mixed - modality label of the mixed - modality feature:

[0107] y mix =αy t +(1 - α)y i ,

[0108] Among them, y mix is the mixed - modality label vector of the mixed - modality feature, α is the modality weight, y t is the one - hot encoded vector of the class labels of the sampled text feature samples, and y i is the one - hot encoded vector of the class labels of the sampled image pseudo - feature samples;

[0109] Then, based on the number of new - class images in each batch of long - tail distribution training data, calculate the long - tail distribution weight of the class labels:

[0110]

[0111] Among them, w c is the long - tail distribution weight of the class label c, C is the number of class labels of new - class images in each batch of long - tail distribution training data, k cis the number of new category images for category label c, and γ is a hyperparameter;

[0112] Finally, calculate the cross-entropy loss of the weighted modality:

[0113]

[0114] where L mix is the cross-entropy loss of the weighted modality, w i is the long-tail distribution weight for category label i, y mix,i is the mixed-modality label for category label i, p mix,i is the output value of the classification head for the mixed-modality feature category label i.

[0115] S5. Input the long-tail distribution data to be recognized into the trained long-tail incremental image recognition model to obtain the recognition result.

[0116] In summary, the present invention introduces multi-modal language knowledge and solves the problems of overfitting of few-shot categories in new data and catastrophic forgetting of old task knowledge encountered in long-tail category incremental learning at the same time. It can be applied to various image classification task datasets (such as CIFAR100 and ImageNet, etc.) to alleviate overfitting and forgetting problems, and the performance is better than existing methods. Next, the present application will be described through a comparative experiment.

[0117] The comparative experiment of the present application is carried out on two widely used image classification datasets: CIFAR-100 and ImageNet-Subset, both of which contain 100 image categories. In addition, for the setting of the long-tail incremental scenario, the exponential decay factor ρ is used to represent the data imbalance degree of each batch of long-tail datasets, and ρ is set to 0.01, 0.05, and 0.1 respectively. The smaller ρ is, the more serious the long-tail distribution of the data is. In addition, the comparative experiment follows two commonly used long-tail incremental learning scenario settings: B50-T5 and B50-T10, that is, first learn a long-tail classification task containing 50 categories, and then in the incremental task, gradually learn the remaining long-tail categories (each incremental task learns 10 or 5 categories respectively). Finally, the test protocol widely used in the incremental learning method is adopted, that is, after each training task, test the current task and the categories seen in all previous tasks. The evaluation index is the average classification accuracy of all incremental tasks.

[0118] Specifically, the method proposed in this application (Ours) is compared with several popular incremental learning methods, including iCaRL [1], EEIL [2], UCIR [3], PODNet [4], PASS [5], and the popular two-stage decoupled learning method [6]. The popular two-stage decoupled learning method is applied to the EEIL, UCIR, and PODNet models respectively and is labeled as EEIL* [6], UCIR* [6], and PODNet* [6] in the experiment. In addition, some long-tail incremental learning methods are also compared, including GrRe [7] and ISpC [8]. The comparison results are shown in Table 1 below:

[0119] Table 1 Comparison Experiment Results

[0120]

[0121] As can be seen from Table 1, in contrast, the method proposed in this application shows significant improvement in all datasets and all different long-tail incremental scenarios, indicating that it effectively retains prior knowledge and can alleviate the catastrophic forgetting problem to a certain extent.

[0122] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0123] Based on the same idea as the language-guided long-tail incremental image recognition method in the above embodiments, the present invention also provides a language-guided long-tail incremental image recognition system, which can be used to execute the above language-guided long-tail incremental image recognition method. For the sake of convenience of description, in the structural schematic diagram of the language-guided long-tail incremental image recognition system embodiment, only the part related to the embodiment of the present invention is shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.

[0124] Please refer to Figure 2 , in another embodiment of this application, a language-guided long-tail incremental image recognition system is provided, which includes a model construction module, a data acquisition module, a text description module, a model training module, and an incremental recognition module;

[0125] Among them, the model construction module is used to construct a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network, and a classification head; the image feature backbone network is used to extract the visual image representation of the input image; the pre-trained text encoder is used to extract the semantic text representation of the input image; the feature remapping network is used to learn an enhanced image representation; the classification head classifies the image based on the enhanced image representation;

[0126] The data acquisition module is used to acquire multi-batch long-tail distribution training data; each batch of long-tail distribution training data contains a fixed number of new category images and their corresponding category labels;

[0127] The text description module is used to design a fixed language template and generate corresponding category text descriptions for the new category images in each batch of long-tail distribution training data through a large language model according to the category labels;

[0128] The model training module is used to input the multi-batch long-tail distribution training data and the corresponding category text descriptions into the long-tail incremental image recognition model batch by batch, and train it using a two-stage training method to obtain a trained long-tail incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; in the image-text feature alignment stage, the semantic text representation of the new category image is generated and feature-aligned with the visual image representation of the image to obtain image-text features; in the balanced training stage, a pseudo-feature space distribution is constructed and the image-text features are enhanced;

[0129] The incremental recognition module is used to input the long-tail distribution data to be recognized into the trained long-tail incremental image recognition model to obtain the recognition result.

[0130] It should be noted that the language-guided long-tail incremental image recognition system of the present invention corresponds one-to-one with the language-guided long-tail incremental image recognition method of the present invention. The technical features and their beneficial effects described in the embodiments of the above language-guided long-tail incremental image recognition method are applicable to the embodiments of the language-guided long-tail incremental image recognition. For specific content, please refer to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.

[0131] In addition, in the implementation manner of the language-guided long-tail incremental image recognition system in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the language-guided long-tail incremental image recognition system is divided into different program modules to complete all or part of the functions described above.

[0132] Please refer to Figure 3, in another embodiment, a computer-readable storage medium is provided, storing a program in a memory. When the program is executed by a processor, the above-mentioned long-tail incremental image recognition method based on language guidance is implemented, specifically as follows:

[0133] Construct a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network, and a classification head; the image feature backbone network is used to extract the visual image representation of the input image; the pre-trained text encoder is used to extract the semantic text representation of the input image; the feature remapping network is used to learn the enhanced image representation; the classification head classifies the image based on the enhanced image representation;

[0134] Obtain multi-batch long-tail distribution training data; each batch of long-tail distribution training data contains a fixed number of new-class images and their corresponding class labels;

[0135] Design a fixed language template and use a large language model to generate corresponding class text descriptions according to the class labels of the new-class images in each batch of long-tail distribution training data;

[0136] Input the multi-batch long-tail distribution training data and the corresponding class text descriptions into the long-tail incremental image recognition model batch by batch, and train it using a two-stage training method to obtain a trained long-tail incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; in the image-text feature alignment stage, the semantic text representation of the new-class image is generated and feature-aligned with the visual image representation of the image to obtain image-text features; in the balanced training stage, a pseudo-feature space distribution is constructed and the image-text features are enhanced;

[0137] Input the long-tail distribution data to be recognized into the trained long-tail incremental image recognition model to obtain the recognition result.

[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0139] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0140] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A language-guided long-tail incremental image recognition method, characterized in that: The steps include: A long-tail incremental image recognition model is constructed, comprising an image feature backbone network, a pre-trained text encoder, a feature remapping network and a classification head; the image feature backbone network is used to extract a visual image representation of an input image; the pre-trained text encoder is used to extract a semantic text representation of an input image; the feature remapping network is used to learn an enhanced image representation; the classification head classifies the image based on the enhanced image representation; Obtain multiple batches of long-tail distribution training data; each batch of long-tail distribution training data contains a fixed number of new category images and their corresponding category labels; Design a fixed language template and use a large language model to generate corresponding category text descriptions based on the category labels of new category images in each batch of long-tail distribution training data; Inputting multiple batches of long-tail distribution training data and corresponding category text descriptions into the long-tail incremental image recognition model in batches, and training them in a two-stage training method to obtain a trained long-tail incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; the image-text feature alignment stage generates a semantic text representation of a new category image and performs feature alignment with the visual image representation of the image to obtain an image-text feature; The balanced training phase constructs pseudo feature space distribution and enhances image-text features; The long-tail distribution data to be identified is input into the trained long-tail incremental image recognition model to obtain the recognition result.

2. According to claim 1, the language-guided long-tail incremental image recognition method is characterized in that: The training steps of the image-text feature alignment stage are: Inputting the category text descriptions corresponding to the new category images in each batch of long-tail distribution training data into the pre-trained text encoder in batches to generate multiple semantic text representations corresponding to the new category images; Input the new category images in the same batch of long-tail distribution training data into the image feature backbone network to extract the visual image representation corresponding to the new category images; Calculate the average value on each dimension of multiple semantic text representations of the new category image to obtain the semantic text representation center of the new category image; The contrastive learning method is used to align the visual image representation of the new category image and the semantic text representation center of the new category image to obtain the image-text feature.

3. The language-guided long-tail incremental image recognition method according to claim 2, characterized in that: The contrastive learning method is used to align the features of the visual image representation of the new category image and the semantic text representation center of the new category image, specifically: Calculate the feature cosine similarity between the visual image representation of the new category image and the center of the semantic text representation of the new category image as the text confidence of the new category image; The text confidence and category label vector of the new category are aligned through the cross entropy loss function to obtain the image-text feature.

4. The language-guided long-tail incremental image recognition method according to claim 1, characterized in that: The training steps of the balance training phase are: Freeze the parameters of the image feature backbone network, then extract the visual image representation of the new category images in each batch of long-tail distribution training data, and calculate the image representation center of the new category images; Calculate the feature covariance matrix of the new category image based on the visual image representation and image representation center of the new category image; The pseudo feature space distribution is constructed using the image representation center of the new category image and the corresponding feature covariance matrix; Sample an image pseudo-feature sample from the pseudo-feature space distribution of the new category image, and sample a text feature sample from the semantic text representation of the new category image, randomly assign a modality weight, and then perform linear weighted interpolation to obtain a mixed modality feature; The mixed modal features are passed through the feature remapping network and the image classification head to obtain the corresponding classification output; The learning of the balanced training phase is supervised using the cross entropy loss with weighted modalities.

5. The language-guided long-tail incremental image recognition method according to claim 4, characterized in that: The feature covariance matrix of the new category image is calculated based on the visual image representation of the new category image, specifically: Get the center of the image representation for images of a new category: Among them, F c is the image representation center of the new category image with category label c, k is the number of visual image representations of the new category image, is the i-th visual image representation of a new category image with category label c; Calculate the feature covariance matrix of the new category image: Among them, K c is the feature covariance matrix of the new category image with category label c, and T is the matrix transpose.

6. The language-guided long-tail incremental image recognition method according to claim 4, characterized in that: The constructed pseudo characteristic normal distribution satisfies the multivariate normal distribution.

7. According to the language-guided long-tail incremental image recognition method according to claim 4, it is characterized in that: The hybrid modal feature obtained by performing the linear weighted difference is expressed as: F mix =αS t +(1-α)S i , Among them, F mix is the mixed modal feature, α is the modal weight, S t is a text feature sample sampled from the semantic text representation of the new category image, S i is the image pseudo feature sample sampled from the pseudo feature space distribution of the new category image.

8. According to the language-guided long-tail incremental image recognition method according to claim 4, it is characterized in that: The cross entropy loss of the weighted modality is calculated as: Get the mixed-modal labels for mixed-modal features: y mix ay t +(1-α)y i , Among them, y mix is the mixed modal label vector of the mixed modal feature, α is the modal weight, y t is the one-hot encoding vector of the category label of the sampled text feature sample, y i is the one-hot encoding vector of the category label of the sampled image pseudo feature sample; Based on the number of new category images in each batch of long-tail distribution training data, the long-tail distribution weight of the category label is calculated: Among them, w c is the long-tail distribution weight of category label c, C is the number of category labels of new category images in each batch of long-tail distribution training data, k c is the number of new category images with category label c, and γ is a hyperparameter; Calculate the cross entropy loss for the weighted modality: Among them, L mix is the cross entropy loss of the weighted modality, w i is the long-tail distribution weight of category label i, y mix,i is the mixed-mode label of category label i, p mix,i is the classification head output value of the mixed modal feature category label i.

9. A language-guided long-tail incremental image recognition system, characterized in that: The language-guided long-tail incremental image recognition method applied to any one of claims 1 to 8 comprises a model building module, a data acquisition module, a text description module, a model training module and an incremental recognition module; The model building module is used to build a long-tail incremental image recognition model, including an image feature backbone network, a pre-trained text encoder, a feature remapping network and a classification head; the image feature backbone network is used to extract a visual image representation of an input image; the pre-trained text encoder is used to extract a semantic text representation of an input image; the feature remapping network is used to learn an enhanced image representation; the classification head classifies the image based on the enhanced image representation; The data acquisition module is used to acquire multiple batches of long-tail distribution training data; each batch of long-tail distribution training data contains a fixed number of new category images and their corresponding category labels; The text description module is used to design a fixed language template and generate corresponding category text descriptions according to the category labels of new category images in each batch of long-tail distribution training data through a large language model; The model training module is used to input multiple batches of long-tail distribution training data and corresponding category text descriptions into the long-tail incremental image recognition model in batches, and adopt a two-stage training method to train, so as to obtain a trained long-tail incremental image recognition model; the two-stage training method includes an image-text feature alignment stage and a balanced training stage; the image-text feature alignment stage generates a semantic text representation of a new category image and performs feature alignment with the visual image representation of the image to obtain an image-text feature; The balanced training phase constructs pseudo feature space distribution and enhances image-text features; The incremental recognition module is used to input the long-tail distribution data to be recognized into the trained long-tail incremental image recognition model to obtain the recognition result.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the language-guided long-tail incremental image recognition method described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Incremental multi-language text recognition method and system based on shared knowledge mining

    CN120561293A