Zero-sample multi-label identification method of double-branch task residual error based on parameter-free attention enhancement

Through the dual-branch task residual method with parameterless attention enhancement, the problem of multi-label image recognition identifying new categories under zero-sample conditions is solved, and more efficient multi-modal information interaction and prior knowledge retention is achieved, which improves the accuracy and flexibility of zero-sample multi-label recognition.

CN120356053APending Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510058320.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing multi-label image recognition method is difficult to effectively identify new categories under zero-sample settings, and the pre-trained model loses prior knowledge and computational overhead during the adaptation process, and lacks cross-modal information interaction.

Method used

The dual-branch task residual method with parameterless attention enhancement is adopted, and prior knowledge is retained from global and local perspectives through the dual-branch task residual module, and the multi-modal information interaction is enhanced by the parameterless attention module, and a network architecture suitable for zero-sample multi-label recognition is built.

Benefits of technology

It improves the accuracy and flexibility of multi-label recognition, especially in zero-sample and partial-label conditions, enhances the learning ability of specific task knowledge and cross-modal information interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356053A_ABST
    Figure CN120356053A_ABST
Patent Text Reader

Abstract

The invention discloses a zero-sample multi-label identification method for a double-branch task residual error based on parameter-free attention enhancement, and the method comprises the steps: firstly providing a double-branch task residual error considering global and local concepts, and greatly improving a final multi-label identification result through a prior irrelevant double-branch task residual error; secondly, in order to enable image / text features to be more concentrated on a target class, a parameter-free attention module is adopted to enhance interactivity of multi-modal information; the dual-branch task residual error is designed to reserve priori knowledge from global and local angles, so that flexibility, expandability and the ability of learning specific task knowledge are enhanced, and a parameter-free attention mechanism aims at providing a bridge for multi-mode information interaction and enhancing a highly interested region of a text in an image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a zero-shot multi-label recognition method based on a dual-branch task residual with parameter-free attention enhancement. Background Art

[0002] Multi-label image recognition (MLR) aims to identify all object categories or concepts present in an input image. Due to the inherent multi-label nature of images, developing specific algorithms to solve the multi-label image recognition problem has great potential value. It is beneficial for a comprehensive understanding of complex scenes and can assist other tasks such as image retrieval. However, in many cases, some categories are not easily collected for training the MLR model, which requires the model to be able to recognize new categories in the "zero-shot" setting.

[0003] Early work used methods such as graph convolutional networks (GCNs) to explore the relationships between labels. For example, SCPNet et al. used GCNs in multi-label classification to learn semantic graph embeddings; MlTr et al. combined pixel attention and cross-window attention to solve the problem of recognizing local small objects; FL-Tran et al. utilized a multi-scale fusion module to learn multi-scale features. However, all of these methods require a large number of annotated images for training, and learning multi-label image recognition remains a challenging problem in cases where images or labels are limited.

[0004] In recent years, with the emergence of large pre-trained vision-language models such as CLIP and ALIGN, which are trained through contrastive learning of billions of image-text pairs, a well-aligned image-text feature space can be obtained. Therefore, pre-trained vision-language models can be endowed with the powerful ability to recognize new classes. To effectively transfer large pre-trained models to downstream tasks, there are currently two popular and effective methods, namely prompt tuning and adapter tuning, which aim to tune the model with a small number of parameters. However, prompt tuning lacks the preservation of prior knowledge. Although the weights of the pre-trained text branch module are frozen in the prompt tuning paradigm, the original well-learned classification boundary is more or less damaged. It abandons the pre-trained text-based classifier and generates a new classifier, which leads to the loss of prior knowledge of VLMs. And it requires the pre-trained text encoder to participate in the whole training, which limits its scalability and increases the computational overhead. Additionally, adapter tuning may not fully explore task-specific knowledge because the input to the adapter is strictly restricted to the old pre-trained features. Regardless of whether the pre-trained features are suitable for the task, the results of the adapter only depend on them, which makes the adapter-style tuning have limited flexibility in learning new knowledge. In existing multi-label recognition methods based on pre-trained base models, text and images are independent of each other before matching, lacking cross-modal information interaction. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the present invention provides a zero-shot multi-label recognition method based on a parameter-free attention-enhanced dual-branch task residual. First, a dual-branch task residual considering global and local concepts is proposed. With the help of the prior-independent dual-branch task residual, the final multi-label recognition result can be greatly improved. Secondly, in order to make the image / text features more concentrated on the target classes, a parameter-free attention module is adopted to enhance the interaction of multi-modal information. The dual-branch task residual in the present invention is designed to retain prior knowledge from both global and local perspectives, thereby enhancing flexibility, scalability, and the ability to learn specific task knowledge. The parameter-free attention mechanism aims to provide a bridge for multi-modal information interaction and enhance the regions of high interest in the text of the image.

[0006] The technical solutions adopted by the present invention to solve its technical problems are as follows:

[0007] Step 1: Data preparation stage;

[0008] Step 1-1: Dataset preparation;

[0009] Use text descriptions instead of images for training:

[0010] Step 1-2: Construct a noun filter;

[0011] Use Python to implement a noun phrase filter based on NLTK; extract the key object nouns from the text description through the noun filter to obtain noun labels;

[0012] Step 1-3: Generate pseudo-labels;

[0013] Perform one-hot mapping on the obtained noun labels to generate pseudo-labels;

[0014] Step 2: Construct a multi-label recognition network architecture;

[0015] Step 2-1: Construct a large pre-trained base model CLIP, which consists of two parts: an image encoder Image Encoder and a text encoder Text Encoder;

[0016] The image encoder is responsible for converting the input image into a high-dimensional vector representation for comparison with the vector generated by the text encoder;

[0017] The text encoder is responsible for converting the input natural language text into a vector representation similar to the output of the image encoder;

[0018] The vectors generated by the two encoders are subjected to contrastive learning in the embedding space;

[0019] Step 2-2: Construct a dual-branch task residual learning module;

[0020] Use two sets of learnable parameters, one for local feature classification and one for global feature classification, and add these two sets of parameters to the learnable parameters in the form of residuals to form a dual-branch task residual learning module; the task residual is expressed as:

[0021] F′ text = F text + α·X

[0022] where X ∈ R K×D represents a set of learnable parameters, F text ∈ R K×D is used as a base classifier to represent the text embedding obtained from the text encoder, K represents the number of classes, D represents the dimension; α is a hyperparameter;

[0023] Scale the learnable parameters X using the hyperparameter α, and then merge α·X into the base classifier to create a new classifier for the target task, denoted as F′ text ;

[0024] Step 2-3: Construct a Parameter Free Attention module;

[0025] Given the image embedding as the query, use the text embedding processed by the dual-branch task residual module as the key and value, and calculate the score matrix; use the Softmax function to calculate the attention weights for the score matrix;

[0026] Weight the text embedding using the attention weights to obtain enhanced image features;

[0027] Step 2-4: Construct a global feature classification head;

[0028] The residual operation in the global branch is:

[0029]

[0030] where, X global represent the global base classifier, global new classifier, and global learnable parameters respectively;

[0031] Step 2-5: Construct a local feature classification head;

[0032] The residual operation in the local branch is:

[0033]

[0034] where X localrespectively represent the local base classifier, the local new classifier, and the local learnable parameters;

[0035] Step 3: Construct a loss function applicable to multi-label recognition;

[0036] The overall objective of parameter optimization is:

[0037] L = L global + L local

[0038]

[0039] where c + represents the positive label, c - represents the negative label, m represents the margin used to control the distance between the positive label and the negative label; p i and p j are the classification probabilities of the positive sample and the negative sample respectively, m is a boundary value used to control the minimum interval between the positive and negative samples; p′ i and p′ j represent the classification probabilities of the positive sample and the negative sample of the local branch respectively;

[0040] L global and L local are obtained by calculating the classification probabilities of the global branch and the local branch respectively, and the pseudo-label is generated by the noun filter;

[0041] Use the ranking loss to measure the difference between the classification probability and the pseudo-label;

[0042] Step 4: Network training;

[0043] Adopt text description for multi-label recognition training;

[0044] In the training stage, the text description is only used to optimize the prior-agnostic double-branch task residuals, while the other parts remain frozen.

[0045] Preferably, the dataset of the text description in the step 1-1 includes:

[0046] (1) Open Images Localized Narratives: Open Images V6 adds new visual relationship annotations, including a new form of local narrative annotation, that is, the image is attached with voice, text, and mouse trajectory annotation information; take its local narrative instead of its picture;

[0047] (2) MS-COCO Captions: The Microsoft Common Objects in Context (MS-COCO) dataset contains detailed annotations for images, one of which is the description of the image. The description is not just a simple label, but a natural language sentence to more accurately express the scene in the image; the total number of image descriptions in the dataset exceeds 600,000.

[0048] Preferably, the image encoder uses the ResNet-50 residual network as the backbone to process the image.

[0049] A computer program that causes a computer to execute the above zero-shot multi-label recognition method.

[0050] An electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the above zero-shot multi-label recognition method.

[0051] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above zero-shot multi-label recognition method is implemented.

[0052] A chip, comprising: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the above zero-shot multi-label recognition method.

[0053] A computer program product, the computer program product includes a computer storage medium, the computer storage medium stores a computer program, the computer program includes instructions that can be executed by at least one processor, and when the instructions are executed by the at least one processor, the above zero-shot multi-label recognition method is implemented.

[0054] The beneficial effects of the present invention are as follows:

[0055] 1) The dual-branch task residual module proposed by the present invention can transfer the pre-trained model for multi-label image recognition tasks from both the global and local aspects.

[0056] 2) In order to make the image / text features more concentrated on the target classes, the present invention adopts a parameter-free attention module to enhance the interaction of multi-modal information.

[0057] 3) The present invention has achieved excellent results on three zero-shot multi-label image recognition datasets, VOC2007, MS-COCO, and NUS-WIDE, exceeding the current state-of-the-art methods. In addition, the method proposed by the present invention also performs very well in the partial label image recognition task. Description of the Drawings

[0058] Figure 1 This is the overall framework diagram of the method of the present invention. Detailed implementation manners

[0059] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0060] The present invention proposes a Dual-branch Task Residual with Parameter-free Attention enhancement (DTRPA) framework. This framework consists of a Dual-branch Task Residual (DBTR) and a Parameter-free Attention (PFA) module, aiming to capture more precise globally complementary and locally complementary features and provide a bridge for multi-modal information interaction. Specifically, first, a dual-branch task residual considering global and local concepts is proposed. With the prior-independent dual-branch task residual, the final multi-label recognition result can be greatly improved. Second, in order to make the image / text features more concentrated on the target classes, we adopt a parameter-free attention module to enhance the interactivity of multi-modal information. The dual-branch task residual in the present invention is designed to retain prior knowledge from both global and local perspectives, thereby enhancing flexibility, scalability, and the ability to learn task-specific knowledge. The parameter-free attention mechanism aims to provide a bridge for multi-modal information interaction and enhance the regions of high interest in the text in the image.

[0061] Specifically, it includes the following steps:

[0062] Step 1: Data preparation stage;

[0063] Step 1-1: Dataset preparation;

[0064] This method uses text descriptions instead of images for training:

[0065] (1) Open Images Localized Narratives: Open Images V6 has added a large number of new visual relationship annotations, including a new annotation form of localized narratives, that is, annotation information such as speech, text, and mouse trajectories are attached to the images. This is a completely new multi-modal annotation form. These local narratives apply to its 500k images. We only take its local narratives and do not use its pictures.

[0066] (2) MS-COCO Captions: The Microsoft Common Objects in Context (MS-COCO) dataset contains detailed annotations of images, one of which is the description (Captions) of the images. The description is not just a simple label, but a relatively detailed natural language sentence to more accurately express the scene in the image. The total number of image descriptions in the dataset exceeds 600,000.

[0067] In the test and verification of the present invention, three widely used data sets, namely VOC2007, MS-COCO and NUS-WIDE, were used to evaluate the method. Among them: (1) VOC2007: This data set includes 20 common categories, which we divided into a training set containing 5,011 images and a test set containing 4,952 images according to the existing literature. (2) MS-COCO: This data set contains 80 categories, with 82,081 training images and 40,504 validation images, which conform to the official division. (3) NUS-WIDE: This data set includes 81 interrelated concepts, providing a comprehensive test set of 107,859 images for evaluating the method of the present invention.

[0068] So far, the data set preparation is completed.

[0069] Step 1-2: Construct a noun filter;

[0070] The present invention uses Python to implement a noun phrase filter based on NLTK. Through the noun filter, key object nouns are extracted from the text description to obtain noun tags.

[0071] Step 1-3: Generate pseudo-labels;

[0072] The obtained noun tags are one-hot mapped to generate pseudo-labels, which are used to assist in the training of the supervised model.

[0073] Step 2: Construct a multi-label recognition network architecture;

[0074] Step 2-1: Construct a large pre-trained basic model, including an Image encoder and a Text encoder;

[0075] The present invention constructs a large pre-trained basic model CLIP, whose function is to convert text into language information and better couple it with image information in UNet using Attention, and then connect text and image. The specific network structure of CLIP consists of two parts: an image Image encoder + a text Text encoder. Among them, the Image encoder is used to extract image features, and the Text encoder is used to extract text features.

[0076] Step 2-2: Construct a dual-branch task residual learning module;

[0077] The present invention uses two sets of learnable parameters, one for local feature classification and one for global feature classification. Adding these two sets of parameters to the learnable parameters in the form of residuals forms a dual-branch task residual learning module. This module is a set of parameters that can be continuously optimized independently of the base classifier (based on text embedding). The task residual can be expressed as:

[0078] F′ text =F text +α·X

[0079] where X ∈ R K×D represents a set of learnable parameters, F text ∈ R K×D serves as the base classifier representing the text embedding obtained from the text encoder, K represents the number of classes, and D represents the dimension. The hyperparameter α is used to scale X, and an adaptive learning method is adopted to learn the appropriate coefficient α. The present invention uses the hyperparameter α to scale the parameters X specifically learned for a given task. Subsequently, α·X is incorporated into the base classifier to create a new classifier for the target task, denoted as F′ text .

[0080] In multi-label recognition tasks, global features often prioritize the main objects and may ignore other important objects. To address this issue, a dual-branch task residual learning module is constructed, that is, a local branch is introduced to mitigate the dominance of global features. It enhances the exploration of fine-grained features and improves the sensitivity to image details. The collaboration between the global and local branch structures enables us to comprehensively capture image information. This structure can improve the accuracy of multi-label recognition, especially in multi-label recognition in complex scenarios.

[0081] Step 2-3: Construct a Parameter Free Attention module;

[0082] Given the image embedding as the query, and using the text embedding processed by the dual-branch task residual module as the key and value, calculate the score matrix. The Softmax function is used to calculate the score matrix to obtain the attention weights. The attention weights are used to weight the text embedding to enhance the regions of interest in the text in the image, obtaining the enhanced image features.

[0083] Step 2-4: Construct a global feature classification head;

[0084] The residual operation in the global branch is:

[0085]

[0086] where X globalrespectively represent the global base classifier, the global new classifier, and the global learnable parameters.

[0087] Step 2 - 5: Construct the local feature classification head;

[0088] In the local branch, the residual operation is:

[0089]

[0090] where X local respectively represent the local base classifier, the local new classifier, and the local learnable parameters.

[0091] Step 3: Construct a loss function suitable for multi - label recognition;

[0092] The overall objective of parameter optimization is:

[0093] L = L global +L local

[0094] where L global and L local are obtained by calculating the classification probabilities of the global branch and the local branch respectively, and the pseudo - labels are generated by the noun filter. The present invention uses the ranking loss to measure the difference between the prediction probability and the pseudo - label, rather than using the most widely used multi - label classification loss function binary cross - entropy loss (BCE).

[0095] L global and L local can be expressed as follows:

[0096]

[0097] where c + represents the positive label, c - represents the negative label, and m represents the margin used to control the distance (cosine similarity) between the positive label and the negative label.

[0098] Step 4: Network training;

[0099] Adopt text descriptions for multi - label recognition training;

[0100] In the training stage, the text description is only used to optimize the prior - irrelevant double - branch task residuals, while the other parts remain frozen. This method enables us to reliably retain prior knowledge and facilitates the flexible exploration of new knowledge.

[0101] Step 5: Network testing;

[0102] Step 5 - 1: Replace the Text encoder in the training stage with the Image encoder;

[0103] In the testing phase, the text encoder of the present invention is replaced with an image encoder. Since the input in the testing phase is an Image and an Image encoder is required, the Text encoder in the training phase needs to be replaced with an Image encoder.

[0104] Step 5-2: Extract features from the input image;

[0105] The present invention extracts local and global features from the input test image respectively.

[0106] Step 5-3: Conduct multi-label recognition testing;

[0107] The present invention conducts testing on multi-label recognition tasks.

Claims

1. A zero-shot multi-label recognition method based on dual-branch task residuals with parameter-free attention enhancement, characterized in that, It includes the following steps: Step 1: Data preparation stage; Step 1-1: Dataset preparation; Use text descriptions instead of images for training: Step 1-2: Construct a noun filter; Implement a noun phrase filter based on NLTK using Python; extract the key object nouns from the text description through the noun filter to obtain noun labels; Step 1-3: Generate pseudo-labels; Perform one-hot mapping on the obtained noun labels to generate pseudo-labels; Step 2: Construct a multi-label recognition network architecture; Step 2-1: Construct a large pre-trained base model CLIP, which consists of two parts: an image encoder Image Encoder and a text encoder Text Encoder; The image encoder is responsible for converting the input image into a high-dimensional vector representation for comparison with the vector generated by the text encoder; The text encoder is responsible for converting the input natural language text into a vector representation similar to the output of the image encoder; The vectors generated by the two encoders are subjected to contrastive learning in the embedding space; Step 2-2: Construct a dual-branch task residual learning module; Use two sets of learnable parameters, one for local feature classification and one for global feature classification, and add these two sets of parameters to the learnable parameters in the form of residuals to form a dual-branch task residual learning module; the task residual is expressed as: F' text = F text + α·X where X ∈ R K×D represents a set of learnable parameters, F text ∈ R K×D as the base classifier represents the text embedding obtained from the text encoder, K represents the number of classes, D represents the dimension; α is a hyperparameter; Scale the learnable parameter X using the hyperparameter α, and then incorporate α·X into the base classifier to create a new classifier for the target task, denoted as F'. text ; Step 2-3: Construct a Parameter Free Attention module; Given the image embedding as a query, use the text embedding processed by the dual-branch task residual module as keys and values to calculate a score matrix; use the Softmax function to calculate the score matrix to obtain attention weights; Weight the text embedding with the attention weights to obtain enhanced image features; Step 2-4: Construct a global feature classification head; The residual operation in the global branch is: Among them, X global respectively represent the global base classifier, the global new classifier, and the global learnable parameters; Step 2-5: Construct a local feature classification head; The residual operation in the local branch is: Among them X local respectively represent the local base classifier, the local new classifier, and the local learnable parameters; Step 3: Construct a loss function suitable for multi-label recognition; The overall goal of parameter optimization is: L = L global + L local where c + represents the positive label, and c - represents the negative label, m represents the margin for controlling the distance between the positive and negative labels; p i and p j are the classification probabilities of the positive and negative samples respectively, and m is a boundary value for controlling the minimum interval between the positive and negative samples; p′ i and p′ j represent the classification probabilities of the positive and negative samples of the local branch respectively; L global and L local are obtained by calculating the classification probabilities of the global branch and the local branch respectively, and the pseudo-label is generated by a noun filter; Use ranking loss to measure the difference between the classification probability and the pseudo-labels; Step 4: Network training; Perform multi-label recognition training using text descriptions; In the training stage, the text description is only used to optimize the dual-branch task residuals independent of the prior, while the other parts remain frozen.

2. The zero-shot multi-label recognition method based on a dual-branch task residual with parameter-free attention enhancement according to claim 1, wherein, The dataset of text descriptions in Step 1-1 includes: (1) Open Images Localized Narratives: Open Images V6 adds new visual relationship annotations, including a new annotation form of local narratives, that is, voice, text, and mouse trajectory annotation information is attached to the image; take its local narrative instead of its picture; (2) MS-COCO Captions: The Microsoft Common Objects in Context (MS-COCO) dataset contains detailed annotations for images, one of which is the description of the image. The description is not just a simple label, but a natural language sentence to more accurately express the scene in the image; the total number of image descriptions in the dataset exceeds 600,000.

3. A zero-shot multi-label recognition method based on a dual-branch task residual with parameter-free attention enhancement according to claim 1, characterized in that The image encoder uses the Residual Network ResNet-50 as the backbone to process the image.

4. A computer program, characterized in that, The computer program causes the computer to execute the method according to any one of claims 1 to 3.

5. An electronic device, characterized in that, Comprising: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 3.

7. A chip, characterized in that, Comprising: A processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method according to any one of claims 1 to 3.

8. A computer program product, characterized in that, The computer program product includes a computer storage medium, the computer storage medium stores a computer program, and the computer program includes instructions that can be executed by at least one processor. When the instructions are executed by the at least one processor, the method according to any one of claims 1 to 3 is implemented.