Weak supervision multi-label image classification method based on high-order semantic coding
By constructing low-order embedding representations and high-order embedding representations of label semantic correlation, combining Transformer decoder and adaptive Top-K strategy, the problem of insufficient utilization of label semantic correlation and local feature in multi-label image classification is solved, and the accuracy and performance of multi-label image classification is improved.
Patent Information
- Application Number
- CN202510432121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multi-label image classification method is difficult to effectively model the semantic correlation and local features of tags, resulting in limited performance, especially in convolutional neural networks, recursive neural networks and graph convolutional neural networks, there are problems such as restricted receptive field, high computational complexity and increased noise areas.
Low-order embedding representations of label semantic correlation are constructed through pre-trained language models, and high-order embedding representations are generated using Transformer decoder, combining class activation mapping and adaptive Top-K strategy, significant target areas are located and extracted, local feature expression is enhanced, and collaborative learning is performed.
Improve the accuracy of multi-label image classification, capture the complex correlation between label semantics and visual information, suppress noise areas, and achieve efficient multi-label classification performance.
Smart Images

Figure CN120339703A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a weakly supervised multi-label image classification method based on high-order semantic encoding, belonging to the technical field of multi-label image classification. Background Art
[0002] With the rapid development of deep learning, convolutional neural networks (CNNs) have achieved remarkable progress in multiple computer vision tasks. As an important task in current computer vision, multi-label image classification aims to assign multiple labels to each image rather than just a single label. This task has wide applications in many fields such as scene understanding, image retrieval, and medical diagnosis, which also makes multi-label image classification gradually become a research problem with both practical value and challenges. However, due to the high-dimensional nature and complex dependencies of labels, most existing methods are difficult to model the label relationships of high-order representations, and these methods ignore the effective expression of local features. At the same time, there are also the following limitations: the receptive field of convolutional neural networks is limited, making it difficult to capture global context; recurrent neural networks rely on a fixed label order and are prone to introducing biases; while graph convolutional neural networks perform excellently in capturing the correlations between labels, their high computational complexity requires high resources. Weakly supervised multi-label learning methods have received extensive attention because they do not require precise region annotations and only rely on image-level label information to locate label-related regions to enhance the expression of local features. However, this type of method also has some limitations: on the one hand, although it can well capture the target regions, it will also introduce a large number of irrelevant noise regions, resulting in increased computational complexity while the performance decreases instead; on the other hand, the target regions provided by these methods lack the effective utilization of label semantic information and fail to fully explore the global correlations contained in the labels, thus limiting the performance of the model.
[0003] Therefore, there is an urgent need for an effective method to make up for the limitations of existing methods, which can both model the high-order relationships of label semantic correlations in multi-label tasks and make full use of the local feature information of samples, thereby improving the performance of multi-label image classification. Summary of the Invention
[0004] The purpose of the present invention is to provide a weakly supervised multi-label image classification method based on high-order semantic encoding, aiming to solve the technical problem that existing methods fail to fully utilize deep semantic relationships in multi-label learning.
[0005] To achieve the above purpose, the technical solution of the present invention is: a weakly supervised multi-label image classification method based on high-order semantic encoding, and the specific steps are as follows:
[0006] Step1: Construct a low-order embedding representation of label semantic correlations for the label categories included in the data set through a pre-trained language model;
[0007] Step 2: The low-order embedding representation is interactively aligned with the global features extracted by the CNN backbone through the Transformer decoder to generate a high-order embedding representation of the label semantic correlation.
[0008] Step 3: Using the global features extracted by the CNN, a heatmap representation of the target region is generated through class activation mapping by means of weak supervised learning, and the heatmap is analyzed to locate the potential target regions in each image.
[0009] Step 4: Based on the located potential target regions, an adaptive Top-K extraction strategy is adopted to select the target regions with the highest saliency, while suppressing the noise regions, and the target regions are re-input into the CNN to obtain an enhanced expression of the local features.
[0010] Step 5: The enhanced expression of the local features is fused and decoded using the high-order embedding representation to achieve co-learning modeling between the label semantics and the target regions, thereby improving the multi-label image classification performance.
[0011] The specific steps of the said Step 1 are as follows:
[0012] The label categories in the dataset are processed by a pre-trained language model, the labels are converted into embedding vectors, and context information is introduced to enhance the semantic relevance between the labels and capture the potential relationships between the labels. The formula is as follows:
[0013] L = Φ T (l)(1)
[0014] where Φ T (·) represents the pre-trained language model, l represents the label categories included in the dataset, and L ∈ R N×d is the low-order embedding representation of the generated label semantic correlation, where N represents the number of label categories and d represents the embedding dimension of the label semantics.
[0015] The specific steps of the said Step 2 are as follows:
[0016] After obtaining the low-order embedding representation through Step 1, a Transformer decoder is introduced, aiming to encode the high-order embedding representation based on the label semantic correlation to effectively capture the complex dependencies between the labels. The Transformer decoder consists of three layers, namely the multi-head self-attention layer, the multi-head cross-attention layer, and the feed-forward neural network layer. Among them, the multi-head self-attention layer and the multi-head cross-attention layer encode the semantic correlation between the low-order embedding representations through the attention distribution calculation. By calculating the dot product of the query matrix Q and the key matrix K, the model can capture the correlation between different positions, thereby obtaining the attention distribution. The attention layer is implemented as shown in the formula:
[0017]
[0018] Among them, Q, K, and V represent the query matrix, the key matrix, and the value matrix respectively, and d k represents the characteristic dimension of the key matrix K, and T represents the transpose;
[0019] For the input image I ∈ R 3×H×W , where H and W represent the height and width of the picture respectively, first the input image extracts the global visual feature representation through the CNN backbone where C represents the number of feature channels, and H1 and W1 represent the height and width of the global features. For the multi-head self-attention layer, where Q ∈ R N×d , K ∈ R N×d and V ∈ R N×d all come from the low-order embedding representation L of the label. For the multi-head cross-attention layer, Q1 ∈ R N×d comes from the self-attention output, and come from the visual feature representation F. Specifically, when processing the visual features, their shapes need to be transformed first to meet the requirements of the multi-head attention input. Finally, through the feed-forward neural network processing, the high-order embedding representation is obtained:
[0020] L' = RELU(xW1 + b1)W2 + b2 (3)
[0021] Among them, represents the output of the final high-order embedding representation, RELU represents the activation function, x represents the output of the cross-attention, W1 and W2 are weight matrices, and b1 and b2 are bias terms.
[0022] The high-order semantic encoding process mainly passes through the two multi-head attention mechanisms of the Transformer decoder, so that the context information of the label embedding can be more accurately aligned with the relevant visual features, thereby strengthening the correlation between the label and the visual information.
[0023] The specific steps of the said Step3 are:
[0024] For the global feature F extracted by the CNN, first perform a 1×1 convolution to map its feature dimension from the number of feature channels C to the number of categories N:
[0025]
[0026] Among them, represents the weight matrix, b3 is the bias term, represents the output containing the category features;
[0027] The output Y passes through the Sigmoid activation function, making the class feature scores within the range of [0, 1]. Then, Y is upsampled to the size of the original input image to obtain the class feature activation score block A ∈ R N×H×W , and the activation score distribution for each class is denoted as a n ∈ R H×W , where n ∈ [1,..., N]. The score of each class n at position (i, j) is a n (i, j). By performing global max pooling on the score distribution of each class along the spatial dimensions H and W, the feature score A of each class is obtained n :
[0028]
[0029] In this way, it can be determined that the class scores that actually exist in the image must be greater than the class scores that do not exist in the image. To select the most prominent class, that is, the class that actually exists, the mean μ and standard deviation σ of the scores of all classes are calculated, and the score threshold θ is obtained through the following formula:
[0030] θ = μ + kσ (6)
[0031] where k is an adjustable parameter used to determine the score threshold;
[0032] According to whether the feature score A n of each class exceeds the threshold, the local blocks containing the target regions in the class feature activation score block A are determined. The heatmap is generated using the local block activation score distribution, and the edge distributions of the heatmap are calculated along the x and y axes and normalized by dividing by the maximum score in their respective directions to obtain the target region. The formula is:
[0033]
[0034] where X d represents the edge distribution along the x direction, and Y d represents the edge distribution along the y direction;
[0035] According to the set fixed threshold δ ∈ (0, 1), the positioning of the target region is achieved and the corresponding coordinates are recorded for the next step of cropping.
[0036] The specific steps of the said Step4 are as follows:
[0037] After obtaining the target regions of several significant classes through the positioning in Step3, in order to extract the target regions, there is a key problem, that is, the number of labels in each picture is different, so the sizes of the obtained local blocks are also different. To solve the said problem, the present invention adopts an adaptive cropping strategy as described below;
[0038] To ensure that all the extracted regions are valid and there is no interference from irrelevant noise regions, each sample is pre-set to crop the top-K target regions, and there are two cases:
[0039] The first case: when the number of target regions is less than Top-K, determine significant regions according to θ, and represent the left and right horizontal coordinates and the lower and upper vertical coordinates of the k-th region respectively;
[0040] When k = 1, supplement by the equal-proportion cropping method:
[0041]
[0042] where t represents the number of target regions to be supplemented, α represents the proportionality factor, and represent the left and right horizontal coordinates and the lower and upper vertical coordinates of the t-th supplemented region;
[0043] When 1 < k < Top-K, supplement by using the permutation and combination of different regions:
[0044]
[0045] Start with the combination of two different region coordinates and increase step by step. If the combined region is still less than the set Top-K quantity, supplement it to Top-K by the equal-proportion cropping method;
[0046] The second case: when the number of target regions is equal to or greater than Top-K, crop according to the actual quantity of the Top-K regions;
[0047] Re-input all the obtained target regions into the CNN backbone to obtain an enhanced expression of the local features
[0048] The specific steps of Step5 are as follows:
[0049] For the obtained high-order semantic embedding representation L' ∈ R d×(H1×W1) decode the local feature F' ∈ R C×H1×W1 using the cross-attention mechanism to first fuse its features, and Q2 ∈ R C×H1×W1 comes from F', K2 ∈ R d×(H1×W1) and V2 ∈ R d ×(H1×W1) comes from L', so that the local feature and the high-order embedding representation achieve more effective feature alignment;
[0050] Concatenate the fused feature and the local feature along the channel dimension;
[0051] After passing through a multi-layer perceptron (MLP) and pooling operations, the predicted scores for each category are output.
[0052] The beneficial effects of the present invention are as follows: A weakly supervised multi-label classification method based on high-order semantic encoding proposed by the present invention aims to improve the accuracy of multi-label image classification tasks. First, the context information of label semantics is aligned with visual features through a Transformer decoder, thereby capturing high-order semantic representations to fully understand the complex correlation between label semantics and visual information. Secondly, the class activation mapping distribution is used to find significant target regions, and an adaptive target region extraction strategy is designed, which can well capture the target regions in the image and suppress irrelevant noise regions. The final co-learning process can achieve deep interaction between high-order semantics and local features, effectively strengthening the dependence relationship between label semantics and target regions, and fully making up for the limitations of existing methods. Experimental results demonstrate the effectiveness of the method and provide a new solution idea for multi-label image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the framework of a weakly supervised multi-label image classification method based on high-order semantic encoding;
[0054] Figure 2 is the edge distribution of the heat map based on the visualization of the class activation mapping distribution;
[0055] Figure 3 is the flow chart of the TOP-K cropping rule;
[0056] Figure 4 is the final visualization result diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Embodiment 1: As Figure 1 shown, a weakly supervised multi-label classification method based on high-order semantic encoding includes the following steps:
[0059] Step1: The label categories included in the data set are passed through a pre-trained language model to construct a low-order embedding representation of label semantic relevance.
[0060] Specifically, for the problem of modeling label relationships in multi-label learning. Through previous research, it is known that modeling the correlation between labels is one of the key factors to improve the model performance. Although current graph neural networks perform excellently in capturing the correlation between labels, they usually need to calculate the co-occurrence matrix first to construct the topological graph structure of label correlation, which brings a large computational overhead. In contrast, the pre-trained language model BERT can be used as an efficient word vector space encoding method to transform labels into embedding vectors, which contain context information, so as to capture the semantic correlation between labels, and the computational overhead is much lower than that of graph neural networks. The formula is as follows:
[0061] L = Φ T (l)(1)
[0062] where Φ T (·) represents the pre-trained language model, l represents the label categories included in the dataset, and L ∈ R N×d is the low-order embedding representation of the generated label semantic correlation, where N represents the number of label categories and d represents the embedding dimension of label semantics.
[0063] Step2: Interact and align the low-order embedding representation with the global features extracted by the CNN backbone through the Transformer decoder to generate the high-order embedding representation of label semantic correlation.
[0064] Specifically, in order to encode the high-order embedding representation of label semantic correlation, after obtaining the low-order embedding representation through Step1, the present invention introduces a Transformer decoder to effectively capture the complex dependencies between labels. As shown in the high-order semantic encoding module in Figure 1 , its input mainly includes two parts, namely the low-order embedding obtained in the above steps and the global visual features obtained from the CNN backbone. The Transformer decoder contains three layers, namely the multi-head self-attention layer, the multi-head cross-attention layer, and the feed-forward neural network layer. Among them, the multi-head self-attention layer and the multi-head cross-attention layer calculate through the attention distribution to encode the semantic correlation between the low-order embedding representations. By calculating the dot product of the query matrix Q and the key matrix K, the model can capture the correlation between different positions, so as to obtain the attention distribution. The implementation of the attention layer is as shown in the formula:
[0065]
[0066] where Q, K, and V represent the query matrix, the key matrix, and the value matrix respectively, and d k represents the feature dimension of the key matrix K, and T represents the transpose;
[0067] For the input image I ∈ R 3×H×W, where H and W respectively represent the height and width of the image. First, the input image extracts the global visual feature representation through the CNN backbone. where C represents the number of feature channels, and H1 and W1 represent the height and width of the global features. For the multi-head self-attention layer, it aims to obtain the high-level semantic representation of the label embedding, where Q ∈ R N×d 、K ∈ R N×d and V ∈ R N×d all come from the low-order embedding representation L of the label. For the multi-head cross-attention layer, Q1 ∈ R N×d comes from the self-attention output, and comes from the visual feature representation F. Specifically, when processing the visual features, their shapes need to be transformed first to meet the requirements of the multi-head attention input. Finally, after being processed by the feed-forward neural network, the high-order embedding representation is obtained:
[0068] L' = RELU(xW1 + b1)W2 + b2 (3)
[0069] where, represents the output of the final high-order embedding representation, RELU represents the activation function, x represents the output of the cross-attention, W1 and W2 are weight matrices, and b1 and b2 are bias terms.
[0070] Furthermore, the high-order semantic encoding process mainly passes through two multi-head attention mechanisms of the Transformer decoder, enabling the context information of the label embedding to be more precisely aligned with the relevant visual features, thereby strengthening the correlation between the label and the visual information.
[0071] Step3: Using the global features extracted by the CNN, adopt the weakly supervised learning method to generate the heatmap representation of the target area through class activation mapping, and analyze the heatmap to locate the potential target areas in each image.
[0072] Specifically, as shown in the weakly supervised object localization module in Figure 1 , in this module, the present invention realizes the target area localization in a simple and effective way, only using the image-level label information. Different from other studies, the present invention abandons using the attention mechanism to focus on local features, but instead focuses more on using the image data of the target area to enhance the expression ability of local features.
[0073] Furthermore, for the global feature F extracted by the CNN, local target area localization is realized through the class activation mapping distribution. First, perform a 1×1 convolution to map its feature dimension from the number of feature channels C to the number of categories N:
[0074]
[0075] where, where \(W_3\) represents the weight matrix and \(b_3\) is the bias term, representing the output containing class features;
[0076] The output \(Y\) passes through the Sigmoid activation function, making the class feature scores within the range \([0, 1]\). Then, \(Y\) is upsampled to the size of the original input image to obtain the class feature activation score block \(A\in\mathbb{R}\) N×H×W , and the activation score distribution for each class is denoted as \(a\) n \(\in\mathbb{R}\) H×W , \(n\in[1,\cdots,N]\), where the score of each class \(n\) at position \((i, j)\) is \(a\) n \((i, j)\). By performing global max - pooling on the score distribution of each class along the spatial dimensions \(H\) and \(W\), the feature score \(A\) of each class is obtained n :
[0077]
[0078] In this way, it can be determined that the class score of the actually existing class in the image must be greater than the class score of the non - existing class in the image. To select the most prominent class, that is, the actually existing class, the mean \(\mu\) and standard deviation \(\sigma\) of the scores of all classes are calculated, and the score threshold \(\theta\) is obtained through the following formula:
[0079] \(\theta=\mu + k\sigma\quad(6)\)
[0080] where \(k\) is an adjustable parameter used to determine the score threshold;
[0081] According to whether the feature score \(A\) n of each class exceeds the threshold, the local blocks containing the target area in the class feature activation score block \(A\) are determined. Using the local block activation score distribution, a heatmap is generated. The edge distributions of the heatmap are calculated along the \(x\) and \(y\) axes and normalized by dividing by the maximum score in their respective directions to obtain the target area, as Figure 2 shown. The formula is:
[0082]
[0083] where \(X\) d represents the edge distribution along the \(x\) - direction, and \(Y\) d represents the edge distribution along the \(y\) - direction;
[0084] According to the set fixed threshold \(\delta\in(0, 1)\), the location of the target area is realized and the corresponding coordinates are recorded for the next step of cropping.
[0085] Step 4: Based on the located potential target regions, an adaptive Top-K extraction strategy is adopted to select the target regions with the highest saliency, while suppressing the noise regions, and the target regions are re-input into the CNN to obtain an enhanced expression of the local features.
[0086] Specifically, after obtaining the target regions of several significant categories through Step 3 localization, in order to extract the target regions, there is a key problem, that is, the number of labels in each picture is different, so the size of the obtained local blocks is also different. To solve the above problem, the present invention adopts an adaptive cropping strategy as described below;
[0087] To ensure that the extracted regions are all valid regions without interference from irrelevant noise regions, it is preset to crop the Top-K target regions for each sample, and there are two cases:
[0088] The first case: when the number of target regions is less than Top-K, determine significant regions according to θ, and respectively represent the left and right abscissas and the lower and upper ordinates of the kth region;
[0089] When k = 1, supplement by the equal-proportion cropping method:
[0090]
[0091] where t represents the number of target regions to be supplemented, α represents the proportionality factor, and represent the left and right abscissas and the lower and upper ordinates of the tth supplemented region;
[0092] When 1 < k < Top-K, supplement by the method of arranging and combining different regions:
[0093]
[0094] Start from the combination of two different region coordinates and increase step by step. If the combined region is still less than the set Top-K quantity, supplement it to Top-K by the equal-proportion cropping method; in addition, this strategy can not only supplement the target regions, but also construct the potential correlation between multiple targets.
[0095] The second case: when the number of target regions is equal to or greater than Top-K, crop according to the actual quantity of the Top-K regions;
[0096] Specifically, the above situation is as Figure 3 shown. Finally, all the obtained target regions are re-input into the CNN backbone to obtain an enhanced expression of the local features Since the features are extracted by the same backbone, its shape is the same as the global features.
[0097] Step 5: Use the enhanced representation of high-order embeddings to fuse and decode the enhanced expression of local features, realize the collaborative learning modeling between label semantics and target regions, so as to improve the performance of multi-label image classification.
[0098] Specifically, for the outputs of the two modules in the previous steps, use the high-order semantic representation generated by the encoding module to decode the local features to complete collaborative learning. Similarly, in this step, the cross-attention mechanism is first used to fuse its features, where comes from F', and comes from L'. Then, the fused features and local features are concatenated along the channel dimension, and finally, through simple multi-layer perceptron (MLP) and pooling operations, the prediction scores for each category are output, and it is specified that a score greater than 0.5 indicates the existence of that category.
[0099] Furthermore, to realize the end-to-end learning process of the entire process, for an input image I, there are N categories, and the true label is represented as y = [y1,..., y N , where y n ∈{0, 1}, n∈[1,..., N]. If a certain category exists, then y n = 1, otherwise it is 0. The final prediction is represented as The present invention adopts the binary cross-entropy loss function (BCELoss) commonly used in multi-label classification tasks, and introduces the class weight w to adjust the importance between different categories to effectively alleviate the influence brought by class imbalance:
[0100]
[0101] where w represents the ratio of the number of negative samples to the number of positive samples for each category. Through this deep collaborative interaction learning mechanism, not only the complex dependence relationship between labels is strengthened, but also the positioning accuracy and classification ability of the target region are gradually improved in the iterative optimization process, making the classification performance of multi-labels more robust.
[0102] Next, based on the specific implementation records, the effectiveness of the technical solution of the present invention will be illustrated by experiments.
[0103] Experimental data and parameter settings: The method model proposed in the present invention was trained using the public dataset Pascal VOC2007. This dataset is one of the most commonly used benchmark datasets in the field of multi-label image classification, containing 20 categories and a total of 9,963 images, of which 5,011 are used for training and 4,952 are used for testing, with an average of 1.4 labels per image. To ensure the fairness of the experiment, the present invention follows the standard evaluation process, trains the model on train-val, and uses test for performance evaluation.
[0104] The present invention uses ResNet-101 pre-trained on the ImageNet-1k large-scale dataset as the backbone network to extract features, and uses the most popular resolution of 448×448 as the input. The model was trained for 40 epochs using the Adam optimizer with a batch size of 16. The learning rate of the pre-trained model was set to 0.00001, and the learning rate of other modules was set to 0.0001. At the 10th and 20th epochs, the learning rate was adjusted to 0.1 times the current value respectively.
[0105] As shown in Table 1, the method proposed in the present invention is compared with the current state-of-the-art methods. It is worth noting that most of the current advanced methods are based on the famous Transformer or GCN, which further verifies the potential advantages of using label correlation modeling in multi-label classification tasks, and the present invention achieves the highest mAP in this task. Specifically, both SSGRL and ML-GCN are based on GCN. At the same resolution, the present invention is superior to them by 1.5% and 0.9% respectively. In addition, at a resolution of 576 used in SST, the present invention still maintains a lead of 0.4%. At the same time, among all methods, the AP index of the present invention achieves the highest score in 9 categories, and the best performance is achieved in categories with relatively high recognition difficulty such as "chair", "potted plant" and "sofa". In addition, for a more intuitive understanding, the present invention Figure 4 shows the heat map of the model output visualization, and uses the CV tool to draw the corresponding cropping rectangle.
[0106] Table 1 Comparison of ap and mAP (%) between the present invention and the state-of-the-art methods on the Pascal VOC2007 dataset
[0107]
[0108] Among them, the best scores are highlighted in bold.
[0109] The present invention proposes a new framework for a weakly supervised multi-label image classification method based on high-order semantic encoding. By leveraging the advantages of Transformer in capturing long-range dependencies and global context relationships, it is used as an encoding module to generate high-order label semantic correlation representations. At the same time, local feature enhancement expressions are obtained by referring to weakly supervised methods and fully integrated with the former to achieve deep interaction between the two. In this framework, the global high-order representation generated by the encoding module is abstracted as a "dictionary" containing complex semantic relationships between labels, where the representation of each label corresponds to an entry in the dictionary. In this way, the semantic relationships between labels are effectively enhanced during the encoding process. Subsequently, the local features are decoded using this high-order representation, enabling each local region to find the most relevant label representation for accurate multi-label classification. Experimental results prove that the present invention is effective and provides new ideas for the multi-label image classification task.
[0110] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A weakly supervised multi-label image classification method based on high-order semantic encoding, characterized in that, It includes the following steps: Step1: Through a pre-trained language model, construct a low-order embedding representation of the semantic relevance of labels for the label categories included in the dataset; Step2: Through a Transformer decoder, interact and align the low-order embedding representation with the global features extracted by the CNN backbone to generate a high-order embedding representation of the semantic relevance of labels; Step3: Using the global features extracted by the CNN, adopt a weakly supervised learning method to generate a heatmap representation of the target region through class activation mapping, and analyze the heatmap to locate the potential target regions in each image; Step4: Based on the located potential target regions, adopt an adaptive Top-K extraction strategy to select the target regions with the highest saliency, while suppressing the noise regions, and re-enter the target regions into the CNN to obtain an enhanced expression of the local features; Step5: Use the high-order embedding representation to fuse and decode the enhanced expression of the local features to achieve collaborative learning modeling between the label semantics and the target regions.
2. The weakly supervised multi-label image classification method based on high-order semantic encoding according to claim 1, wherein The specific steps of Step1 are as follows: Process the label categories in the dataset through a pre-trained language model, convert the labels into embedding vectors, and introduce context information to capture the potential relationships between the labels. The formula is as follows: L = Φ T (l) (1) Among them, Φ T (·) represents the pre-trained language model, l represents the label categories included in the dataset, L ∈ R N×d is the low-order embedding representation of the generated label semantic relevance, where N represents the number of label categories and d represents the embedding dimension of the label semantics.
3. The weakly supervised multi-label image classification method based on high-order semantic encoding according to claim 1, wherein The specific steps of Step2 are as follows: After obtaining the low-order embedding representation through Step1, introduce a Transformer decoder. The Transformer decoder contains three layers, namely the multi-head self-attention layer, the multi-head cross-attention layer, and the feed-forward neural network layer. Among them, the multi-head self-attention layer and the multi-head cross-attention layer calculate through the attention distribution to encode the semantic relevance between the low-order embedding representations. The attention layer is implemented as shown in the formula: Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and d k represents the characteristic dimension of the key matrix K, and T represents the transpose; For the input image I ∈ R 3×H×W , where H and W represent the height and width of the picture respectively. First, the input image extracts the global visual feature representation through the CNN backbone , where C represents the number of feature channels, and H1 and W1 represent the height and width of the global features. For the multi-head self-attention layer, where Q ∈ R N×d , K ∈ R N×d and V ∈ R N×d all come from the low-order embedding representation L of the label. For the multi-head cross-attention layer, Q1 ∈ R N×d comes from the self-attention output and comes from the visual feature representation F. After being processed by the feed-forward neural network, the high-order embedding representation is obtained: L' = RELU(xW1 + b1)W2 + b2 (3) Among them, represents the output of the final high-order embedding representation, RELU represents the activation function, x represents the output of cross-attention, W1 and W2 are weight matrices, and b1 and b2 are bias terms.
4. The weakly supervised multi-label image classification method based on high-order semantic encoding according to claim 1, wherein The specific steps of Step3 are as follows: For the global feature F extracted by the CNN, first perform a 1×1 convolution to map its feature dimension from the number of feature channels C to the number of categories N: Among them, represents the weight matrix, and b3 is the bias term. represents the output containing class features; The output Y passes through the Sigmoid activation function, making the class feature scores within the range of [0, 1]. Then Y is upsampled to the size of the original input image to obtain the class feature activation score block A ∈ R N×H×W , and the activation score distribution for each class is denoted as a n ∈ R H×W , n ∈ [1,..., N], where the score of each class n at position (i, j) is a n (i, j). By performing global max pooling on the score distribution of each class along the spatial dimensions H and W, the feature score A of each class is obtained n : Calculate the mean μ and standard deviation σ of the scores of all categories, and obtain the score threshold θ through the following formula: θ = μ + kσ (6) Among them, k is an adjustable parameter used to determine the score threshold; According to the feature score A of each category n Determine the local block containing the target area in the class feature activation score block A based on whether it exceeds the threshold. Generate a heat map using the local block activation score distribution. Calculate the edge distribution of the heat map in both the x and y axis directions and normalize it by dividing by the maximum score in their respective directions to obtain the target area. The formula is as follows: Among them, X d represents the edge distribution along the x-direction, and Y d represents the edge distribution along the y-direction; According to the set fixed threshold δ ∈ (0, 1), locate the target region and record the corresponding coordinates for the next step of cropping.
5. The weakly supervised multi-label image classification method based on high-order semantic encoding according to claim 1, wherein, The specific steps of Step4 are as follows: After obtaining the target regions of several significant categories located through Step3, preset to crop Top-K target regions for each sample. There are two cases: The first case: When the target area is smaller than Top-K, determine the number of significant areas according to θ, and respectively represent the left and right abscissas and the lower and upper ordinates of the k-th area; When k = 1, supplement by the equal-proportion cropping method: where t represents the number of target regions to be supplemented, and α represents the scaling factor, and represent the left and right abscissas and the lower and upper ordinates of the t-th supplementary region; When 1 < k < Top-K, supplement by using the combination method of different regions: Start from the combination of two different region coordinates and increase step by step. If the combined region is still less than the set Top-K quantity, supplement it by the equal-proportion cropping method until Top-K; The second case: When the number of target regions is equal to or greater than Top-K, crop according to the actual number of Top-K regions; Re-input all the obtained target regions into the CNN backbone to obtain an enhanced expression of local features 6. The weakly supervised multi-label image classification method based on high-order semantic encoding according to claim 1, wherein The specific steps of Step5 are as follows: For the obtained high-order semantic embedding representation For the local features Decode them, first fuse their features using the cross-attention mechanism, and from F', and from L'; Concatenate the fused features and the local features along the channel dimension; After passing through a multi-layer perceptron (MLP) and pooling operations, the predicted scores for each class are output.