Multi-label image classification method based on graph semantic interaction

By introducing a multi-scale retention mechanism and a cross-modal attention mechanism in multi-label image classification, the problems of insufficient fusion of semantic features and visual features and unreasonable construction of label relationship maps in the existing technology are solved, and the classification accuracy and feature expression capabilities are significantly improved.

CN120147696APending Publication Date: 2025-06-13XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510172949.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-13

Smart Images

  • Figure CN120147696A_ABST
    Figure CN120147696A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-label image classification method based on graph semantic interaction, and solves the problems that the classification performance is poor and the complex relationship between labels cannot be fully utilized due to the lack of dynamic adaptive capacity in the prior art. The method comprises the following steps: acquiring a high-dimensional visual feature map corresponding to an image to be classified, a description text corresponding to each label and a label embedding vector; calculating an initial similarity matrix among the labels, introducing a multi-scale retention mechanism model to generate a label relation graph, and further mapping to obtain high-dimensional label semantic features; performing feature interaction on the high-dimensional visual feature map and the high-dimensional tag semantic features, generating a semantic channel attention vector through a cross-modal attention mechanism, and dynamically adjusting image features according to semantic channel attention to obtain output image features of a current layer; and inputting the information into a classification layer to obtain the classification probability of each label, and completing a multi-label classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a multi-label image classification method based on graph semantic interaction. Background Art

[0002] Multi-label image classification is an important task in the field of computer vision, aiming to identify multiple object labels in an image. With the development of deep learning technology, methods based on graph convolutional networks have gradually become the mainstream methods for multi-label image classification. These methods capture the dependency relationships between labels by constructing label relationship graphs, thereby improving the classification performance.

[0003] Existing multi-label classification methods are mainly divided into two categories: methods based on region of interest localization and methods based on label relationship modeling. The former improves the classification accuracy by localizing specific regions in the image, but may ignore the overall context information of the image; the latter improves the classification performance by modeling the dependency relationships between labels. In particular, methods based on graph convolutional networks can effectively capture the complex relationships between labels and have gradually become the mainstream. However, the existing methods based on graph convolutional networks still have the following problems: insufficient fusion of semantic features and visual features: existing methods usually map the semantic features to the classifier through simple dot product operations, lacking dynamic adaptability, resulting in poor classification performance; unreasonable construction of the label relationship graph: existing methods fail to fully utilize the complex relationships between labels when constructing the label relationship graph. Summary of the Invention

[0004] By providing a multi-label image classification method based on graph semantic interaction, the present invention solves the problems in the prior art of lacking dynamic adaptability, resulting in poor classification performance, and not being able to fully utilize the complex relationships between labels, and realizes accurately capturing the complex dependency relationships between labels, dynamically embedding label semantics into the visual feature extraction process, and enhancing the saliency of category-related features.

[0005] The present invention provides a multi-label image classification method based on graph semantic interaction, and the method includes:

[0006] Obtain a high-dimensional visual feature map corresponding to the image to be classified Description text D corresponding to each label l and label embedding vectors corresponding to each label

[0007] According to the label embedding vectors corresponding to each label Calculate an initial similarity matrix S between each label i,j , and perform binary quantization and reweighting strategies on the initial similarity matrix S i,jProcess to obtain the label adjacency matrix A between each label; introduce a multi-scale retention mechanism module to adaptively modify the label adjacency matrix A to generate a label relationship graph A final ; where the multi-scale retention mechanism module includes: x multi-scale retention units; each of the multi-scale retention units includes: h single-scale retention layers;

[0008] Map the label relationship graph A through a graph convolutional network final to obtain high-dimensional label semantic features

[0009] Perform feature interaction between the high-dimensional visual feature map and the high-dimensional label semantic features to generate a semantic channel attention vector Z through a cross-modal attention mechanism c , and insert the semantic channel attention vector Z c into any layer of the CNN network to dynamically adjust the image features of the high-dimensional visual feature map to obtain the output image features F of the current layer final ;

[0010] Input the output image features F of the current layer final into the classification layer to obtain the classification probabilities of each label and complete the multi-label classification task.

[0011] In a possible implementation, the obtaining of the high-dimensional visual feature map corresponding to the image to be classified the description text D corresponding to each label l and the label embedding vectors corresponding to each label include:

[0012] Extract high-dimensional features from the image to be classified through a pre-trained CNN network to obtain a high-dimensional visual feature map

[0013] Generate the corresponding description text D for multiple labels on the image to be classified through a pre-trained large language model l ;

[0014] Encode the description text through a pre-trained vision-language model to obtain the label embedding vectors corresponding to each label

[0015] In a possible implementation, the processing of the initial similarity matrix S i,j through a binarization and reweighting strategy to obtain the label adjacency matrix A between each label includes:

[0016] Through a preset threshold τ for the initial similarity matrix Si,j Perform binarization processing to obtain the adjacency matrix A' i,j ;

[0017] For the adjacency matrix A' i,j Perform reweighting to obtain the label adjacency matrix A between each label.

[0018] In a possible implementation manner, introduce a multi-scale retention mechanism to adaptively modify the label adjacency matrix A to generate a label relationship graph A final , including:

[0019] Input the label adjacency matrix A into each of the multi-scale retention units to generate a first query matrix Q through three independent linear layers 1 , a first key matrix K 1 and a first value matrix V 1 ;

[0020] Based on the retention mechanism of the multi-scale retention mechanism module, use the first query matrix Q 1 , the first key matrix K 1 , the first value matrix V 1 and the relative distance matrix D to obtain the single-scale retention feature Retention(Q 1 ,K 1 ,V 1 ,γ h ) corresponding to each single-scale retention layer;

[0021] Based on the multi-scale mechanism of the multi-scale retention mechanism module, aggregate the single-scale retention features Retention(Q 1 ,K 1 ,V 1 ,γ h ) corresponding to each single-scale retention layer to generate a single subgraph adjacency matrix G corresponding to each multi-scale retention unit x ;

[0022] Concatenate multiple single subgraph adjacency matrices G x to obtain an initial label relationship graph A' final ;

[0023] Perform symmetrization processing on the initial label relationship graph A' final to obtain the label relationship graph A final .

[0024] In a possible implementation manner, map the label relationship graph A final and the label embedding vector through a graph convolutional network to obtain high-dimensional label semantic features including:

[0025] Embed the label into a vector Input the label relationship graph A final into the first-layer graph convolutional layer to generate intermediate features

[0026] Aggregate the intermediate features in the second-layer graph convolutional layer to obtain high-dimensional label semantic features

[0027] In a possible implementation, the high-dimensional label semantic features are represented as:

[0028] H (L) = GCN(E l , A final );

[0029] where c represents the number of label categories; E l represents the label embedding vectors corresponding to each label; A final represents the label relationship graph; GCN(·) represents the graph convolutional network; L represents the number of layers of the graph convolutional network; H (L) represents the high-dimensional label semantic features of the L-th layer of the graph convolutional network.

[0030] In a possible implementation, performing feature interaction between the high-dimensional visual feature map and the high-dimensional label semantic features to generate a semantic channel attention vector Z c , including:

[0031] Performing global average pooling operation on the high-dimensional visual feature map to obtain the global visual features after dimensionality reduction

[0032] Using the high-dimensional label semantic features as the second query matrix Q in the cross-modal attention mechanism 2 , and using the global visual features as the second key matrix K in the cross-modal attention mechanism 2 and the second value matrix V in the cross-modal attention mechanism 2 ;

[0033] Calculating the similarity between the second query matrix Q 2 and the second key matrix K 2 through dot product to obtain the attention score matrix

[0034] Performing on the attention score matrix Perform a Softmax operation to generate an attention weight matrix

[0035] Multiply the attention weight matrix with the second value matrix V 2 to generate a semantic channel attention vector

[0036] In a possible implementation, the output image feature F of the current layer final is expressed as:

[0037] F final =(F v '⊙Z c )+F v ';

[0038] where F v ' represents the global visual feature after dimensionality reduction; Z c represents the semantic channel attention vector, ⊙ represents the Hadamard product; c represents the number of label categories.

[0039] In a possible implementation, the classification probability is expressed as:

[0040] P = Sigmoid(W c F final +b c );

[0041] where W c represents the learnable weight; F final represents the output image feature of the current layer; b c represents the learnable bias; Sigmoid(·) represents the Sigmoid activation function.

[0042] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0043] By introducing a multi-scale retention mechanism model, the present invention adaptively learns the structure of the label relationship graph, can dynamically capture the complex dependencies between labels, generates more expressive semantic features, and significantly improves the expressive ability of semantic features. The present invention introduces a cross-modal attention mechanism to implicitly interact the high-dimensional visual feature map with the high-dimensional label semantic features, generates a semantic channel attention vector, can dynamically adjust the representation of image features, and can effectively enhance the model's ability to extract category-related features, significantly improving the classification accuracy and generalization ability, and realizing the deep fusion of semantic features and visual features. The present invention solves the key problems of insufficient label dependence modeling, shallow cross-modal fusion, and positive and negative sample imbalance in multi-label image classification, and realizes more accurate capture of the complex dependencies between labels. At the same time, through the multi-subgraph mechanism, the label semantics are dynamically embedded in the visual feature extraction process to enhance the salience of category-related features. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of the steps of the multi-label image classification method based on graph semantic interaction provided by an embodiment of the present invention;

[0045] Figure 2 It is a comparison diagram of the results of the simulation experiment provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0047] The present invention provides a multi-label image classification method based on graph semantic interaction, as Figure 1 shown, the method includes the following steps S101 to S105.

[0048] S101, obtaining a high-dimensional visual feature map corresponding to the image to be classified description text D corresponding to each label l and label embedding vectors corresponding to each label

[0049] Specifically, in step S101, obtaining a high-dimensional visual feature map corresponding to the image to be classified description text D corresponding to each label l and label embedding vectors corresponding to each label includes the following steps S1011 to S1013.

[0050] S1011. Extract high-dimensional features from the image to be classified through a pre-trained CNN network to obtain a high-dimensional visual feature map

[0051] S1012. Generate corresponding descriptive text D for multiple labels on the image to be classified through a pre-trained large language model l ;

[0052] S1013. Encode the descriptive text D through a pre-trained vision-language model l to obtain label embedding vectors corresponding to each label

[0053] Exemplarily, in step S1011, when obtaining the high-dimensional visual feature map corresponding to the image to be classified , the pre-trained CNN network can use ResNet-101 or EfficientNet etc. as the backbone network, and uniformly adjust the size of the image to be classified to 448×448 pixels. The high-dimensional visual feature map where h represents the height of the high-dimensional visual feature map; w represents the width of the high-dimensional visual feature map; d represents the number of channels of the high-dimensional visual feature map. In the present invention, the number of channels d of the high-dimensional visual feature map can be specifically set to 2048.

[0054] In step S1012, the prompt template of the pre-trained large language model is: "Generate a detailed description of the visual appearance and context of {label}." Example: Input "horse", generate "A large four-legged mammal with a mane, tail, and hooves, commonly used for riding or pulling carriages."

[0055] In step S1013, encode the generated descriptive text D l again through the text encoder in the pre-trained vision-language model to obtain label embedding vectors corresponding to each label Finally, perform L2 normalization on the label embedding vectors to prevent the curse of dimensionality.

[0056] S102. Calculate the initial similarity matrix S between each label according to the label embedding vectors corresponding to each label , and perform binarization and reweighting strategies on the initial similarity matrix S i,j i,j ​Process to obtain the label adjacency matrix A between each label; introduce a multi-scale retention mechanism module to adaptively modify the label adjacency matrix A to generate the label relationship graph A final Among them, the multi-scale retention mechanism module includes: x multi-scale retention units; each multi-scale retention unit includes: h single-scale retention layers;

[0057] Specifically, in step S102, the initial similarity matrix S i,j is processed to obtain the label adjacency matrix A between each label, including the following steps S1021 and S1022.

[0058] S1021, binarize the initial similarity matrix S i,j through a preset threshold τ to obtain the adjacency matrix A' i,j ;

[0059] S1022, re-weight the adjacency matrix A' i,j to obtain the label adjacency matrix A between each label.

[0060] Exemplarily, the initial similarity matrix S i,j is expressed as: Among them, E i represents the label embedding vector corresponding to the i-th label; E j represents the label embedding vector corresponding to the j-th label.

[0061] Binarize the initial similarity matrix S i,j through a preset threshold τ. The preset threshold τ is set manually to filter out the noise edges. The specific formula is:

[0062]

[0063] Re-weight the binarized adjacency matrix A' i,j to obtain the label adjacency matrix; adjust the edge weights:

[0064]

[0065] Among them, c represents the number of label categories, p is the weight parameter, which is optimized by grid search and the value range is [0.2, 0.8]. Finally, p is selected as 0.6. The self-connection weight (i = j) is set to 1 - p to enhance the importance of the label's own semantics.

[0066] Specifically, in step S102, introduce a multi-scale retention mechanism to adaptively modify the label adjacency matrix A to generate the label relationship graph A final , including:

[0067] (1) Input the label adjacency matrix A into each multi-scale retention unit, and generate the first query matrix Q through three independent linear layers. 1 , the first key matrix K 1 , and the first value matrix V 1 ;

[0068] (2) Based on the retention mechanism of the multi-scale retention mechanism module, utilize the first query matrix Q 1 , the first key matrix K 1 , the first value matrix V 1 , and the relative distance matrix D to obtain the single-scale retention feature Retention(Q 1 , K 1 , V 1 , γ h ) corresponding to each single-scale retention layer;

[0069] (3) Based on the multi-scale mechanism of the multi-scale retention mechanism module, aggregate the single-scale retention features Retention(Q 1 , K 1 , V 1 , γ h ) corresponding to each single-scale retention layer to generate the single subgraph adjacency matrix G x corresponding to each multi-scale retention unit;

[0070] (4) Concatenate multiple single subgraph adjacency matrices G x to obtain the initial label relationship graph A' final ;

[0071] (5) Perform symmetrization processing on the initial label relationship graph A' final to obtain the label relationship graph A final .

[0072] Exemplarily, introduce the multi-scale retention mechanism and input the label adjacency matrix where c represents the number of label categories; generate the first query matrix Q 1 , the first key matrix K 1 , and the first value matrix V 1 through three independent linear layers:

[0073]

[0074] where represents the learnable weight matrix of the query of the label adjacency matrix; represents the learnable weight matrix of the key of the label adjacency matrix; represents the learnable weight matrix of the value of the label adjacency matrix.

[0075] The single-scale retention layer adaptively changes the edges of the label adjacency matrix A by introducing an exponentially decaying relative distance matrix D, and obtains the single-scale retention feature Retention(Q 1 ,K 1 ,V 1 ,γ h ) corresponding to each single-scale retention layer. The specific expression of the single-scale retention feature Retention(Q 1 ,K 1 ,V 1 ,γ h ) is as follows:

[0076] Retention(Q 1 ,K 1 ,V 1 ,γ h )=(Q 1 K 1 T ⊙D)V 1 ;

[0077] Among them, γ h represents the decay coefficient of the h-th single-scale retention layer, which is used to control the decay speed of historical information; ⊙ represents the Hadamard product; D is the relative distance matrix, and the exponential decay of this matrix can simulate the influence of the distance between nodes. Nodes with a longer distance have a smaller influence on the current node. By controlling the decay speed with γ h , the importance of local neighbors can be emphasized while still allowing a certain degree of long-distance connection. The definition of the relative distance matrix D is as follows:

[0078]

[0079] Among them, n represents the row coordinate of the distance matrix D; m represents the column coordinate of the distance matrix D.

[0080] Subgraph structure of each multi-scale retention unit:

[0081] G x =Concat(Retention(Q x ,K x ,V x ,γ x,1 ),...,Retention(Q x ,K x ,V x ,γ x,h ));

[0082] Among them, Concat represents the concatenation operation, Retention represents the single-scale retention feature, h represents the number of single-scale retention layers, and each single-scale retention layer corresponds to an independent decay coefficient γh , with a value range of [0.8, 0.95], and dynamically learned through backpropagation.

[0083] The multi-subgraph mechanism is adopted. In this experimental example, the number of multi-scale retention units is 4, and each multi-scale retention unit outputs a single-subgraph adjacency matrix G x ∈R n×n , where x represents the x-th multi-scale retention unit. After concatenation, it is fused into the initial label relationship graph A' through a fully connected layer final :

[0084] A' final = Linear(Concat(G 1 , G 2 , G 3 , G 4 ));

[0085] Among them, Linear represents the fully connected operation. In this experimental example, for the stability of training, the initial label relationship graph A' final is also symmetrized to avoid directional errors, obtaining the label relationship graph denotes the transpose operation of A' final .

[0086] S103, map the label relationship graph A final through a graph convolutional network to obtain high-dimensional label semantic features

[0087] Specifically, in step S103, map the label relationship graph A final through a graph convolutional network to obtain high-dimensional label semantic features including the following steps S1031 to S1032.

[0088] S1031, input the label embedding vector and the label relationship graph A final into the first-layer graph convolutional layer to generate intermediate features

[0089] S1032, aggregate the intermediate features in the second-layer graph convolutional layer to obtain high-dimensional label semantic features

[0090] Exemplarily, after obtaining the label embedding vector E l in S101 and the label relationship graph A final in S102, use them as the input of the graph convolutional network. By stacking multiple layers of graph convolutional networks, gradually aggregate higher-order neighbor information to generate more expressive semantic features. The output of the L-th layer of the final graph convolutional network is H(L) That is the high-dimensional label semantic feature H (L) , specifically expressed as:

[0091] H (L) = GCN(E l , A final );

[0092] In this experiment, a two-layer graph convolutional network (GCN) was adopted, and its output dimension was gradually expanded to match the dimension of visual features. Specifically, the dimension of the network increased step by step from 512 to 1024 and finally reached 2048 to ensure the full expression and fusion of features.

[0093] Specifically, the first-layer graph convolutional layer combines the label embedding vector (c is the number of label categories) with the label relationship graph to generate intermediate features

[0094] The second-layer graph convolutional layer further aggregates the intermediate features and outputs the high-dimensional label semantic feature

[0095] After each layer of graph convolution, a LeakyReLU activation function (negative slope 0.2) is connected to enhance the non-linear expression ability. At the same time, batch normalization (BatchNorm) is added to reduce the internal covariate shift and accelerate the training convergence.

[0096] S104. Perform feature interaction between the high-dimensional visual feature map and the high-dimensional label semantic feature to generate the semantic channel attention vector Z c through a cross-modal attention mechanism, and insert the semantic channel attention vector Z c into any layer of the CNN network to dynamically adjust the image features of the high-dimensional visual feature map to obtain the output image feature F final ;

[0097] Here, the high-dimensional label semantic feature is expressed as:

[0098] H (L) = GCN(E l , A final );

[0099] Among them, E l represents the label embedding vector corresponding to each label; A final represents the label relationship graph; GCN(·) represents the graph convolutional network; L represents the number of layers of the graph convolutional network; H (L)Represents the high-dimensional label semantic features of the L-th layer graph convolutional network.

[0100] Here, the output image feature F of the current layer final Is expressed as:

[0101] F final =(F v '⊙Z c ) + F v ';

[0102] Among them, F v ' represents the global visual feature after dimensionality reduction; Z c Represents the semantic channel attention vector.

[0103] Specifically, in step S104, the high-dimensional visual feature map And the high-dimensional label semantic feature Perform feature interaction to generate the semantic channel attention vector Z c , including the following steps S1041 to S1045.

[0104] S1041, perform global average pooling operation on the high-dimensional visual feature map To obtain the global visual feature after dimensionality reduction

[0105] S1042, use the high-dimensional label semantic feature As the second query matrix Q in the cross-modal attention mechanism 2 , the global visual feature As the second key matrix K in the cross-modal attention mechanism 2 And the second value matrix V in the cross-modal attention mechanism 2 ;

[0106] S1043, calculate the similarity between the second query matrix Q 2 And the second key matrix K 2 To obtain the attention score matrix

[0107] S1044, perform Softmax operation on the attention score matrix To generate the attention weight matrix

[0108] S1045, multiply the attention weight matrix With the second value matrix V 2 To generate the semantic channel attention vector

[0109] Exemplarily, for the high-dimensional visual feature map F extracted by S101v Perform global average pooling (GAP) to obtain the global visual features after dimensionality reduction F v ′ = GAP(F v ).

[0110] The second query matrix Q in the cross-modal attention mechanism 2 is the high-dimensional label semantic features The second key matrix K 2 and the second value matrix V 2 are the global visual features Specifically, it is expressed by the formula:

[0111]

[0112] where represents the learnable weight matrix of the query in the cross-modal attention mechanism; represents the learnable weight matrix of the key in the cross-modal attention mechanism; represents the learnable weight matrix of the value in the cross-modal attention mechanism.

[0113] Map the global visual features to the same space through the learnable weight matrices and .

[0114] Calculate the similarity between the second query matrix Q 2 and the second key matrix K 2 through dot product to obtain the attention score matrix where d k = d is the dimension of the second key matrix K 2 , and the scaling factor is used to prevent the Softmax gradient from vanishing due to the too large dot product result.

[0115] Perform Softmax operation on the attention score matrix M to generate the attention weight matrix

[0116] M s = Softmax(M);

[0117] where the Softmax operation is performed along the dimension of the second key matrix K 2 to ensure the normalization of the attention weights. Each row of M s represents the attention weight distribution of a label for all image channels.

[0118] Weighted aggregate the second value matrix V s through the attention weight matrix M 2 to generate the semantic channel attention vector corresponding to each label:

[0119]

[0120] The attention vector of each label encodes the importance of the label for each channel of the image. For example, the label "dog" may have high weights on the channels corresponding to "hair texture" and "limb shape".

[0121] S105, input the output image feature F of the current layer final into the classification layer to obtain the classification probabilities of each label, and complete the multi-label classification task.

[0122] Here, the classification probability is expressed as:

[0123] P = Sigmoid(W c F final + b c );

[0124] where, W c represents the learnable weight; F final represents the output image feature of the current layer; b c represents the learnable bias; Sigmoid(·) represents the Sigmoid activation function, which is used to compress each output value between 0 and 1, representing the probability that the label is a positive class. In multi-label image classification, the present invention believes that a classification probability above 0.5 represents the existence of this category.

[0125] Generally speaking, the present invention solves the key problems of insufficient label dependence modeling, shallow cross-modal fusion, and positive and negative sample imbalance in multi-label image classification through adaptive label relationship graph construction and cross-modal feature-semantic interaction. Compared with the existing methods, this method proposes a dynamically optimized label relationship graph, combines a multi-scale semantic retention mechanism with semantic descriptions generated by a large language model, and more accurately captures the complex dependence relationships between labels; at the same time, through the semantic channel attention mechanism, the label semantics are dynamically embedded in the visual feature extraction process to enhance the saliency of category-related features.

[0126] In a simulation experiment provided by the present invention.

[0127] 1. Simulation experiment conditions:

[0128] Operating system: Ubuntu 20.02, Python3.9;

[0129] Experiment platform: Pytorch-2.1.0;

[0130] Processor: Intel Xeon Gold 6226R CPU, 64GB RAM, 1T SSD;

[0131] Graphics card: NVIDIA Tesla A100 GPU;

[0132] Memory: 64GB;

[0133] 2. Contents of simulation experiment: Network;

[0134] Simulation Experiment 1: Multi-label image classification accuracy experiment.

[0135] It should be noted that the following experiments were all conducted in the same experimental environment. Both Dataset 1 and Dataset 2 are classic datasets for multi-label image classification tasks. The baseline method and the multi-label image classification method based on graph semantic interaction proposed in the present invention both belong to the multi-label image classification algorithm based on graph convolution.

[0136] Table 1 Comparison of accuracy metrics between the baseline method and the method proposed in the present invention under Dataset 1

[0137]

[0138] Table 2 Comparison of accuracy metrics between the baseline method and the method proposed in the present invention under Dataset 2

[0139] model mAP OP OR CF1 OF1 baseline method 83.4 82.9 76.6 78.0 79.6 the present invention 93.632 87.769 85.320 85.057 86.527

[0140] As can be seen from Table 1 and Table 2 according to the experimental results, the present model performs excellently in multiple evaluation metrics. Especially in terms of the mean average precision (mAP), it successfully surpasses the previous models. More significantly, the present model exceeds other similar models in the recognition accuracy of the vast majority of categories, showing excellent classification ability. For some common categories, the accuracy of the present model even reaches 1, which undoubtedly proves its advantages in feature extraction, classification determination, and model optimization, further verifying the superiority and practicality of the model.

[0141] Compared with Dataset 1, Dataset 2 is larger in scale, has more annotation types, and has higher universality. For this more complex and diverse dataset, our model significantly improves the recall rate on the premise of maintaining the accuracy. From the perspective of global evaluation metrics such as mAP, CF1, and OF1, the performance of our model comprehensively surpasses the previous models, showing excellent performance. It is worth noting that even under the condition of lower-resolution images, the model can still achieve better performance, which further proves its excellent effect in multi-label classification tasks.

[0142] Simulation Experiment 2: Visualization experiment of the multi-label image classification method based on graph semantic interaction and the benchmark method, as Figure 2 shown.

[0143] In the first image (within the blue dashed box), the baseline model successfully recognized "horse" and "person", but its heatmap focused on irrelevant regions, while the present invention focuses on the positions related to the labels through the semantic channel attention mechanism, eliminating redundant information. In the second image (within the red dashed box), although the baseline model recognized "sofa", its recognition of "person" and "TV monitor" was poor, and the heatmap contained a lot of irrelevant details; the present invention accurately classified all labels and concentrated attention on the relevant regions, demonstrating stronger semantic recognition ability.

[0144] In the third image (within the yellow dashed box), the baseline model misjudged due to the similarity between glass and bottle, while the present invention can correctly distinguish between the two. When recognizing the table, the baseline model focused on the objects on the table, while the present invention captured the characteristics of the table itself, which benefited from the label relationship graph generated using image features and focused on category-specific features. In the fourth image (within the green dashed box), although there is no chair in the scene, the present invention placed attention on the leg and back regions of the person, implying the absence of a chair, indicating that it can effectively reduce semantic noise and capture semantic information more accurately. Generally speaking, through the semantic channel attention mechanism and the label relationship graph, the present invention significantly improves the accuracy of multi-label classification, can more precisely focus on the regions related to the labels, reduce the interference of irrelevant information, and at the same time enhance the model's understanding ability of semantic relationships in complex scenes.

[0145] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. All or part of the present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0146] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the present invention; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present invention.

Claims

1. A multi-label image classification method based on graph semantic interaction, characterized in that: include: Get the high-dimensional visual feature map corresponding to the image to be classified Description text D corresponding to each label l The label embedding vector corresponding to each label According to the label embedding vector corresponding to each label Calculate the initial similarity matrix S between each label i,j , the initial similarity matrix S is transformed into i,j Processing is performed to obtain a label adjacency matrix A between each label; a multi-scale retention mechanism module is introduced to adaptively modify the label adjacency matrix A to generate a label relationship graph A final ; Wherein, the multi-scale retention mechanism module includes: x multi-scale retention units; each of the multi-scale retention units includes: h single-scale retention layers; The label relationship graph A is transformed through the graph convolution network final Mapping to obtain high-dimensional label semantic features The high-dimensional visual feature map With the high-dimensional label semantic features Perform feature interaction and generate a semantic channel attention vector Z through a cross-modal attention mechanism c , and the semantic channel attention vector Z c , inserted into any layer of the CNN network to dynamically adjust the high-dimensional visual feature map The image features of the current layer are obtained by final ; The output image feature F of the current layer final Input into the classification layer to obtain the classification probability of each label and complete the multi-label classification task.

2. The multi-label image classification method based on graph semantic interaction according to claim 1 is characterized in that: The step of obtaining a high-dimensional visual feature map corresponding to the image to be classified Description text D corresponding to each label l The label embedding vector corresponding to each label include: The pre-trained CNN network is used to extract high-dimensional features of the classified image to obtain a high-dimensional visual feature map. Generate corresponding description text D for multiple labels on the image to be classified through a pre-trained large language model l ; The description text is encoded through the pre-trained visual language model to obtain the label embedding vector corresponding to each label 3. The multi-label image classification method based on graph semantic interaction according to claim 1 is characterized in that: The initial similarity matrix S is processed by binarization and reweighting strategy. i,j After processing, the label adjacency matrix A between each label is obtained, including: The initial similarity matrix S is calculated by using a preset threshold τ. i,j Perform binarization to obtain the adjacency matrix A' i,j ; For the adjacency matrix A' i,j Reweighting is performed to obtain the label adjacency matrix A between each label.

4. The multi-label image classification method based on graph semantic interaction according to claim 3 is characterized in that: The multi-scale retention mechanism is introduced to adaptively modify the label adjacency matrix A to generate a label relationship graph A. final ,include: Input the label adjacency matrix A into each of the multi-scale preservation units, and generate a first query matrix Q1, a first key matrix K1 and a first value matrix V1 through three independent linear layers; Based on the retention mechanism of the multi-scale retention mechanism module, the single-scale retention feature Retention (Q1, K1, V1, γ h ); Based on the multi-scale mechanism of the multi-scale retention mechanism module, the single-scale retention feature Retention (Q1, K1, V1, γ h ), generate a single subgraph adjacency matrix G corresponding to each multi-scale retention unit x ; Multiple single subgraph adjacency matrices G x Splice and get the initial label relationship graph A' final ; The initial label relationship graph A' final After symmetry processing, we get the label relationship diagram A final .

5. The multi-label image classification method based on graph semantic interaction according to claim 1, characterized in that: The label relationship graph A is transformed into final and the label embedding vector Mapping to obtain high-dimensional label semantic features include: Embed the label into a vector A diagram showing the relationship between the label final Input into the first layer of graph convolution layer to generate intermediate features The intermediate feature Aggregation is performed in the second graph convolution layer to obtain high-dimensional label semantic features 6. The multi-label image classification method based on graph semantic interaction according to claim 5 is characterized in that: The high-dimensional label semantic features It is expressed as: H (L) =GCN(E l ,A final ); Among them, c represents the number of label categories, E l Represents the label embedding vector corresponding to each label; A final represents the label relationship graph; GCN(·) represents the graph convolutional network; L represents the number of layers of the graph convolutional network; H (L) Represents the high-dimensional label semantic features of the L-th layer graph convolutional network.

7. The multi-label image classification method based on graph semantic interaction according to claim 1, characterized in that: The high-dimensional visual feature map With the high-dimensional label semantic features Perform feature interaction and generate a semantic channel attention vector Z through a cross-modal attention mechanism c ,include: For the high-dimensional visual feature map Perform global average pooling operation to obtain the global visual features after dimensionality reduction The high-dimensional label semantic features As the second query matrix Q2 in the cross-modal attention mechanism, the global visual features As the second key matrix K2 in the cross-modal attention mechanism and the second value matrix V2 in the cross-modal attention mechanism; The similarity between the second query matrix Q2 and the second key matrix K2 is calculated by dot product to obtain the attention score matrix For the attention score matrix Perform Softmax operation to generate attention weight matrix The attention weight matrix Multiply it with the second value matrix V2 to generate the semantic channel attention vector 8. The multi-label image classification method based on graph semantic interaction according to claim 1, characterized in that: The output image feature F of the current layer final It is expressed as: F final =(F v ′⊙Z c )+F v ′; Among them, F v ′ represents the global visual feature after dimensionality reduction; Z c represents the semantic channel attention vector, ⊙ represents the Hadamard product; c represents the number of label categories.

9. The multi-label image classification method based on graph semantic interaction according to claim 1, characterized in that: The classification probability is expressed as: P=Sigmoid(W c F final +b c ); Among them, W c represents the learnable weight; F final Represents the output image features of the current layer; b c represents the learnable bias; Sigmoid(·) represents the Sigmoid activation function; c represents the number of label categories.