Multi-label text classification method based on dual feature fusion and multi-relational graph convolutional network

By combining Chinese-BERT-wwm, TextCNN, and CompGCN models, the problems of data imbalance and label correlation in multi-label text classification are solved, improving the accuracy and recall of multi-label text classification and achieving a more comprehensive revelation of semantic information.

CN120950696APending Publication Date: 2025-11-14TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511132526.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Multi-label text classification suffers from problems such as data imbalance, insufficient text feature extraction, and inadequate modeling of the correlation between labels, resulting in poor model performance.

Method used

We use Chinese-BERT-wwm and TextCNN models to extract text features, perform feature fusion through cross-attention gating residual mechanism, construct a multi-relationship graph structure, use CompGCN to extract graph features, and design head and tail label classifiers and comprehensive loss function to optimize the problem of imbalanced data distribution.

Benefits of technology

It improves the model's performance when processing low-frequency labeled text, enhances the understanding of label associations, improves the accuracy and recall of multi-label text classification, and achieves a more comprehensive revelation of semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950696A_ABST
    Figure CN120950696A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-label text classification, and discloses a multi-label text classification method based on dual feature fusion and a multi-relational graph convolutional network, which comprises the following steps: firstly, extracting text features, extracting text global features by adopting a Chine-BERT-wwm pre-training model, extracting text local features by utilizing a TextCNN model, and extracting text global features by adopting a TextCNN model; effective fusion of global features and local features is realized through a cross attention gating residual mechanism; secondly, mining the relevance of the labels, representing the complex relevance between the labels by constructing a multi-relation graph structure, and performing graph feature learning through CompGCN; meanwhile, aiming at the problem of unbalanced data distribution, a head and tail label classifier and a comprehensive loss function are designed; and finally, by taking the ceramic comments as research objects, constructing data sets of three scales, and comparing and analyzing the data sets with nine reference models, so that the effect of a multi-label text classification task is effectively improved, and the effectiveness of the method is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning, natural language processing, and multi-label text classification. More specifically, it designs a multi-label text classification method based on dual feature fusion and multi-relation graph convolutional networks. Background Technology

[0002] With the development of internet technology, text data has experienced explosive growth. Extracting valuable information from this massive and complex dataset has become a crucial issue in Natural Language Processing (NLP). Text classification, as one of the core tasks of NLP, aims to automatically categorize text data according to predefined classes, serving as a vital means of information extraction and organization. Through text classification, massive amounts of text data can be processed efficiently, in applications such as news categorization, sentiment analysis, spam filtering, and topic tagging.

[0003] Multi-label text classification differs from traditional single-label text classification tasks in that it aims to assign text data to multiple predefined categories simultaneously. This classification approach better reflects the complexity and diversity of text content in the real world, and can reveal the semantic information of the text more comprehensively and meticulously.

[0004] Despite the immense potential of multi-label text classification in both theory and application, it still faces numerous challenges in practical applications. First, data imbalance severely restricts model performance. In multi-label text classification tasks, the uneven distribution of labels leads to overfitting of high-frequency labels and underfitting of low-frequency labels during training, resulting in a significant performance drop when processing low-frequency label text. Second, insufficient extraction of text features further impairs model performance. Text data exhibits high heterogeneity in grammatical structure, semantic expression, and sentiment, significantly increasing the difficulty of semantic understanding. This prevents models from fully learning effective text features during training, leading to poor model performance. Finally, insufficient modeling of inter-label relationships further degrades model performance. Labels in multi-label classification tasks are not independent but possess complex structures such as semantic associations, hierarchical dependencies, and co-occurrence relationships. However, existing methods often use independent binary classifiers to process labels, neglecting the inter-label relationships and preventing the model from fully utilizing the semantic constraints between labels. This makes it difficult for the model to make a comprehensive judgment based on label association when making label predictions, resulting in a deviation between the classification results and the actual situation, and failing to give full play to the advantages of multi-label classification. Summary of the Invention

[0005] To address the above problems, this invention provides a multi-label text classification method based on dual feature fusion and multi-relation graph convolutional networks, comprising the following steps:

[0006] Step 1: Extract text features. Use the Chinese-BERT-wwm pre-trained model to extract global text features, and use the TextCNN model to extract local text features. Then, use the cross-attention gating residual mechanism to effectively fuse global and local features.

[0007] Step 2: Construct a multi-relationship graph structure to represent the complex relationships between tags;

[0008] Step 3: Use CompGCN to extract features from the multi-relationship graph;

[0009] Step four: Perform multimodal feature fusion of text features and graph features;

[0010] Step 5: To address the issue of uneven data distribution, a head and tail label classifier was designed.

[0011] Step six: To address the issue of uneven data distribution, a comprehensive loss function was designed.

[0012] Step 7: Taking ceramic reviews as the research object, three datasets of different sizes were constructed and compared with nine benchmark models to verify the superiority of the proposed method under different datasets.

[0013] Furthermore, the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional networks has the following model framework diagram: Figure 1 As shown:

[0014] Furthermore, step one specifically includes:

[0015] Step 1.1: Extract global features of the text using the Chinese-BERT-wwm pre-trained model. Chinese-BERT-wwm is a pre-trained language model for Chinese that uses a technique called Whole Word Masking. Its model structure is as follows: Figure 2 As shown. This technique simultaneously masks all Chinese characters that make up a complete word, avoiding the semantic fragmentation problem caused by traditional character masking. It can better handle complex words and phrases in Chinese text, thereby improving the model's performance.

[0016] Step 1.2: Extract local text features using the TextCNN (Text Convolutional Neural Network) model. TextCNN is a deep learning model widely used in natural language processing tasks, and its model structure is as follows: Figure 3 As shown. Its core idea is to extract local features from text through convolution operations and pool features from different locations through pooling layers to generate a fixed-length feature vector, providing an effective feature representation for NLP tasks;

[0017] Step 1.3: A cross-attention gating residual fusion module is used to fuse global and local features, ultimately generating an enhanced semantic representation. Specifically:

[0018] First, global and local features are unified to the same feature dimension through linear projection, so that they can interact effectively in the latent space, as shown in the following formula:

[0019]

[0020]

[0021] in, and It is the result of linear projection transformation of global and local features. and It is a weight matrix. and It is a bias term;

[0022] Then, a cross-attention mechanism is introduced to enhance the interaction between global and local features, as shown in the following formula:

[0023]

[0024]

[0025] in, It is a local feature ( ), It is a global feature ( ), These are target features. It is a scaling factor used to prevent the gradient from vanishing due to an excessively large dot product result. It is the attention weight matrix. It is a global feature ( ), It is attention output;

[0026] Finally, global and local features are fused using a dynamic gated residual fusion mechanism, as shown in the following formula:

[0027]

[0028]

[0029] in, It is a weight matrix. It is a bias term. It is the sigmoid activation function. It's a splicing operation. It is the generated gating signal. and These are adaptive residual parameters, which control the proportions of global and local features in the fused features, respectively.

[0030] Furthermore, step two specifically involves:

[0031] Step 2.1: Extract specific label columns from the dataset and construct a label co-occurrence matrix. The label co-occurrence matrix is ​​a binary matrix where rows and columns represent different labels, and element values ​​represent the number of times each label appears simultaneously.

[0032] Step 2.2: Convert the tag co-occurrence matrix into a co-occurrence probability matrix. Each element in the co-occurrence probability matrix represents the conditional probability of tags appearing together, calculated as follows:

[0033]

[0034] in, These are elements in the label co-occurrence matrix, representing the labels. and tags The number of times they occur simultaneously; It is a tag Total number of occurrences; It is a very small constant, to prevent division by zero;

[0035] Step 2.3: Based on the co-occurrence probability matrix, discretize the continuous probability values ​​into relation types using predefined threshold intervals: each probability interval is mapped to a specific semantic edge type. The edge types are defined as follows:

[0036]

[0037] in, These are elements in the co-occurrence probability matrix, representing the label. and tags The probability of them occurring simultaneously;

[0038] Step 2.4: Construct a multi-relationship graph based on the generated weighted edges and nodes. Nodes represent labels, edges represent co-occurrence relationships between labels, edge weights quantify co-occurrence probability values, and edge types represent the discretized relationship strength levels.

[0039] Furthermore, step three specifically includes:

[0040] Step 3.1: CompGCN aggregates the neighbor information of nodes through a message passing mechanism. During the computation of each layer of the network, the model traverses all its neighbor nodes for each node in the graph and aggregates the feature information of the neighbor nodes in a targeted manner based on the differences in the relationship types of the edges.

[0041] Step 3.2: CompGCN uses relational combination operators. Features of neighboring nodes and relational embedding Deep recombination is performed; commonly used combinatorial operators include addition, multiplication, and scalar multiplication. In this way, CompGCN can deeply mine the hidden semantic connections between nodes under different relation types, enabling the model to dynamically and flexibly adjust the information fusion strategy based on the semantic differences of edge relations when processing multi-relation graph data.

[0042] Step 3.3: After completing message passing and composite operations, CompGCN achieves deep learning of the hierarchical structure information of graph data through node feature updates. It adopts a multi-layered stacked architecture, updating node features layer by layer to capture information from the original nodes to the deeper graph structure, enabling it to better handle complex graph data. The update formula is shown below:

[0043]

[0044] in, It is the weight matrix of the node's own feature transformation; It is a correspondence. The neighborhood weight matrix is ​​used to determine the aggregate weights of the features of neighboring nodes with different relationships; It is an activation function. It is a node In the Layer feature representation, It is a node In the Layer feature representation, It is a node In the Layer neighbor nodes Feature representation; It is a node In relation types The set of neighboring nodes.

[0045] Furthermore, step four specifically involves:

[0046] The fused text features are concatenated with the graph features processed by CompGCN to form a new joint feature vector. This joint vector serves as the input to the fully connected layer, effectively integrating textual semantic information and graph structural information, and providing a suitable feature representation for subsequent multi-label text classification tasks.

[0047] Furthermore, step five specifically includes:

[0048] Step 5.1: Based on the category distribution characteristics of the data, the dataset is hierarchically divided. By counting the number of samples corresponding to each category, the boundary between the head and tail categories is determined according to the preset division ratio, dividing all data categories into a head category set with abundant samples and a tail category set with sparse samples;

[0049] Step 5.2: For the head category data, this study constructed a head label classifier, which includes two fully connected layers, a GeLU activation function, and a dropout layer. This structure, through nonlinear transformation and regularization mechanisms, can effectively mine common features in the head category data, and finally output a feature vector equal to the number of head categories. The design formula is as follows:

[0050]

[0051] in, The feature vector obtained after fusing multimodal features; It is the first fully connected layer, which maps the input feature vector; It is a regularization technique used in neural networks to prevent model overfitting; Activation functions can effectively improve the generalization ability and performance of the model; the second fully connected layer is used to map the activated output to a feature vector equal to the number of head categories;

[0052] Step 5.3: For the tail category data, considering the difficulty of feature learning due to sample scarcity, a tail label classifier was designed. This classifier first uses an attention mechanism to adaptively weight the input features, strengthening the representation ability of key features; simultaneously, it extracts long-distance dependencies through a Transformer encoder and uses a fully connected layer to output feature vectors matching the number of tail categories, thereby effectively improving the model's generalization performance for low-frequency categories. Its design formula is as follows:

[0053]

[0054] in, The feature vector obtained after fusing multimodal features; Activation functions can maintain numerical stability while introducing nonlinearity; yes Encoder layer.

[0055] Furthermore, step six specifically includes:

[0056] Step 6.1: The cross-entropy loss introduces a dynamic weighting mechanism. By dynamically balancing the learning weights of positive and negative samples, it enhances the model's attention to tail-class data. The specific formula is as follows:

[0057]

[0058]

[0059] in, It is the first The weights of each sample, weights It is determined by calculating the imbalance ratio of positive and negative samples, thereby increasing the model's attention to the tail category data and reducing the model's bias towards the head category data. It is the Sigmoid function. These are the raw scores output by the model;

[0060] Step 6.2, Contrastive Loss: The core of contrastive loss is prototype labels. By comparing the similarity between samples and prototype labels, it guides the model to learn features between categories. The logic is to aggregate features of samples from the same category and decouple features of samples from different categories, thereby strengthening the model's ability to distinguish between different categories. Its calculation formula is as follows:

[0061]

[0062]

[0063] in, It is the similarity between the sample and the positive class prototype label; It is the similarity between the sample and the negative class prototype label; It is the first Contrast loss for each sample; It is the sample size; It is the average contrast loss of all samples;

[0064] Step 6.3: The overall loss is the weighted sum of the cross loss and the comparison loss, and its calculation formula is as follows:

[0065]

[0066] in, It is a weighting parameter used to balance the losses of the two parts.

[0067] Through collaborative optimization, cross-entropy loss ensures the model's basic fitting ability to the label distribution, while contrastive loss enhances the class discrimination ability of the feature space. Especially for tail categories, it can improve the robustness of their feature representation by learning from similar categories, ultimately optimizing the overall classification accuracy.

[0068] Furthermore, step seven specifically includes:

[0069] Step 7.1: Taking ceramic products as the background, we crawled the review data of ceramic products and constructed three multi-label datasets with sample sizes of 1,000, 2,000 and 3,000 respectively. The review content was divided into six dimensions (quality, service, price, packaging, transportation and suggestions) and labeled with multiple tags.

[0070] Step 7.2: Select nine models in the field of multi-label text classification and conduct comparative experiments with the proposed method to verify the superiority of the proposed method under different datasets. Attached Figure Description

[0071] Figure 1 For technology roadmap

[0072] Figure 2 Model framework diagram

[0073] Figure 3 BERT model diagram

[0074] Figure 4 For the TextCNN model diagram

[0075] Figure 5 This is a comparison chart of the experimental results of the present invention. Detailed Implementation Plan

[0076] The multi-label text classification problem in this embodiment is described as follows:

[0077] Given a text dataset ,in Indicates the first A text sample, This represents the set of tags associated with the text. , (For the set of all possible labels). The task of multi-label text classification is to learn a model. This makes it possible for any text The model can predict its associated set of labels. ;

[0078] In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, taking ceramic products as the background, the review data of ceramic products is crawled, and the review content is divided into six dimensions (quality, service, price, packaging, transportation, and suggestions) for multi-label annotation;

[0079] In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, three datasets of different sizes were constructed, as shown in the table below:

[0080] Table 1 Dataset Information

[0081] Dataset training set test set Total word count Number of tags Dataset-S 800 200 12966 6 Dataset-M 1600 400 26287 6 Dataset-L 2400 600 53458 6

[0082] In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, in order to verify the effectiveness of the proposed algorithm model, nine related models in the field of multi-label text classification are selected for comparative experiments, specifically including the following models: BERT, DistilBERT, Bi-LSTM, BBN, CNN, CNN-RNN, HTTN, S2S-Attn, and S2S-LSAM.

[0083] In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, the superiority of the proposed method has not been verified. Instead, it is compared with other related models on three different datasets. The comparison results are as follows: Figure 5 As shown. In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, in order to comprehensively evaluate the performance of the proposed algorithm in the multi-label text classification task, the micro-averaged evaluation index is adopted as the core evaluation standard, which specifically includes micro-averaged precision, micro-averaged recall and micro-averaged F1 score.

[0084] In the multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network disclosed in this invention, in order to comprehensively evaluate the performance of the proposed algorithm in the multi-label text classification task, the micro-averaged evaluation index is adopted as the core evaluation standard, which specifically includes micro-averaged precision, micro-averaged recall and micro-averaged F1 score.

[0085] Micro-precision represents the global percentage of true positives among all predicted positives by the model, reflecting the model's overall ability to avoid false positives; the calculation formula is as follows:

[0086] Micro-recall represents the global percentage of all true positives that are correctly identified, reflecting the model's overall ability to avoid false negatives.

[0087] The Micro-F1 score, as the harmonic mean of Micro-precision and Micro-recall, comprehensively measures the balanced performance of the model's global precision and recall; the calculation formula is as follows:

[0088]

[0089]

[0090]

[0091] in, Indicates the number of categories of the label. This represents the number of samples that were predicted to be positive and were actually positive. This represents the number of samples that were actually positive examples but were predicted as negative examples.

[0092] In comparative experiments, such as Figure 5 As can be observed, our proposed method achieves superior performance compared to other benchmark models. It is noteworthy that while some baseline models excel in recall, their precision and F1 scores are significantly lower than our proposed method. This indicates that such models tend to overpredict positive examples (high recall) but contain a large number of errors (low precision), resulting in poor overall performance in distinguishing the true class of samples. Conversely, our proposed method demonstrates significant advantages in all three core metrics: microprecision, microrecall, and microF1 score, validating its excellent overall performance and superior balance.

[0093] The above detailed description further illustrates the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-label text classification method based on dual feature fusion and multi-relation graph convolutional networks, characterized in that, Includes the following steps: Step 1: Extract text features. Use the Chinese-BERT-wwm pre-trained model to extract global text features, and use the TextCNN model to extract local text features. Then, use the cross-attention gating residual mechanism to effectively fuse global and local features. Step 2: Construct a multi-relationship graph structure to represent the complex relationships between tags; Step 3: Use CompGCN to extract features from the multi-relationship graph; Step four: Perform multimodal feature fusion of text features and graph features; Step 5: To address the issue of uneven data distribution, a head and tail label classifier was designed. Step six: To address the issue of uneven data distribution, a comprehensive loss function was designed. Step 7: Taking ceramic reviews as the research object, three datasets of different sizes were constructed and compared with nine benchmark models to verify the superiority of the proposed method under different datasets.

2. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The specific implementation process of step one is as follows: Step 1.1: Use the Chinese-BERT-wwm model to extract features from the input text, and obtain the global feature representation and word-level vector representation of the text; Step 1.2: Using the TextCNN model, multi-scale convolution kernels are used to perform convolution operations on word-level vectors to obtain local feature representations in the text; Step 1.3: Employ a cross-attention gating residual mechanism to fuse global and local text features, thereby obtaining a more comprehensive text feature representation. The cross-attention gating residual calculation process is as follows: ; ; in, and It is the result of linear projection transformation of global and local features. and It is a weight matrix. and It is a bias term; ; ; in, It is a local feature ( ), It is a global feature ( ), These are target features. It is a scaling factor used to prevent the gradient from vanishing due to an excessively large dot product result. It is the attention weight matrix. It is a global feature ( ), It is attention output; ; ; in, It is a weight matrix. It is a bias term. It is the sigmoid activation function. It's a splicing operation. It is the generated gating signal. and These are adaptive residual parameters, which control the proportions of global and local features in the fused features, respectively.

3. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The specific implementation process of step two is as follows: Step 2.1: Extract specific label columns from the dataset and construct a label co-occurrence matrix. The label co-occurrence matrix is ​​a binary matrix where rows and columns represent different labels, and element values ​​represent the number of times each label appears simultaneously. Step 2.2: Convert the tag co-occurrence matrix into a co-occurrence probability matrix. Each element in the co-occurrence probability matrix represents the conditional probability of tags appearing together. The calculation formula is as follows: ; in, These are elements in the label co-occurrence matrix, representing the labels. and tags The number of times they occur simultaneously; It is a tag Total number of occurrences; It is a very small constant, to prevent division by zero; Step 2.3: Based on the co-occurrence probability matrix, discretize the continuous probability values ​​into relation types using predefined threshold intervals: each probability interval maps to a specific semantic edge type. The edge types are defined as follows: ; in, These are elements in the co-occurrence probability matrix, representing the label. and tags The probability of them occurring simultaneously; Step 2.4: Construct a multi-relationship graph based on the generated weighted edges and nodes. Nodes represent labels, edges represent co-occurrence relationships between labels, edge weights quantify co-occurrence probability values, and edge types represent the discretized relationship strength levels.

4. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The specific implementation process of step three is as follows: Step 3.1: CompGCN aggregates the neighbor information of nodes through a message passing mechanism. During the computation of each layer of the network, the model traverses all its neighbor nodes for each node in the graph and aggregates the feature information of the neighbor nodes in a targeted manner based on the differences in the relationship types of the edges. Step 3.2: CompGCN uses relational combination operators. Features of neighboring nodes and relational embedding Perform in-depth recombination; Step 3.3: After completing message passing and composite operations, CompGCN achieves deep learning of the hierarchical structure information of the graph data through node feature updates. The update formula is shown below: ; in, It is the weight matrix of the node's own feature transformation; It is a correspondence. The neighborhood weight matrix is ​​used to determine the aggregate weights of the features of neighboring nodes with different relationships; It is an activation function. It is a node In the Layer feature representation, It is a node In the Layer feature representation, It is a node In the Layer neighbor nodes Feature representation; It is a node In relation types The set of neighboring nodes.

5. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The text features fused in step one and the graph features extracted in step three are concatenated and spliced ​​together to form a new joint feature vector.

6. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The specific implementation process of step five is as follows: Step 5.1: Based on the category distribution characteristics of the data, the dataset is divided into hierarchical segments. By counting the number of samples corresponding to each category, the boundary between the head and tail categories is determined according to the preset division ratio, and all data categories are divided into a head category set with abundant samples and a tail category set with sparse samples. Step 5.2: For the head category data, a head label classifier was constructed. This classifier consists of two fully connected layers, a GeLU activation function, and a dropout layer. Its design formula is as follows: ; in, The feature vector obtained after fusing multimodal features; It is the first fully connected layer, which maps the input feature vector; It is a regularization technique used in neural networks to prevent model overfitting; Activation functions can effectively improve the generalization ability and performance of the model; the second fully connected layer is used to map the activated output to a feature vector equal to the number of head categories; Step 5.3: For the tail category data, a tail label classifier was designed. This classifier first uses an attention mechanism to adaptively weight the input features, strengthening the representation ability of key features. Simultaneously, it extracts long-distance dependencies through a Transformer encoder and uses a fully connected layer to output feature vectors matching the number of tail categories, thereby effectively improving the model's generalization performance for low-frequency categories. Its design formula is as follows: ; in, The feature vector obtained after fusing multimodal features; Activation functions can maintain numerical stability while introducing nonlinearity; yes Encoder layer.

7. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, The specific implementation process of step six is ​​as follows: Step 6.1: The cross-entropy loss introduces a dynamic weighting mechanism. By dynamically balancing the learning weights of positive and negative samples, it enhances the model's attention to tail-class data. The specific formula is as follows: ; ; in, It is the first The weights of each sample, weights It is determined by calculating the imbalance ratio of positive and negative samples, thereby increasing the model's attention to the tail category data and reducing the model's bias towards the head category data. It is the Sigmoid function. These are the raw scores output by the model; Step 6.2: Contrastive Loss introduces prototype labels. By comparing the similarity between the sample and the prototype labels, the model is guided to learn features between categories. The specific formula is as follows: ; ; in, It is the similarity between the sample and the positive class prototype label; It is the similarity between the sample and the negative class prototype label; It is the first Contrast loss for each sample; It is the sample size; It is the average contrast loss of all samples; Step 6.3: The overall loss is the weighted sum of the cross loss and the comparison loss, and its calculation formula is as follows: ; in, It is a weighting parameter used to balance the losses of the two parts.

8. The multi-label text classification method based on dual feature fusion and multi-relation graph convolutional network according to claim 1, characterized in that, We crawled user review data from the JD.com e-commerce platform and constructed three multi-label datasets with sample sizes of 1000, 2000, and 3000, respectively. We then conducted comparative experiments with nine multi-label text classification models to verify the superiority of the proposed method on datasets of different sizes.