A multi-label text classification method and system based on label relationships
By adopting a multi-stage re-ranking method based on label relationships in multi-label text classification, combining label co-occurrence information and frequency distribution information, and establishing semantic relationships between text and labels through attention mechanisms, the problem of difficulty in capturing label relationships in the existing technology is solved, and the accuracy of classification and the ability to predict tail labels are improved.
Patent Information
- Application Number
- CN202510020631.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The prior art is difficult to effectively capture the relationship between labels and between labels and text in multi-label text classification, especially in the long-tail phenomenon, and the classification ability of tail labels is limited.
Through a multi-stage re-ranking method based on label relationships, combining tag co-occurrence information and tag frequency distribution information, a frequency-integrated tag sequence is generated, and a semantic relationship between text and tags is established through attention mechanisms, and finally classified.
It improves the accuracy and relevance of multi-label text classification, effectively alleviates the long-tail problem, and enhances the prediction ability of tail labels.
Smart Images

Figure CN119938924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a multi-label text classification method and system based on label relationships. Background Art
[0002] Multi-label text classification is a key task in natural language processing. Its goal is to assign appropriate labels to given texts. It also has wide applications in other fields such as recommendation systems, sentiment analysis, rumor detection, and question-answering tasks.
[0003] The latest progress in deep learning has promoted the development of models based on convolutional neural networks and recurrent neural networks, which have demonstrated excellent feature extraction capabilities. However, these models face challenges in capturing the context relationships of words in long texts. For example, although convolutional neural networks can effectively capture local patterns, they struggle in dealing with long-distance dependencies. Recurrent neural networks, although better at handling sequential data, are often troubled by problems such as vanishing gradients and thus it is difficult to effectively learn long-term dependencies. The introduction of the attention mechanism has alleviated some of these problems, and models combine the attention mechanism to improve the semantic representation and classification performance of texts. And models based on Transformer and powerful pre-trained models have further enhanced the text feature extraction ability. These models use self-attention mechanisms to capture the relationships between words in the sequence, thus being able to better handle long-distance dependencies and improve the overall semantic understanding of texts.
[0004] Despite these advancements, the existing technologies mainly focus on extracting text features and basically follow the pattern of extracting text features through neural networks and then classifying with the text features through a linear layer. The improvement points also focus on strengthening text feature extraction and often ignore the relationships between labels and between labels and texts. And in some cases, relying solely on text features is not sufficient to predict the correct labels, especially when the text features cannot fully represent the label semantics. At this time, label relationships can help predict these labels. And for tail labels that are difficult to establish the connection between features and labels due to the lack of training samples, additional label information is even more helpful for their prediction. For example, in medical document classification, the presence of one disease label may imply the possibility of another disease occurring simultaneously, and many existing models fail to fully capture these relationships.
[0005] The long-tail problem in multi-label text classification is a crucial and challenging task. The long-tail phenomenon refers to the fact that there are few instances of a large number of labels, which makes it difficult to accurately classify labels with lower frequencies. In some cases, relying solely on text features is not sufficient to predict the correct labels, especially when the text features cannot fully represent the semantic meanings of the labels. This problem is particularly evident for tail labels because tail labels lack sufficient training samples and are thus more difficult to associate with text features. For example, when classifying scientific papers based on abstracts, the short text of the abstract may not fully convey the semantics of all labels, or the text may express in a more obscure and ambiguous way.
[0006] In this case, it is necessary to utilize the relationships between labels to identify tail labels that cannot be directly seen from the text semantics. Although some methods attempt to incorporate label information into the classification process. For example, some use attention mechanisms to explore the relationships between label semantics and text. However, they do not consider the relationships between labels. Other researchers use graph networks to model labels and explore the relationships between labels. But in these studies, the expression of the same label information is consistent across different samples, ignoring the fact that the importance of different label information should vary for different texts.
[0007] In the first prior art, it is proposed to enhance the adaptation ability of the pre-trained model for classifying a large number of labels by using multiple special tokens [CLS]. For a classification task with a large number of labels, the feature size may be insufficient. For example, the typical embedding dimension of a model is 256, then BERT will also output a 256-dimensional hidden state for the token [CLS]. Suppose there are 10,000 different labels, then the information contained in the 256-dimensional feature vector may not be sufficient to perfectly support the 10,000-dimensional label prediction. And increasing the feature dimension will significantly increase the model size and reduce the model efficiency because the embeddings, hidden states, and intermediate self-attention results will also become larger. Therefore, the first prior art uses multiple special tokens to adapt to classification tasks with a larger number of labels. BERT's multi-head self-attention mechanism allows each special token to traverse the entire input sequence, and the final hidden state of each special token can be regarded as a specific feature vector. Specifically, X symbols [CLS_1], [CLS_2],..., [CLS_X] for classification can be added in front of each input sequence, and the concatenation of their final hidden states is used as the aggregated sequence representation. The cost of using multiple [CLS] tokens to obtain a wider feature vector is much smaller than increasing the feature dimension. Finally, this feature vector is used for classification.
[0008] However, there are the following drawbacks in the prior art one: This technology mainly strengthens the feature extraction ability of the pre-trained model by adding special marks of CLS, and obtains more comprehensive feature information in a way that has a relatively small impact on the model efficiency to assist classification. However, this method is still limited to feature extraction. However, some labels do not directly reflect from the semantics, especially for the tail labels. Due to the lack of training samples, the model fails to recognize the association between the features and the labels. Strengthening feature extraction alone has little effect. Moreover, this technology does not integrate the label relationship information other than features to assist classification, which limits the classification ability for tail labels.
[0009] In the prior art two, an interpretable graph convolutional network model is proposed. This model first uses the pre-trained model to extract the features of tokens, and then embeds the tokens and labels as nodes in a heterogeneous graph, and constructs edges of mark-mark, mark-label, and label-label. Finally, graph convolution is applied to graph-level classification. It calculates the cosine similarity between label embeddings to capture the relationship between labels, and solves the multi-label classification problem as a link prediction task. It has good effects in semi-supervised learning and alleviates the impact of data imbalance. And since the mark-label relationship is shown in the graph, it is easy to identify the tokens that trigger specific categories, thus providing good interpretability for multi-label classification.
[0010] However, there are the following drawbacks in the prior art two: Although it fuses label information by constructing the graph structure of tokens and labels and applying the graph convolutional network for graph-level classification. However, due to the complexity of the graph structure, its support for classification tasks with a large number of labels and a large amount of text is not ideal. The rapid decrease in efficiency makes its training and use quite difficult. Secondly, in this method, the expression of the same label information is consistent in different samples. It mainly relies on text features to establish the connection with the labels. For different samples, the different distribution of the importance of label information is also a potential information, and this technology fails to fuse this potential information to assist classification. Summary of the Invention
[0011] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a multi-label text classification method and system based on label relationships, which improves the accuracy and relevance of the final classification.
[0012] The present invention adopts the following technical solutions to solve the above technical problems:
[0013] In the first aspect, a multi-label text classification method based on label relationships proposed according to the present invention includes:
[0014] Capturing the text feature P in the text dataset through the pre-trained model, obtaining the initial classification ranking according to the text feature, and obtaining the first label sequence S1 of the predicted label probability;
[0015] Obtain a second tag sequence S2 according to the head tag in the first tag sequence S1;
[0016] Combine S2 with the tag frequency co-occurrence matrix M from a given text dataset to obtain a third tag sequence S3, take the union of S2 and S3 to obtain a fourth tag sequence S4, reorder the tags in S4 through tag frequency distribution information to obtain a frequency-integrated tag sequence S, and generate a tag feature sequence based on S
[0017] Through the attention mechanism Establish a semantic relationship with the text to obtain the final feature f cat ;
[0018] Adopt the final feature f cat Perform final classification.
[0019] As a further optimization scheme of a multi-label text classification method based on tag relationships described in the present invention, capture the text features P in the text dataset through a pre-trained model, obtain an initial classification ranking according to the text features, and obtain a first tag sequence S1 of predicted tag probabilities; including:
[0020] Tokenize and preprocess the initial text to obtain a word sequence T of size n; where, T = {t cls , t1, t2, t3…, t n-2 , t sep}, t cls is the first special token, and the first special token is used to represent the features of the text through a pre-trained model, t a is the a-th non-special token, 1 ≤ a ≤ n - 2, t sep is the second special token, and the second special token is used for the pre-trained model to identify the end of the text;
[0021] Encode T using a pre-trained model, and the output of the pre-trained model is expressed as:
[0022] H = RoBERTa(W RoBERTa , T)
[0023] P = Tanh(h cls W1 + b1)
[0024] where, H is the character-level feature extracted by the pre-trained model, Tanh(*) is the activation function, W RoBERTa is the training parameter of the pre-trained model, P is the text feature, h cls is t clsThe feature representation of the text encoded by the pre-trained model, where W1 is the first training parameter, b1 is the first training bias parameter, and RoBERTa(*) is the RoBERTa model;
[0025] Obtain S1 according to P;
[0026] S1 = sigmoid(PW2 + b2)
[0027] where sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter;
[0028] S1 = {s1, s2, … s K}, where s i is the prediction probability of the i-th label, K is the number of all labels, and 1 ≤ i ≤ K.
[0029] As a further optimization scheme of the multi-label text classification method based on label relationships described in the present invention, obtain the second label sequence S2 according to the head label in the first label sequence S1; including:
[0030] Select the label prediction probabilities in S1 that are greater than the threshold hyperparameter α, and use the labels corresponding to the selected label prediction probabilities as the second label sequence S2.
[0031] As a further optimization scheme of the multi-label text classification method based on label relationships described in the present invention, combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3, take the union of S2 and S3 to obtain the fourth label sequence S4, reorder the labels in S4 through label frequency distribution information to obtain the frequency-integrated label sequence S, and generate a label feature sequence based on S Specifically as follows:
[0032] For a given text dataset with K labels, the co-occurrence frequency matrix M = {m1, m2, …, m K}, where m i is the co-occurrence frequency sequence of all labels corresponding to the i-th label, that is
[0033] m i = {m i,1 , m i,2 , …, m i,K}, m i,j represents the frequency of the j-th label when the i-th label exists, and 1 ≤ j ≤ K;
[0034] The second label sequence S2 includes V labels, S2 = {l1, l2, … l V}, l vFor the v-th label, where 1 ≤ v ≤ V; the set of co-occurrence frequency sequences of labels is the co-occurrence frequency sequence of all labels corresponding to the v-th label;
[0035] The sum of the set of co-occurrence frequency sequences of labels is obtained
[0036]
[0037] Among them, is the label co-occurrence frequency distribution of the second label sequence S2 for K labels, is the frequency of co-occurrence of the k-th label when the v-th label exists;
[0038] According to the descending order of the label co-occurrence frequencies of K labels from select the top γ - V labels that do not belong to S2 to form the third label sequence S3, where γ is a hyperparameter used to control the expected length of the third label sequence S3;
[0039] According to S2 and S3, calculate the fourth label sequence S4:
[0040] S4 = S2 ∪ S3
[0041] Sort S4 according to the label frequency distribution information of the entire given text dataset, and the sorting order is from high to low frequency of labels, to obtain the frequency-integrated label sequence S;
[0042] Design a label position feature matrix for training and a label feature matrix Among them, δ is the size of the hidden layer in the pre-trained model, is a representation matrix with γ rows and δ columns over the real number field, is a representation matrix with K rows and δ columns over the real number field;
[0043] Use S to select the label features corresponding to the labels in S from M f and add them to M pos to obtain the feature representation F that integrates label position information;
[0044] Integrate the feature sequences of the position information and label information brought by and in F, which is represented as follows
[0045]
[0046] Among them, is obtained by further fusing the information from and The label feature sequence integrating position information and label information, ReLU(*) is the activation function, Drop(*) is the function that randomly makes some parameters not participate in training, W3 is the third training parameter, b3 is the third training bias parameter, W4 is the fourth training parameter, and b4 is the fourth training bias parameter.
[0047] As a further optimization scheme of the multi-label text classification method based on label relationships described in the present invention,
[0048] S = Rank(S4)
[0049] where Rank(*) sorts S4 according to the frequency of the label appearing in the given text dataset, and the sorting order is from high to low frequency of the label;
[0050] ReLU(*) is the activation function, i.e., the rectified linear unit function;
[0051] The label frequency distribution information is the frequency of the label appearing in the given text dataset.
[0052] As a further optimization scheme of the multi-label text classification method based on label relationships described in the present invention, through the attention mechanism, To establish a semantic relationship with the text to obtain the final feature f cat ; including:
[0053] Utilize the label feature sequence Perform masked attention learning on H to extract features including extended label information where, is all individual features in in are concatenated to obtain the final feature f cat ;
[0054]
[0055] where Concat(*) is the feature concatenation operation;
[0056] Adopt the final feature f cat For final classification; including:
[0057] f cat Is mapped to the corresponding probabilities of K labels through linear transformation and the sigmoid function;
[0058] Y = sigmoid(f cat W5 + b5)
[0059] where Y is the final label prediction result, W5 is the fifth training parameter, and b5 is the fifth training bias parameter;
[0060] Implement multi-label text classification according to Y.
[0061] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, the pre-trained model is the RoBERTa model, and sigmoid(*) is the Logistic function;
[0062] Among them, the final loss L of the RoBERTa model is:
[0063] L = β × L1 + (1 - β) × L2
[0064] Among them, β is a hyperparameter used to balance the influence of the prediction loss L1 and the final classification loss L2 on Y;
[0065] The binary cross-entropy loss function is used to calculate L1, and the formula is as follows:
[0066]
[0067] Among them, is the true value of the i-th label, and s i is the predicted probability value of the i-th label;
[0068] For Y = {y1, y2,... y K}, the binary cross-entropy loss function is used as the loss function for the second prediction, and L2 is:
[0069]
[0070] Among them, is the true value of the i-th label, and y i is the predicted probability value of the i-th label.
[0071] In a second aspect, an embodiment of the present invention further provides a multi-label text classification system based on label relationship, including:
[0072] A first prediction module, configured to capture text features P in a text data set through a pre-trained model, obtain an initial classification ranking according to the text features, and obtain a first label sequence S1 of predicted label probabilities;
[0073] A label feature sequence generation module, configured to obtain a second label sequence S2 according to the head label in the first label sequence S1; combine S2 with a label frequency co-occurrence matrix M from a given text data set to obtain a third label sequence S3, obtain a fourth label sequence S4 by taking the union of S2 and S3, reorder the labels in S4 through label frequency distribution information to obtain a frequency-integrated label sequence S, and generate a label feature sequence based on S
[0074] A final feature generation module, configured to obtain a final feature f by establishing a semantic relationship with the text through an attention mechanism ; cat
[0075] A classification module, configured to perform final classification by using the final feature f cat
[0076] In a third aspect, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the multi-label text classification method based on label relationships according to the first aspect or any corresponding implementation manner thereof are implemented.
[0077] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multi-label text classification method based on label relationships according to the first aspect or any corresponding implementation manner thereof are implemented.
[0078] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0079] (1) Through a two-stage re-ranking method, the present invention organically combines the neglected label co-occurrence information and label frequency distribution information. In a relatively efficient manner, it not only strengthens the feature representation, but also flexibly integrates the label-label and label-text relationships, and assigns different label importances to different samples to assist in prediction; the present invention can effectively alleviate the long-tail problem and improve the accuracy and relevance of the final classification;
[0080] (2) The LabelCoRank of the present invention can effectively learn and distinguish relevant and irrelevant label information without introducing too much noise from additional labels, and has strong robustness even in tasks with fewer labels.
[0081] (3) The present invention has good adaptability to both datasets with a small number of labels and datasets with a large number of labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a structural diagram of the LabelCoRank model.
[0083] Figure 2 is a schematic flowchart of the present invention. DETAILED DESCRIPTION
[0084] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0085] The present invention proposes LabelCoRank, a new method inspired by ranking principles. LabelCoRank utilizes label co-occurrence relationships to improve initial label classification through a two-stage re-ranking process. In the first stage, the initial classification results are used to form a preliminary ranking. In the second stage, the preliminary results are re-ranked using the label co-occurrence frequency matrix to enhance the accuracy and relevance of the final classification. The model also incorporates the frequency distribution of labels, enabling it to assign different importance to labels based on their occurrences in the dataset. Through label embedding and attention mechanisms, LabelCoRank establishes semantic relationships between labels and text features. This two-stage approach ensures that even uncommon labels, which are usually underrepresented, receive sufficient attention during the classification process.
[0086] The LabelCoRank model consists of three main modules. The first module uses the text features captured by a pre-trained model to obtain an initial classification ranking. In the second module, label reordering is performed by leveraging the head labels in the initial classification and combining them with the label frequency co-occurrence matrix and label frequency distribution information from the dataset. This generates a sequence of label features containing various additional information, which is then used to establish semantic relationships with the text through an attention mechanism. The third module uses these features for the final classification. Figure 1 Illustrates the overall architecture of the LabelCoRank model. Figure 2 Illustrates the flow schematic of the method.
[0087] Initially, the RoBERTa model is used for feature extraction and initial label prediction based on its text features. The initial text is tokenized and preprocessed to obtain a sequence of words T of size n; where, T = {t cls , t1, t2, t3…, t n-2 , t sep}. t cls is the first special token, which is used to represent the features of the text through the pre-trained model, t a is the a-th non-special token, 1 ≤ a ≤ n - 2, and t sep is the second special token, which is used by the pre-trained model to identify the end of the text;
[0088] The pre-trained model is used to encode T, and the output of the pre-trained model is expressed as:
[0089] H = RoBERTa(W RoBERTa , T)
[0090] P = Tanh(h cls W1 + b1)
[0091] Among them, H is the character-level feature extracted by the pre-trained model, Tanh(*) is the activation function, W RoBERTa are the training parameters of the pre-trained model, P is the text feature, h cls is the feature representation of the text encoded by the pre-trained model, W1 is the first training parameter, and b1 is the first training bias parameter; cls According to P, obtain S1;
[0092] S1 = sigmoid(PW2 + b2)
[0093] where sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter;
[0094] S1 = {s1, s2, … s
[0095] }, where s K is the prediction probability of the i-th label, K is the number of all labels, and 1 ≤ i ≤ K. i For loss calculation, use the binary cross-entropy loss function, and its formula is as follows:
[0096] where K is the number of all labels,
[0097]
[0098] is the true value of the i-th label, and s is the predicted probability value of the i-th label. The prediction loss of this instance is denoted as L1, which will be combined with the final classification loss L2 as the final loss L. i After the preliminary prediction, obtain the first label sequence S1 of the predicted label probabilities. Use the threshold hyperparameter α to select the part of the label prediction probabilities in S1 that are greater than α, and use the corresponding labels as the second label sequence S2. The purpose is to obtain the labels more relevant to the text predicted by RoBERTa, thereby reducing the introduction of irrelevant noise. For the co-occurrence frequency matrix M, its content is defined as follows: For a text dataset with K labels, m
[0099] represents the frequency of the j-th label when the i-th label exists, where 1 ≤ j ≤ K. m i,j is the co-occurrence frequency sequence of all labels corresponding to the i-th label, that is, m i = {m i , m i,1 , …, m i,2}. Therefore, for the second label sequence S2 with V labels, S2 = {l1, l2, … l i,K}, and the set of label co-occurrence frequency sequences V Summing the tag frequency sequences gives which is defined as follows:
[0100]
[0101] where is the tag co - occurrence frequency distribution of the second tag sequence S2 for K tags, is the frequency of co - occurrence of the k - th tag when the v - th tag is present.
[0102] The hyperparameter γ is used to control the expected length of the tag sequence.
[0103] Retrieve the corresponding tags from and select the top γ - V tags that do not belong to S2 according to the frequency in descending order to form the third tag sequence S3. S2 ∪ S3 gives the fourth tag sequence S4 that is more relevant to each sample:
[0104] S4 = S2 ∪ S3
[0105] This is the first Label Reranking. In terms of effect, using the co - occurrence frequency matrix can expand the samples to more relevant tags, thus obtaining more relevant information. But from another perspective, through the co - occurrence frequency matrix, the important information that may be ignored in the initial tag probability sequence can be reordered according to the co - occurrence information and the priorities can be redetermined. Then, sort S4 according to the tag frequency distribution of the entire dataset, placing the high - frequency tags at the beginning of the sequence and the low - frequency tags at the end of the sequence. In this way, the frequency - integrated tag sequence S
[0106] S = Rank(S4)
[0107] Design a tag position feature matrix for training and a tag feature matrix
[0108] where δ is the size of the RoBERTa hidden layer, is a representation matrix with γ rows and δ columns over the real number field, is a representation matrix with K rows and δ columns over the real number field;. Select the corresponding tag features from M f using S and add them to M pos to obtain the feature representation F that integrates tag position information. Then, further integrate the position information and tag information in F using a feed - forward neural network module, which is expressed as follows:
[0109]
[0110] where is obtained by further fusing with The label feature sequence integrating position information and label information, ReLU(*) is the activation function, Drop(*) is the function that randomly makes some parameters not participate in training, W3 is the third training parameter, b3 is the third training bias parameter, W g is the fourth training parameter, and b4 is the fourth training bias parameter.
[0111] This is the second Label Reranking. By sorting the labels according to their distribution in the dataset and combining the label frequency distribution information, all relevant labels are reordered, thus effectively organizing the semantically highly relevant head labels initially predicted by RoBERTa and the highly relevant labels obtained through co-occurrence relationships. Sort the sequence by frequency to obtain an ordered label sequence S containing frequency information. Map this label sequence to label features, fuse the position information, and further fit it through a linear layer to obtain the fused label features so that the same label has different feature representations at different positions in the sequence. Then, use this label feature sequence to perform masked attention learning on the character-level features extracted by RoBERTa to extract features containing extended label information For all individual features in in are concatenated to obtain the final feature f cat。
[0112]
[0113] where Concat(*) is the feature concatenation operation;
[0114] These features are mapped to the corresponding probabilities of the K labels through linear transformation and the sigmoid function.
[0115] Y = sigmoid(f cat W5 + b5)
[0116] For Y = {y1, y2, … y K}, the binary cross-entropy loss function is used as the loss function for the second prediction, and the result is the loss L2. The final loss is defined as:
[0117]
[0118] L = β × L1 + (1 - β) × L2
[0119] where K is the number of all labels, is the true value of the i-th label, and y i is the predicted probability value of the i-th label, and β is the hyperparameter used to balance the influence of the two losses on the final result.
[0120] RoBERTa (*) is an open-source general pre-trained model, and so is BERT. Here, RoBERTa can be replaced by other open-source pre-trained models trained with data from various fields using the BERT or RoBERTa model structure.
[0121] The present invention also discloses a multi-label text classification system based on label relationships, including:
[0122] A first prediction module for capturing text features P in a text dataset through a pre-trained model, obtaining an initial classification ranking based on the text features, and obtaining a first label sequence S1 of predicted label probabilities;
[0123] A label feature sequence generation module for obtaining a second label sequence S2 according to the head labels in the first label sequence S1; combining S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain a third label sequence S3, taking the union of S2 and S3 to obtain a fourth label sequence S4, reordering the labels in S4 through label frequency distribution information to obtain a frequency-integrated label sequence S, and generating a label feature sequence based on S
[0124] A final feature generation module for obtaining a final feature f by establishing a semantic relationship with the text through an attention mechanism cat ;
[0125] A classification module for performing final classification using the final feature f cat
[0126] The technical solution proposed by the present invention has been evaluated on three publicly available datasets. The following are the statistics of the three datasets:
[0127] MAG-CS: This dataset consists of 705,407 papers from the Microsoft Academic Graph (MAG), which were selected from 105 well-known CS conferences held between 1990 and 2020. It contains 15,808 unique labels, providing a comprehensive collection of scientific literature in this field.
[0128] PubMed: This dataset contains 898,546 papers from PubMed, representing 150 leading medical journals between 2010 and 2020. It contains 17,963 labels corresponding to MeSH terms, providing valuable insights for biomedical research.
[0129] AAPD: This dataset contains the English abstracts of computer science papers from arxiv.org, and each abstract is paired with a relevant topic. It contains a total of 55,840 abstracts covering various related disciplines.
[0130] Table 1 shows the dataset information, where N trn and N tst represent the number of documents in the training set and the test set respectively. D represents the vocabulary size of all documents. L n represents the number of labels, and L avg represents the average number of labels per document, and W avg represents the average number of words per document.
[0131] Table 1. Dataset Information Table
[0132]
[0133] Two ranking-based evaluation metrics are used in the experiment: top-K label accuracy (P@K) and top-K label normalized discounted cumulative gain (NDCG@K).
[0134] Nine methods are used for comparison in the experiment:
[0135] XML-CNN uses a dynamic max pooling scheme to capture richer information from different regions of the document, adopts a binary cross-entropy loss function to handle multi-label problems, and introduces a hidden bottleneck layer to obtain better document representations and reduce the model size.
[0136] MeSHProbeNet is an end-to-end deep learning model designed for MeSH indexing, which assigns MeSH terms to MEDLINE citations. It won the first place in the latest Batch A of the 2018 BioASQ Challenge.
[0137] AttentionXML is a label-tree-based deep learning model designed specifically for extreme multi-label text classification. It introduces two key features: a multi-label attention mechanism that can capture the relevant text parts of each label, and a shallow and wide probabilistic label tree that can effectively handle millions of labels.
[0138] Transformer is a network architecture and is the first sequence transduction model entirely based on attention mechanisms. It replaces the commonly used recurrent layers with multi-head self-attention mechanisms and abandons the common recurrent and CNN architectures.
[0139] Star-Transformer is a lightweight alternative to Transformer for NLP tasks. It uses a star topology to reduce the complexity from quadratic to linear and solves the problem of high computational requirements.
[0140] BertXML is a customized version of BERT designed specifically for XMTC. It overcomes the limitation of a single [CLS] token in BERT by merging multiple [CLS] tokens at the beginning of each input sequence.
[0141] MATCH learns improved text and metadata representations by jointly embedding text and metadata into the same space, enabling higher-order interactions between words and metadata.
[0142] LiGCN introduces an interpretable graph convolutional network model that models tokens and labels as nodes in a heterogeneous graph. It calculates the cosine similarity between label embeddings to capture the relationships between labels.
[0143] GUDN utilizes label semantics and a deep pre-trained model, combined with a label reinforcement strategy for fine-tuning to improve classification performance. The model shows sensitivity to label semantics and demonstrates significant efficacy on datasets with semantically rich labels.
[0144] Tables 2, 3, and 4 summarize the results of different models on the MAG-CS, PubMed, and AAPD datasets, respectively.
[0145] In the MAG-CS dataset (Table 2), the model outperforms all baselines on almost all metrics, except for P@1, where it lags behind MATCH by 0.0072. However, for P@3, P@5, NDCG@3, and NDCG@5, the model improves over MATCH by 0.0047, 0.0131, 0.0007, and 0.0079, respectively. This indicates an enhancement in the prediction performance of tail labels. This improvement is attributed to the introduction of a large amount of relevant label information, which improves the prediction of tail labels but slightly weakens the attention to head labels. Additionally, MATCH uses additional meta-information not available in LabelCoRank, which can explain the difference in P@1.
[0146] In the PubMed dataset (Table 3), LabelCoRank achieves the best performance among all metrics.
[0147] Notably, compared to the MAG dataset, LabelCoRank performs significantly better than MATCH on the PubMed dataset. The improvements in P@3, P@5, NDCG@3, and NDCG@5 are all over 2.8 percentage points, indicating significant progress in predicting difficult-to-predict tail labels. The PubMed dataset has the highest average number of labels per instance, which may enhance the relevance of supplementary label information, leading to superior performance improvement.
[0148] In the AAPD dataset (Table 4), LabelCoRank also achieved the best performance in all metrics. Although the average number of labels per document is only 2.41, the total number of labels is 54, allowing most label features to participate in the model's calculation. LabelCoRank can effectively learn and distinguish relevant and irrelevant label information without introducing too much noise from additional labels. This demonstrates the strong robustness of the model even in tasks with fewer labels.
[0149] Table 2. Experimental Results Table of MAG-CS Dataset
[0150]
[0151] Table 3. Experimental Results Table of PubMed Dataset
[0152]
[0153] Table 4. Experimental Results Table of AAPD Dataset
[0154]
[0155] Through the results of ablation experiments (Tables 5 - 10), the effectiveness of the module used to integrate label information in LabelCoRank was verified. Four design elements in LabelCoRank need to have their effectiveness verified: label selection (correlation matrix), label sequence ranking, integration of position information, and selection of the number of labels.
[0156] Label selection directly affects which information in the text will be extracted under the attention mechanism, thus affecting the final label prediction. To verify the effectiveness of using the frequency correlation matrix for label selection, it was compared with directly selecting the same number of labels from the predictions of RoBERTa without using the frequency correlation matrix. The experimental results in Tables 5, 6, and 7 show that using the frequency correlation matrix for label selection is more effective than directly using the predicted labels of RoBERTa. This is understandable because the correlation of the predicted labels gradually decreases. In addition, the predicted labels of RoBERTa do not consider co-occurrence relationships outside of text semantics. For some labels, these connections are more hidden and do not directly reflect in the semantic content of the text.
[0157] The label sequences are sorted according to the frequencies of the labels in the entire dataset, aiming to introduce the frequency distribution information of the labels on the dataset. The features are extended through the multi-label attention mechanism, which naturally places the features of high-frequency labels at the front, enabling the classifier to naturally pay different degrees of attention to labels at different positions. The results in Tables 5, 6, and 7 show its effectiveness.
[0158] Such label sequences containing frequency distribution information are regarded as a sequence with priorities. For the same label, its meaning varies according to its position; the more forward it is, the more attention it should receive, and conversely, the more backward it is, the less attention it should receive. Therefore, integrating the position information into the label sequences enables different feature expressions at different positions, realizing the dynamic representation of label features. The results in Tables 5, 6, and 7 show its effectiveness.
[0159] Different selections of the number of labels will introduce different degrees of relevant label information and noise, having different impacts on different datasets and samples. The influence of the selection of the number of labels was studied through experiments and analysis. The results show that in the MAG dataset, 35 labels are optimal (Table 8); in the Mesh dataset, 30 labels are appropriate (Table 9); in the AAPD dataset, 20 labels are appropriate (Table 10).
[0160] Table 5. Results of ablation experiments on the MAG-CS dataset
[0161]
[0162] Table 6. Results of ablation experiments on the PubMed dataset
[0163]
[0164] Table 7. Results of ablation experiments on the AAPD dataset
[0165]
[0166] Table 8. Results of selecting different numbers of labels on the MAG-CS dataset
[0167]
[0168] Table 9. Results of selecting different numbers of labels on the PubMed dataset
[0169]
[0170] Table 10. Results of selecting different numbers of labels on the AAPD dataset
[0171]
[0172] The present invention proposes a novel general multi-label text classification framework, which can effectively alleviate the long-tail problem. The core lies in that, through a two-stage re-ranking method, the neglected label co-occurrence information and label frequency distribution information are organically combined. It strengthens the feature representation in a relatively efficient manner, and also flexibly integrates the label-label and label-text relationships, and assigns different label importance to different samples to assist prediction. And this technology has good adaptability to both datasets with a small number of labels and those with a large number of labels.
[0173] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the steps of the multi-label text classification method based on label relationships as described in the above first aspect or any corresponding implementation manner thereof.
[0174] An embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-label text classification method based on label relationships as described in the above first aspect or any corresponding implementation manner thereof.
[0175] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0176] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0179] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0180] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-label text classification method based on label relationship, characterized in that: include: Capture the text features P in the text dataset through the pre-training model, obtain the initial classification ranking based on the text features, and obtain the first label sequence S1 of the predicted label probability; According to the header tag in the first tag sequence S1, a second tag sequence S2 is obtained; Combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3. Take the union of S2 and S3 to obtain the fourth label sequence S4. Reorder the labels in S4 according to the label frequency distribution information to obtain the frequency-integrated label sequence S. Generate a label feature sequence based on S Through the attention mechanism Establish a semantic relationship with the text to obtain the final feature f cat ; Take the final feature f cat Perform final classification.
2. According to the multi-label text classification method based on label relationship of claim 1, it is characterized in that: Capture the text features P in the text dataset through the pre-training model, obtain the initial classification ranking based on the text features, and obtain the first label sequence S1 of the predicted label probability; include: The initial text is tokenized and preprocessed to obtain a word sequence T of size n; where T = {t cls , t1, t2, t3…, t n-2 , t sep }, t cls is the first special marker word, which is used to represent the characteristics of the text through the pre-trained model. a is the ath non-special tag word, 1≤a≤n-2, t sep is the second special marker word, which is used by the pre-trained model to recognize as the end of the text; The pre-trained model is used to encode T, and the output of the pre-trained model is expressed as: H=RoBERTa(W RoBERTa ,T) P=Fish(h) cls W1+b1) Among them, H is the character-level feature extracted by the pre-training model, Tanh(*) is the activation function, and W ROBERTa is the training parameter of the pre-trained model, P is the text feature, h cls Yes cls Feature expression of text encoded by the pre-trained model, W1 is the first training parameter, b1 is the first training bias parameter, and RoBERTa(*) is the RoBERTa model; Obtain S1 according to P; S1=sigmoid(PW2+b2) Among them, sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter; S1={s1,s2,...s K }, where s i is the predicted probability of the i-th label, K is the number of all labels, 1≤i≤K.
3. The multi-label text classification method based on label relationship according to claim 1 is characterized in that: According to the header tag in the first tag sequence S1, a second tag sequence S2 is obtained; include: Select the label prediction probability in S1 that is greater than the threshold hyperparameter α, and use the label corresponding to the selected label prediction probability as the second label sequence S2.
4. The multi-label text classification method based on label relationship according to claim 2 is characterized in that: Combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3. Take the union of S2 and S3 to obtain the fourth label sequence S4. Reorder the labels in S4 according to the label frequency distribution information to obtain the frequency-integrated label sequence S. Generate a label feature sequence based on S The details are as follows: For a given text dataset with K labels, the co-occurrence frequency matrix M = {m1, m2, ..., m K }, where m i is the co-occurrence frequency sequence of all tags corresponding to the i-th tag, that is, m i ={m i,1 , m i,2 , ..., m i,K },m i,j Indicates the frequency of the jth label when the ith label exists, 1≤j≤K; The second tag sequence S2 includes V tags, S2 = {l1, l2, ...l V }, l v is the vth label, 1≤v≤V; label co-occurrence frequency sequence set is the co-occurrence frequency sequence of all tags corresponding to the vth tag; Sum the tag co-occurrence frequency sequence set to get in, is the label co-occurrence frequency distribution of the second label sequence S2 for K labels, is the frequency of co-occurrence of the kth tag when the vth tag exists; According to the co-occurrence frequency of the K tags, Select the first γ-V tags that do not belong to S2 to form a third tag sequence S3, where γ is a hyperparameter for controlling the expected length of the third tag sequence S3; According to S2 and S3, the fourth tag sequence S4 is calculated: S4=S2∪S3 Sort S4 according to the label frequency distribution information of the entire given text data set, and the order of sorting is from high to low label frequency, and obtain the frequency-integrated label sequence S; Design a label position feature matrix for training and a label feature matrix Among them, δ is the size of the hidden layer in the pre-trained model, To represent the matrix as γ rows and δ columns over the real number field, To represent the matrix as K rows and δ columns in the real number field; Using S from M f Select the label features corresponding to the labels in S and add them to M pos In the above example, we obtain the feature representation F that integrates the label position information; Integrated F and The feature sequence of location information and label information is expressed as follows in, For further integration by and The label feature sequence of the integrated position information and label information, ReLU(*) is the activation function, Drop(*) is a function that randomly makes some parameters not participate in training, W3 is the third training parameter, b3 is the third training bias parameter, W4 is the fourth training parameter, and b4 is the fourth training bias parameter.
5. The multi-label text classification method based on label relationship according to claim 4 is characterized in that: S=Rank(S4) Among them, Rank(*) is to sort S4 by the frequency of the label in a given text dataset, and the order of sorting is from high to low frequency of the label; ReLU(*) is the activation function, i.e., the linear rectification function; The label frequency distribution information is the frequency with which the label appears in a given text dataset.
6. The multi-label text classification method based on label relationship according to claim 5 is characterized in that: Through the attention mechanism To establish a semantic relationship with the text, we obtain the final feature f cat ; include: Using label feature sequence Perform masked attention learning on H to extract features including extended label information in, for All individual features in In are concatenated to obtain the final feature f cat ; Among them, Concat(*) is a feature concatenation operation; Take the final feature f cat Conduct final classification; including: f cat Mapped to the corresponding probabilities of K labels through linear transformation and sigmoid function; Y=sigmoid(f cat W5+b5) Among them, Y is the final label prediction result, W5 is the fifth training parameter, and b5 is the fifth training bias parameter; Implement multi-label text classification based on Y.
7. The multi-label text classification method based on label relationship according to claim 6 is characterized in that: The pre-trained model is the RoBERTa model, and sigmoid(*) is the Logistic function; Among them, the final loss L of the RoBERTa model is: L=β×L1+(1-β)×L2 Among them, β is a hyperparameter used to balance the impact of prediction loss L1 and final classification loss L2 on Y; The binary cross entropy loss function is used to calculate L1, and the formula is as follows: in, is the true value of the i-th label, s i is the predicted probability value of the i-th label; For Y = {y1, y2, ...y K The binary cross entropy loss function of} is used as the loss function for the second prediction, and L2 is: in, is the true value of the i-th label, and yi is the predicted probability value of the i-th label.
8. A multi-label text classification system based on label relations, characterized in that: include: A first prediction module is used to capture text features P in a text dataset through a pre-trained model, obtain an initial classification ranking based on the text features, and obtain a first label sequence S1 for predicting label probabilities; The tag feature sequence generation module is used to obtain the second tag sequence S2 according to the head tag in the first tag sequence S1; combine S2 with the tag frequency co-occurrence matrix M from a given text data set to obtain the third tag sequence S3, obtain the fourth tag sequence S4 by taking the union of S2 and S3, reorder the tags in S4 according to the tag frequency distribution information, obtain the frequency-integrated tag sequence S, and generate a tag feature sequence based on S The final feature generation module is used to convert Establish a semantic relationship with the text to obtain the final feature f cat ; Classification module, used to adopt the final feature f cat Perform final classification.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the multi-label text classification method based on label relations as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multi-label text classification method based on label relations as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Public opinion text classification method and system based on multi-label embedding, terminal and medium
CN113987187A
Label processing method and device, electronic equipment and computer readable storage medium
CN114970548A