Multi-label text classification method and system based on label relationship
By adopting a multi-stage re-ranking method based on label relationships in multi-label text classification, combining label co-occurrence information and frequency distribution information, and establishing semantic relationships between text and labels through attention mechanisms, the problem of failure to effectively capture label relationships in the existing technology is solved, and the accuracy of classification and the ability to predict tail labels are improved.
Patent Information
- Application Number
- CN202510020631.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The prior art is difficult to effectively capture the relationship between labels and between labels and text in multi-label text classification, especially in the long-tail problem and tail label prediction.
Through a multi-stage re-ranking method based on label relationships, combining tag co-occurrence information and tag frequency distribution information, a frequency-integrated tag sequence is generated, and a semantic relationship between text and tags is established through attention mechanisms, and finally classified.
It improves the accuracy and relevance of multi-label text classification, effectively alleviates the long-tail problem, and can better identify and predict tail labels.
Smart Images

Figure CN119938924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a multi-label text classification method and system based on label relationship. Background Art
[0002] Multi-label text classification is a key task in natural language processing. Its goal is to assign appropriate labels to a given text. It also has wide applications in other fields such as recommender systems, sentiment analysis, rumor detection, question answering tasks, etc.
[0003] Recent advances in deep learning have driven the development of models based on convolutional neural networks and recurrent neural networks, which have demonstrated excellent feature extraction capabilities. However, these models face challenges in capturing the contextual relationships of words in long texts. For example, while convolutional neural networks can effectively capture local patterns, they struggle to handle long-distance dependencies. Although recurrent neural networks are better at processing sequential data, they often suffer from problems such as gradient vanishing, making it difficult to effectively learn long-term dependencies. The introduction of attention mechanisms has alleviated some of these problems, and models have combined attention mechanisms to improve the semantic representation and classification performance of text. Transformer-based models and powerful pre-trained models have further enhanced text feature extraction capabilities. These models use self-attention mechanisms to capture the relationship between words in a sequence, which can better handle long-distance dependencies and improve the overall semantic understanding of the text.
[0004] Despite these advances, existing technologies have mainly focused on extracting text features, basically following the pattern of extracting text features through neural networks and then using text features for classification through linear layers. Improvements have also focused on strengthening text feature extraction, while often ignoring the relationship between labels and between labels and text. In some cases, text features alone are not enough to predict the correct label, especially when text features cannot fully represent the label semantics. At this time, label relationships can help predict these labels. For tail labels where it is difficult to establish a feature-label connection due to a lack of training samples, additional label information is even more helpful for their prediction. For example, in medical document classification, the presence of a disease label may mean the possibility of another disease occurring at the same time, and many existing models fail to fully capture these relationships.
[0005] The long-tail problem in multi-label text classification is a critical and challenging task. The long-tail phenomenon refers to a large number of labels with very few instances, which poses difficulties in accurately classifying labels with lower frequencies. In some cases, text features alone are not enough to predict the correct label, especially when the text features cannot fully represent the semantic meaning of the label. This problem is particularly evident for tail labels, which lack sufficient training samples and are therefore more difficult to associate with text features. For example, when classifying scientific papers based on their abstracts, the short text of the abstract may not fully convey the semantics of all labels, or the text may be expressed in a more obscure and ambiguous way.
[0006] In this case, it is necessary to exploit the relationship between labels to identify tail labels that cannot be directly seen from the text semantics. Although some methods try to incorporate label information into the classification process. For example, some use attention mechanisms to explore the relationship between label semantics and text, however, they do not consider the relationship between labels. Other researchers use graph networks to model labels and explore the relationship between labels. However, in these studies, the expression of the same label information is consistent in different samples, ignoring the fact that different label information should have different importance for different texts.
[0007] In the prior art 1, it is proposed to enhance the adaptability of the pre-trained model to a large number of label classifications by using multiple special tags [CLS]. For classification tasks with a large number of labels, the feature size may be insufficient. For example, the typical embedding dimension of the model is 256, then BERT will also output a 256-dimensional hidden state for token [CLS]. Assuming there are 10,000 different labels, the information contained in the 256-dimensional feature vector may not be sufficient to fully support the 10,000-dimensional label prediction. Increasing the feature dimension will greatly increase the model size and reduce the model efficiency because the embedding, hidden state and intermediate self-attention results will also become larger. Therefore, the prior art 1 uses multiple special tags to adapt to classification tasks with a larger number of labels. The multi-head self-attention mechanism of BERT allows each special tag to traverse the entire input sequence, and the final hidden state of each special tag can be regarded as a specific feature vector. Specifically, X symbols [CLS_1], [CLS_2], ..., [CLS_X] for classification can be added in front of each input sequence, and the connection of their final hidden states can be used as an aggregate sequence representation. The cost of using multiple [CLS] tags to obtain a wider feature vector is much smaller than increasing the feature dimension. This feature vector is finally used for classification.
[0008] However, the first existing technology has the following disadvantages: This technology mainly enhances the feature extraction capability of the pre-trained model by adding special CLS tags, and obtains more comprehensive feature information to assist classification in a way that has less impact on model efficiency. However, this method is still limited to feature extraction, and some labels are not directly reflected in the semantics, especially for tail labels, because the lack of training samples makes the model recognize the association between features and labels, and enhancing feature extraction alone has little effect. And this technology does not integrate label relationship information other than features to assist classification, which limits the classification ability of tail labels.
[0009] In the second prior art, an interpretable graph convolutional network model is proposed, which first uses a pre-trained model to extract token features, then models token and label embeddings as nodes in a heterogeneous graph, and constructs label-label, label-label, and label-label edges, and finally applies graph convolution to graph-level classification. It calculates the cosine similarity between label embeddings to capture the relationship between labels, and solves the multi-label classification problem as a link prediction task. It works well in semi-supervised learning and mitigates the impact of data imbalance. And because the label-label relationship is shown in the graph, it is easy to identify the label that triggers a specific category, providing good interpretability for multi-label classification.
[0010] However, the second prior art has the following shortcomings: Although it integrates label information by constructing a graph structure of tokens and labels and using a graph convolutional network for graph-level classification. However, due to the complexity of the graph structure, its support for classification tasks with a large amount of labels and text is not ideal. The rapid reduction in efficiency makes its training and use quite difficult. Secondly, in this method, the expression of the same label information is consistent in different samples, and it relies more on text features to establish a connection with the label. For different samples, the different distribution of the importance of label information is also a kind of potential information. This technology fails to integrate this potential information to help classification. Summary of the invention
[0011] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a multi-label text classification method and system based on label relationships, and the present invention improves the accuracy and relevance of the final classification.
[0012] The present invention adopts the following technical solutions to solve the above technical problems:
[0013] In a first aspect, a multi-label text classification method based on label relationship proposed in the present invention includes:
[0014] Capture the text features P in the text dataset through the pre-training model, obtain the initial classification ranking based on the text features, and obtain the first label sequence S1 of the predicted label probability;
[0015] According to the header tag in the first tag sequence S1, a second tag sequence S2 is obtained;
[0016] Combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3. Take the union of S2 and S3 to obtain the fourth label sequence S4. Reorder the labels in S4 according to the label frequency distribution information to obtain the frequency-integrated label sequence S. Generate a label feature sequence based on S
[0017] Through the attention mechanism Establish a semantic relationship with the text to obtain the final feature f cat ;
[0018] Take the final feature f cat Perform final classification.
[0019] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, the text features P in the text data set are captured by a pre-trained model, an initial classification ranking is obtained according to the text features, and a first label sequence S1 for predicting label probabilities is obtained; comprising:
[0020] The initial text is tokenized and preprocessed to obtain a word sequence T of size n; where T = {t cls ,t1,t2,t3…,t n-2 ,t sep}, t cls is the first special marker word, which is used to represent the characteristics of the text through the pre-trained model. a is the ath non-special tag word, 1≤a≤n-2, t sep is the second special marker word, which is used by the pre-trained model to recognize as the end of the text;
[0021] The pre-trained model is used to encode T, and the output of the pre-trained model is expressed as:
[0022] H = RoBERTa(W RoBERTa ,T)
[0023] P = Tanh(h cls W1+b1)
[0024] Among them, H is the character-level feature extracted by the pre-training model, Tanh(*) is the activation function, and W RoBERTa is the training parameter of the pre-trained model, P is the text feature, h cls Yes clsFeature expression of text encoded by the pre-trained model, W1 is the first training parameter, b1 is the first training bias parameter, and RoBERTa(*) is the RoBERTa model;
[0025] Obtain S1 according to P;
[0026] S1=sigmoid(PW2+b2)
[0027] Among them, sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter;
[0028] S1={s1,s2,…s K}, where s i is the predicted probability of the i-th label, K is the number of all labels, 1≤i≤K.
[0029] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, a second label sequence S2 is obtained according to the head label in the first label sequence S1; comprising:
[0030] Select the label prediction probability in S1 that is greater than the threshold hyperparameter α, and use the label corresponding to the selected label prediction probability as the second label sequence S2.
[0031] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, S2 is combined with the label frequency co-occurrence matrix M from a given text data set to obtain a third label sequence S3, and a fourth label sequence S4 is obtained by taking the union of S2 and S3. The labels in S4 are reordered according to the label frequency distribution information to obtain a frequency-integrated label sequence S, and a label feature sequence is generated based on S The details are as follows:
[0032] For a given text dataset with K labels, the co-occurrence frequency matrix M = {m1,m2,…,m K}, where m i is the co-occurrence frequency sequence of all tags corresponding to the i-th tag, that is,
[0033] m i ={m i,1 ,m i,2 ,…,m i,K},m i,j Indicates the frequency of the jth label when the ith label exists, 1≤j≤K;
[0034] The second tag sequence S2 includes V tags, S2 = {l1, l2, ... l V}, l vis the vth label, 1≤v≤V; label co-occurrence frequency sequence set is the co-occurrence frequency sequence of all tags corresponding to the vth tag;
[0035] Sum the tag co-occurrence frequency sequence set to get
[0036]
[0037] in, is the label co-occurrence frequency distribution of the second label sequence S2 for K labels, is the frequency of co-occurrence of the kth tag when the vth tag exists;
[0038] According to the co-occurrence frequency of the K tags, Select the first γ-V tags that do not belong to S2 to form a third tag sequence S3, where γ is a hyperparameter for controlling the expected length of the third tag sequence S3;
[0039] According to S2 and S3, the fourth tag sequence S4 is calculated:
[0040] S4=S2∪S3
[0041] Sort S4 according to the label frequency distribution information of the entire given text data set, and the order of sorting is from high to low label frequency, and obtain the frequency-integrated label sequence S;
[0042] Design a label position feature matrix for training and a label feature matrix Among them, δ is the size of the hidden layer in the pre-trained model, To represent the matrix as γ rows and δ columns over the real number field, To represent the matrix as K rows and δ columns in the real number field;
[0043] Using S from M f Select the label features corresponding to the labels in S and add them to M pos In the above example, we obtain the feature representation F that integrates the label position information;
[0044] Integrated F and The feature sequence of location information and label information is expressed as follows
[0045]
[0046] in, For further integration by and The label feature sequence of the integrated position information and label information, ReLU(*) is the activation function, Drop(*) is a function that randomly makes some parameters not participate in training, W3 is the third training parameter, b3 is the third training bias parameter, W4 is the fourth training parameter, and b4 is the fourth training bias parameter.
[0047] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention,
[0048] S=Rank(S4)
[0049] Among them, Rank(*) is to sort S4 by the frequency of the label in a given text dataset, and the order of sorting is from high to low frequency of the label;
[0050] ReLU(*) is the activation function, i.e., the linear rectification function;
[0051] The label frequency distribution information is the frequency with which the label appears in a given text dataset.
[0052] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, the attention mechanism is used to To establish a semantic relationship with the text, we obtain the final feature f cat ;include:
[0053] Using label feature sequence Perform masked attention learning on H to extract features including extended label information in, for All individual features in In are concatenated to obtain the final feature f cat ;
[0054]
[0055] Among them, Concat(*) is a feature concatenation operation;
[0056] Take the final feature f cat Conduct final classification; including:
[0057] f cat Mapped to the corresponding probabilities of K labels through linear transformation and sigmoid function;
[0058] Y = sigmoid(f cat W5+b5)
[0059] Among them, Y is the final label prediction result, W5 is the fifth training parameter, and b5 is the fifth training bias parameter;
[0060] Implement multi-label text classification based on Y.
[0061] As a further optimization scheme of the multi-label text classification method based on label relationship described in the present invention, the pre-training model is the RoBERTa model, and sigmoid(*) is the Logistic function;
[0062] Among them, the final loss L of the RoBERTa model is:
[0063] L=β×L1+(1-β)×L2
[0064] Among them, β is a hyperparameter used to balance the impact of prediction loss L1 and final classification loss L2 on Y;
[0065] The binary cross entropy loss function is used to calculate L1, and the formula is as follows:
[0066]
[0067] in, is the true value of the i-th label, s i is the predicted probability value of the i-th label;
[0068] For Y = {y1, y2, ...y K The binary cross entropy loss function of} is used as the loss function for the second prediction, and L2 is:
[0069]
[0070] in, is the true value of the i-th label, y i is the predicted probability value of the i-th label.
[0071] In a second aspect, an embodiment of the present invention further provides a multi-label text classification system based on label relations, comprising:
[0072] A first prediction module is used to capture text features P in a text dataset through a pre-trained model, obtain an initial classification ranking based on the text features, and obtain a first label sequence S1 for predicting label probabilities;
[0073] The tag feature sequence generation module is used to obtain the second tag sequence S2 according to the head tag in the first tag sequence S1; combine S2 with the tag frequency co-occurrence matrix M from a given text data set to obtain the third tag sequence S3, obtain the fourth tag sequence S4 by taking the union of S2 and S3, reorder the tags in S4 according to the tag frequency distribution information, obtain the frequency-integrated tag sequence S, and generate a tag feature sequence based on S
[0074] The final feature generation module is used to convert Establish a semantic relationship with the text to obtain the final feature f cat ;
[0075] Classification module, used to adopt the final feature f cat Perform final classification.
[0076] In a third aspect, an embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein when the processor executes the computer program, the steps of the multi-label text classification method based on label relations as described in the first aspect above or any corresponding embodiment thereof are implemented.
[0077] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multi-label text classification method based on label relationships as described in the first aspect above or any corresponding embodiment thereof are implemented.
[0078] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:
[0079] (1) The present invention organically combines neglected label co-occurrence information and label frequency distribution information through a two-stage re-ranking method. It strengthens feature representation in a relatively efficient way, flexibly integrates label-label and label-text relationships, and assigns different label importances to different samples to help prediction. The present invention can effectively alleviate the long-tail problem and improve the accuracy and relevance of the final classification.
[0080] (2) The LabelCoRank of the present invention can effectively learn and distinguish relevant and irrelevant label information without introducing excessive noise from additional labels, and is highly robust even in tasks with fewer labels.
[0081] (3) The present invention has good adaptability to both data sets with small and large label quantities. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 It is the structural diagram of the LabelCoRank model.
[0083] Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0085] This paper proposes LabelCoRank, a new method inspired by the ranking principle. LabelCoRank uses label co-occurrence relations to improve the initial label classification through a two-stage re-ranking process. The first stage uses the initial classification results to form a preliminary ranking. In the second stage, the preliminary results are re-ranked using the label co-occurrence frequency matrix to improve the accuracy and relevance of the final classification. The model also incorporates the frequency distribution of labels, which enables it to assign different importance to labels based on their occurrence in the dataset. Through label embedding and attention mechanisms, LabelCoRank establishes semantic relationships between labels and text features. This two-stage approach ensures that even uncommon labels, which are usually underrepresented, receive sufficient attention during the classification process.
[0086] The LabelCoRank model consists of three main modules. The first module uses text features captured by the pre-trained model to obtain the initial classification ranking. In the second module, label re-ranking is performed by leveraging the head labels from the initial classification and combining them with the label frequency co-occurrence matrix and label frequency distribution information from the dataset. This produces a sequence of label features containing various additional information, which is then used to establish semantic relationships with the text through an attention mechanism. The third module uses these features for the final classification. Figure 1 The overall architecture of the LabelCoRank model is illustrated. Figure 2 A schematic flow chart illustrating the method.
[0087] Initially, the RoBERTa model is used for feature extraction and initial label prediction based on its text features. The initial text is tokenized and preprocessed to obtain a word sequence T of size n; where T = {t cls ,t1,t2,t3…,t n-2 ,t sep}. cls is the first special marker word, which is used to represent the characteristics of the text through the pre-trained model. a is the ath non-special tag word, 1≤a≤n-2, t sep is the second special marker word, which is used by the pre-trained model to recognize as the end of the text;
[0088] The pre-trained model is used to encode T, and the output of the pre-trained model is expressed as:
[0089] H = RoBERTa(W RoBERTa ,T)
[0090] P = Tanh(h cls W1+b1)
[0091] Among them, H is the character-level feature extracted by the pre-training model, Tanh(*) is the activation function, and W RoBERTa is the training parameter of the pre-trained model, P is the text feature, h cls Yes cls The feature expression of the text encoded by the pre-trained model, W1 is the first training parameter, and b1 is the first training bias parameter;
[0092] Obtain S1 according to P;
[0093] S1=sigmoid(PW2+b2)
[0094] Among them, sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter;
[0095] S1={s1,s2,…s K}, where s i is the predicted probability of the i-th label, K is the number of all labels, 1≤i≤K.
[0096] For loss calculation, the binary cross entropy loss function is used, and its formula is as follows:
[0097]
[0098] Where K is the number of all labels, is the true value of the i-th label, s i is the predicted probability value of the i-th label. The prediction loss of this instance is denoted as L1, which will be combined with the final classification loss L2 as the final loss L.
[0099] After the initial prediction, the first label sequence S1 with predicted label probabilities is obtained. The threshold hyperparameter α is used to select the part of S1 with label prediction probabilities greater than α, and the corresponding labels are used as the second label sequence S2. The purpose is to obtain labels predicted by RoBERTa that are more relevant to the text, thereby reducing the introduction of irrelevant noise. For the co-occurrence frequency matrix M, its content is defined as follows: For a text dataset with K labels, m i,j Indicates the frequency of the jth label appearing when the ith label exists, 1≤j≤K. m i is the co-occurrence frequency sequence of all tags corresponding to the i-th tag, that is, m i ={m i,1 ,m i,2 ,…,m i,K Therefore, for the second tag sequence S2 with V tags, S2 = {l1, l2, ... l V}, label co-occurrence frequency sequence set Summing the label frequency sequence gives The definition is as follows:
[0100]
[0101] in, is the label co-occurrence frequency distribution of the second label sequence S2 for K labels, is the frequency of co-occurrence of the kth tag when the vth tag exists.
[0102] The hyperparameter γ is used to control the expected length of the label sequence.
[0103] from Get the corresponding labels from S2, and select the first γ-V labels that do not belong to S2 in descending order of frequency to form the third label sequence S3. S2∪S3 obtains the fourth label sequence S4 that is more relevant to each sample:
[0104] S4=S2∪S3
[0105] This is the first Label Reranking. Effectively, the use of the co-occurrence frequency matrix allows the sample to be expanded to more relevant labels, thereby obtaining more relevant information. But from another perspective, the co-occurrence frequency matrix can be used to reorder and re-prioritize important information that may be ignored in the initial label probability sequence based on the co-occurrence information. Then, S4 is sorted according to the label frequency distribution of the entire data set, with high-frequency labels placed at the beginning of the sequence and low-frequency labels placed at the end of the sequence. In this way, the frequency-integrated label sequence S is obtained.
[0106] S=Rank(S4)
[0107] Design a label position feature matrix for training and a label feature matrix
[0108] Where δ is the size of the RoBERTa hidden layer, To represent the matrix as γ rows and δ columns over the real number field, To represent the matrix as K rows and δ columns in the real number field;. Using S from M f Select the corresponding label feature and add it to M pos In the above example, we get the feature representation F that integrates the label position information. Then we use the feedforward neural network module to further integrate the position information and label information in F, as shown below:
[0109]
[0110] in, For further integration by and The label feature sequence of the integrated position information and label information, ReLU(*) is the activation function, Drop(*) is a function that randomly excludes some parameters from training, W3 is the third training parameter, b3 is the third training bias parameter, W g is the fourth training parameter, and b4 is the fourth training bias parameter.
[0111] This is the second Label Reranking. By sorting the labels according to their distribution in the dataset and combining the label frequency distribution information, all relevant labels are re-ranked, thereby effectively sorting out the semantically highly relevant head labels initially predicted by RoBERTa and the highly relevant labels obtained through co-occurrence relationships. The sequence is sorted by frequency to obtain an ordered label sequence S containing frequency information. This label sequence is mapped to the label feature, the position information is fused, and further fitted through the linear layer to obtain the fused label feature This allows the same label to have different feature representations at different positions in the sequence. The label feature sequence is then used to perform masked attention learning on the character-level features extracted by RoBERTa to extract features containing extended label information. for All individual features in In are concatenated to obtain the final feature f cat。
[0112]
[0113] Among them, Concat(*) is a feature concatenation operation;
[0114] These features are mapped to the corresponding probabilities of K labels through linear transformation and sigmoid function.
[0115] Y = sigmoid(f cat W5+b5)
[0116] For Y = {y1, y2, ...y K The binary cross entropy loss function of} is used as the loss function for the second prediction, and the result is the loss L2. The final loss is defined as:
[0117]
[0118] L=β×L1+(1-β)×L2
[0119] Where K is the number of all labels, is the true value of the i-th label, y i is the predicted probability value of the i-th label, and β is a hyperparameter used to balance the impact of the two losses on the final result.
[0120] RoBERTa(*) is an open source general pre-trained model, as is BERT. Here, RoBERTa can be replaced by other open source pre-trained models trained with data from different fields using the BERT or RoBERTa model structure.
[0121] The present invention also discloses a multi-label text classification system based on label relationship, comprising:
[0122] A first prediction module is used to capture text features P in a text dataset through a pre-trained model, obtain an initial classification ranking based on the text features, and obtain a first label sequence S1 for predicting label probabilities;
[0123] The tag feature sequence generation module is used to obtain the second tag sequence S2 according to the head tag in the first tag sequence S1; combine S2 with the tag frequency co-occurrence matrix M from a given text data set to obtain the third tag sequence S3, obtain the fourth tag sequence S4 by taking the union of S2 and S3, reorder the tags in S4 according to the tag frequency distribution information, obtain the frequency-integrated tag sequence S, and generate a tag feature sequence based on S
[0124] The final feature generation module is used to convert Establish a semantic relationship with the text to obtain the final feature f cat ;
[0125] Classification module, used to adopt the final feature f cat Perform final classification.
[0126] The technical solution proposed in this invention is evaluated on three publicly available datasets. The following are the statistics of the three datasets:
[0127] MAG-CS: This dataset consists of 705,407 papers from the Microsoft Academic Graph (MAG), selected from 105 prominent CS conferences held between 1990 and 2020. It contains 15,808 unique tags, providing a comprehensive collection of scientific literature in the field.
[0128] PubMed: This dataset contains 898,546 papers from PubMed, representing 150 leading medical journals between 2010 and 2020. It contains 17,963 tags corresponding to MeSH terms, providing valuable insights into biomedical research.
[0129] AAPD: This dataset contains English abstracts of computer science papers from arxiv.org, each of which is paired with a related topic. A total of 55,840 abstracts covering various relevant disciplines are included.
[0130] Table 1 is the dataset information, N trn and N tst Represents the number of documents in the training set and test set respectively. D represents the vocabulary size of all documents. L n Indicates the number of labels, L avg represents the average number of tags per document, W avg Represents the average number of words per document.
[0131] Table 1. Dataset information table
[0132]
[0133] Two ranking-based evaluation metrics are used in the experiments, the top-K label accuracy (P@K) and the normalized discounted cumulative gain of top-K labels (NDCG@K).
[0134] Nine methods were used for comparison in the experiment:
[0135] XML-CNN utilizes a dynamic maximum pooling scheme to capture richer information from different regions of the document, adopts a binary cross entropy loss function to handle multi-label problems, and introduces a hidden bottleneck layer to obtain better document representation and reduce the model size.
[0136] MeSHProbeNet is an end-to-end deep learning model designed for the MeSH index that assigns MeSH terms to MEDLINE citations. It won first place in the latest batch of Task A of the 2018 BioASQ Challenge.
[0137] AttentionXML is a label tree-based deep learning model designed for extreme multi-label text classification. It introduces two key features: a multi-label attention mechanism that captures the relevant text parts of each label, and a shallow and wide probabilistic label tree that can efficiently handle millions of labels.
[0138] Transformer is a network architecture that is the first sequence conversion model based entirely on the attention mechanism. It replaces the commonly used recurrent layer with a multi-head self-attention mechanism and abandons the common recurrent and CNN structures.
[0139] Star-Transformer is a lightweight alternative to Transformer for NLP tasks. It uses a star topology to reduce the complexity from quadratic to linear and solves the problem of high computational requirements.
[0140] BertXML is a customized version of BERT designed specifically for XMTC. It overcomes the limitation of BERT's single [CLS] token by merging multiple [CLS] tokens at the beginning of each input sequence.
[0141] MATCH learns improved text and metadata representations by jointly embedding them into the same space, thereby enabling high-order interactions between words and metadata.
[0142] LiGCN introduces an interpretable graph convolutional network model that models tokens and labels as nodes in a heterogeneous graph. It calculates the cosine similarity between label embeddings to capture the relationship between labels.
[0143] GUDN utilizes label semantics and deep pre-trained models, combined with label enhancement strategies for fine-tuning to improve classification performance. The model exhibits sensitivity to label semantics and shows significant efficacy on datasets with semantically rich labels.
[0144] Tables 2, 3, and 4 summarize the results of different models on the MAG-CS, PubMed, and AAPD datasets, respectively.
[0145] In the MAG-CS dataset (Table 2), the proposed model outperforms all baselines in almost all metrics except P@1, where it lags behind MATCH by 0.0072. However, for P@3, P@5, NDCG@3, and NDCG@5, the proposed model improves over MATCH by 0.0047, 0.0131, 0.0007, and 0.0079, respectively. This indicates that the prediction performance of tail labels is enhanced. This improvement is attributed to the introduction of a large amount of relevant label information, which improves the prediction of tail labels but slightly weakens the focus on head labels. In addition, MATCH uses additional meta-information that LabelCoRank does not have, which can explain the difference in P@1.
[0146] In the PubMed dataset (Table 3), LabelCoRank achieved the best performance in all metrics.
[0147] It is worth noting that LabelCoRank significantly outperforms MATCH on the PubMed dataset compared to the MAG dataset. The improvements in P@3, P@5, NDCG@3, and NDCG@5 are all over 2.8 percentage points, indicating significant progress in predicting the hard-to-predict tail labels. The PubMed dataset has the highest average number of labels per instance, which may enhance the relevance of the supplementary label information, leading to the superior performance improvement.
[0148] In the AAPD dataset (Table 4), LabelCoRank also achieved the best performance in all indicators. Although the average number of labels per document is only 2.41, the total number of labels is 54, allowing most label features to participate in the calculation of the model. LabelCoRank can effectively learn and distinguish between relevant and irrelevant label information without introducing too much noise from additional labels. This proves that the model is very robust even in tasks with fewer labels.
[0149] Table 2. Experimental results of MAG-CS dataset
[0150]
[0151] Table 3. PubMed dataset experimental results
[0152]
[0153] Table 4. AAPD dataset experimental results
[0154]
[0155] The results of ablation experiments (Tables 5-10) verify the effectiveness of the module used to integrate label information in LabelCoRank. Four design elements in LabelCoRank need to be verified for their effectiveness: label selection (correlation matrix), label sequence sorting (label ranking), integration of position information (position information) and selection of the number of labels (Number of labels).
[0156] Label selection directly affects what information in the text will be extracted under the attention mechanism, thus affecting the final label prediction. In order to verify the effectiveness of using the frequency correlation matrix for label selection, a comparison was made with selecting the same number of labels directly from RoBERTa's predictions without using the frequency correlation matrix. The experimental results in Tables 5, 6, and 7 show that using the frequency correlation matrix for label selection is more effective than directly using RoBERTa's predicted labels. This is understandable because the relevance of the predicted labels is gradually reduced. In addition, RoBERTa's predicted labels do not consider co-occurrence relationships outside of text semantics. For some labels, these connections are more hidden and are not directly reflected in the semantic content of the text.
[0157] The label sequence is sorted according to the frequency of the label in the entire dataset, aiming to introduce the frequency distribution information of the label in the dataset. The features are expanded through the multi-label attention mechanism, which naturally puts the high-frequency label features at the forefront, so that the classifier naturally pays different degrees of attention to labels in different positions. The results in Tables 5, 6, and 7 show that it is effective.
[0158] This tag sequence containing frequency distribution information is regarded as a sequence with priority. For the same tag, its meaning varies according to its position; the closer it is to the front, the more attention it should receive, and vice versa, the closer it is to the back, the less attention it should receive. Therefore, incorporating position information into the tag sequence makes it possible to express different features at different positions, and realizes the dynamic representation of tag features. The results in Tables 5, 6, and 7 show that it is effective.
[0159] Different choices of the number of labels will introduce different degrees of relevant label information and noise, and have different effects on different datasets and samples. The impact of the choice of the number of labels is studied through experiments and analysis. The results show that in the MAG dataset, 35 labels are optimal (Table 8); in the Mesh dataset, 30 labels are appropriate (Table 9); in the AAPD dataset, 20 labels are appropriate (Table 10).
[0160] Table 5. Ablation experiment results of MAG-CS dataset
[0161]
[0162] Table 6. PubMed dataset ablation experiment results
[0163]
[0164] Table 7. AAPD dataset ablation experiment results
[0165]
[0166] Table 8. Results of selecting different numbers of labels for the MAG-CS dataset
[0167]
[0168] Table 9. Results of selecting different numbers of labels for the PubMed dataset
[0169]
[0170] Table 10. Results of selecting different numbers of labels for the AAPD dataset
[0171]
[0172] This paper proposes a novel universal multi-label text classification framework that can effectively alleviate the long-tail problem. The core lies in that, through a two-stage re-ranking method, the neglected label co-occurrence information and label frequency distribution information are organically combined. It strengthens the feature representation in a relatively efficient way, and also flexibly integrates the relationship between label-label and label-text, and assigns different label importance to different samples to help prediction. And this technology has good adaptability to both small and large label data sets.
[0173] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the multi-label text classification method based on label relationships as described in the first aspect above or any corresponding embodiment thereof are implemented.
[0174] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the multi-label text classification method based on label relationships as described in the first aspect or any corresponding embodiment thereof.
[0175] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0176] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0177] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0179] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0180] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A multi-label text classification method based on label relationship, characterized in that: include: Capture the text features P in the text dataset through the pre-training model, obtain the initial classification ranking based on the text features, and obtain the first label sequence S1 of the predicted label probability; According to the header tag in the first tag sequence S1, a second tag sequence S2 is obtained; Combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3. Take the union of S2 and S3 to obtain the fourth label sequence S4. Reorder the labels in S4 according to the label frequency distribution information to obtain the frequency-integrated label sequence S. Generate a label feature sequence based on S Through the attention mechanism Establish a semantic relationship with the text to obtain the final feature f cat ; Take the final feature f cat Perform final classification.
2. According to the multi-label text classification method based on label relationship of claim 1, it is characterized in that: Capture the text features P in the text dataset through the pre-training model, obtain the initial classification ranking based on the text features, and obtain the first label sequence S1 of the predicted label probability; include: The initial text is tokenized and preprocessed to obtain a word sequence T of size n; where T = {t cls , t1, t2, t3…, t n-2 , t sep }, t cls is the first special marker word, which is used to represent the characteristics of the text through the pre-trained model. a is the ath non-special tag word, 1≤a≤n-2, t sep is the second special marker word, which is used by the pre-trained model to recognize as the end of the text; The pre-trained model is used to encode T, and the output of the pre-trained model is expressed as: H=RoBERTa(W RoBERTa ,T) P=Fish(h) cls W1+b1) Among them, H is the character-level feature extracted by the pre-training model, Tanh(*) is the activation function, and W ROBERTa is the training parameter of the pre-trained model, P is the text feature, h cls Yes cls Feature expression of text encoded by the pre-trained model, W1 is the first training parameter, b1 is the first training bias parameter, and RoBERTa(*) is the RoBERTa model; Obtain S1 according to P; S1=sigmoid(PW2+b2) Among them, sigmoid(*) is the activation function, W2 is the second training parameter, and b2 is the second training bias parameter; S1={s1,s2,...s K }, where s i is the predicted probability of the i-th label, K is the number of all labels, 1≤i≤K.
3. The multi-label text classification method based on label relationship according to claim 1 is characterized in that: According to the header tag in the first tag sequence S1, a second tag sequence S2 is obtained; include: Select the label prediction probability in S1 that is greater than the threshold hyperparameter α, and use the label corresponding to the selected label prediction probability as the second label sequence S2.
4. The multi-label text classification method based on label relationship according to claim 2 is characterized in that: Combine S2 with the label frequency co-occurrence matrix M from a given text dataset to obtain the third label sequence S3. Take the union of S2 and S3 to obtain the fourth label sequence S4. Reorder the labels in S4 according to the label frequency distribution information to obtain the frequency-integrated label sequence S. Generate a label feature sequence based on S The details are as follows: For a given text dataset with K labels, the co-occurrence frequency matrix M = {m1, m2, ..., m K }, where m i is the co-occurrence frequency sequence of all tags corresponding to the i-th tag, that is, m i ={m i,1 , m i,2 , ..., m i,K },m i,j Indicates the frequency of the jth label when the ith label exists, 1≤j≤K; The second tag sequence S2 includes V tags, S2 = {l1, l2, ...l V }, l v is the vth label, 1≤v≤V; label co-occurrence frequency sequence set is the co-occurrence frequency sequence of all tags corresponding to the vth tag; Sum the tag co-occurrence frequency sequence set to get in, is the label co-occurrence frequency distribution of the second label sequence S2 for K labels, is the frequency of co-occurrence of the kth tag when the vth tag exists; According to the co-occurrence frequency of the K tags, Select the first γ-V tags that do not belong to S2 to form a third tag sequence S3, where γ is a hyperparameter for controlling the expected length of the third tag sequence S3; According to S2 and S3, the fourth tag sequence S4 is calculated: S4=S2∪S3 Sort S4 according to the label frequency distribution information of the entire given text data set, and the order of sorting is from high to low label frequency, and obtain the frequency-integrated label sequence S; Design a label position feature matrix for training and a label feature matrix Among them, δ is the size of the hidden layer in the pre-trained model, To represent the matrix as γ rows and δ columns over the real number field, To represent the matrix as K rows and δ columns in the real number field; Using S from M f Select the label features corresponding to the labels in S and add them to M pos In the above example, we obtain the feature representation F that integrates the label position information; Integrated F and The feature sequence of location information and label information is expressed as follows in, For further integration by and The label feature sequence of the integrated position information and label information, ReLU(*) is the activation function, Drop(*) is a function that randomly makes some parameters not participate in training, W3 is the third training parameter, b3 is the third training bias parameter, W4 is the fourth training parameter, and b4 is the fourth training bias parameter.
5. The multi-label text classification method based on label relationship according to claim 4 is characterized in that: S=Rank(S4) Among them, Rank(*) is to sort S4 by the frequency of the label in a given text dataset, and the order of sorting is from high to low frequency of the label; ReLU(*) is the activation function, i.e., the linear rectification function; The label frequency distribution information is the frequency with which the label appears in a given text dataset.
6. The multi-label text classification method based on label relationship according to claim 5, characterized in that: Through the attention mechanism To establish a semantic relationship with the text, we obtain the final feature f cat ; include: Using label feature sequence Perform masked attention learning on H to extract features including extended label information in, for All individual features in In are concatenated to obtain the final feature f cat ; Among them, Concat(*) is a feature concatenation operation; Take the final feature f cat Conduct final classification; including: f cat Mapped to the corresponding probabilities of K labels through linear transformation and sigmoid function; Y=sigmoid(f cat W5+b5) Among them, Y is the final label prediction result, W5 is the fifth training parameter, and b5 is the fifth training bias parameter; Implement multi-label text classification based on Y.
7. The multi-label text classification method based on label relationship according to claim 6 is characterized in that: The pre-trained model is the RoBERTa model, and sigmoid(*) is the Logistic function; Among them, the final loss L of the RoBERTa model is: L=β×L1+(1-β)×L2 Among them, β is a hyperparameter used to balance the impact of prediction loss L1 and final classification loss L2 on Y; The binary cross entropy loss function is used to calculate L1, and the formula is as follows: in, is the true value of the i-th label, s i is the predicted probability value of the i-th label; For Y = {y1, y2, ...y K The binary cross entropy loss function of} is used as the loss function for the second prediction, and L2 is: in, is the true value of the i-th label, and yi is the predicted probability value of the i-th label.
8. A multi-label text classification system based on label relations, characterized in that: include: A first prediction module is used to capture text features P in a text dataset through a pre-trained model, obtain an initial classification ranking based on the text features, and obtain a first label sequence S1 for predicting label probabilities; The tag feature sequence generation module is used to obtain the second tag sequence S2 according to the head tag in the first tag sequence S1; combine S2 with the tag frequency co-occurrence matrix M from a given text data set to obtain the third tag sequence S3, obtain the fourth tag sequence S4 by taking the union of S2 and S3, reorder the tags in S4 according to the tag frequency distribution information, obtain the frequency-integrated tag sequence S, and generate a tag feature sequence based on S The final feature generation module is used to convert Establish a semantic relationship with the text to obtain the final feature f cat ; Classification module, used to adopt the final feature f cat Perform final classification.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the multi-label text classification method based on label relations as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multi-label text classification method based on label relations as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Public opinion text classification method and system based on multi-label embedding, terminal and medium
CN113987187A
Label processing method and device, electronic equipment and computer readable storage medium
CN114970548A
Text classification method, system and equipment based on multi-label association and medium
CN118227790A
Method, system and computer program product for learning classification model
US20170061330A1
Systems and methods for large scale semantic indexing with deep level-wise extreme multi-label learning
US20200356851A1