Layered text classification method based on label information fusion
By using BERT and Graphormer models to process hierarchical text and fuse label semantics and structural information, the problem of inaccurate classification results in the prior art is solved, and more efficient hierarchical text classification performance is achieved.
Patent Information
- Application Number
- CN202510074034.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing global classification model has shortcomings in utilizing hierarchical information and label semantic information, resulting in inaccurate classification results.
The BERT model is used to model the semantic feature of the tag text, and the Graphormer graph encoder is used to model the hierarchical structure information of the tag. Text embedding and label embedding are fused through direct addition to generate the final text-label hybrid embedding. At the same time, the sigmoidF1 loss function is added to form a joint loss function with the BCE loss function to improve classification performance.
By fusing semantic information and structural information, richer label embeddings are generated, which improves the accuracy and performance of hierarchical text classification, especially in multi-label classification tasks.
Smart Images

Figure CN119988627A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hierarchical text classification, and in particular to a hierarchical text classification method integrating label information and hierarchical structure information. Background Art
[0002] Hierarchical text classification (HTC) is a special type of multi-label classification problem. In real-world scenarios, most text classification tasks are multi-label tasks, and there is a hierarchical structure between labels, which is hierarchical text classification. For example, hierarchical classification of scientific literature, protein function prediction, and patent management in international patent libraries. Conceptually, the hierarchical text classification problem means that the text is associated with multiple labels, and these labels are organized in a structured hierarchy. Mathematically, hierarchical multi-label classification can be defined as a function f:X→Y * , where X is the sample space, Y * is the power set of all possible label sets. * Each element in is an ordered sequence of labels (y1, y2, ..., y n ), represents the hierarchical path from the root node to the leaf node. The goal of hierarchical text classification is to learn a function f so that for any given sample x∈X, the function f(x) can output one or more labels that are most relevant to x. Depending on whether the hierarchical class label information is used and how the hierarchical information is used, the HTC algorithm can be divided into non-hierarchical methods and hierarchical methods: non-hierarchical methods are planar methods; hierarchical methods can be divided into local methods, global methods, and hybrid methods that combine local and global methods. Among them, the global method is currently the most mainstream method, but the existing global classification model does not make sufficient use of the information in the local hierarchical structure, and there is still room for performance improvement. The existing global classification model also has the following defects:
[0003] Lack of effective use of hierarchical information: Although the global method uses only one classifier, avoiding the problems of high computational cost and exposure to bias encountered in the local method, how to more effectively utilize the hierarchy of labels using only a single classifier remains a challenge for the global method.
[0004] Lack of utilization of label semantic information: In addition, the semantics of category label words helps to distinguish different categories. In real classification systems, each category has a unique label in the hierarchy with clear semantic meaning. However, almost all existing HTC methods ignore the semantics of category label words. Summary of the invention
[0005] The purpose of the present invention is to more comprehensively extract and fuse the semantic information and structural information in the hierarchical labels, and propose a hierarchical text classification method to solve the problem of inaccurate classification results of the existing methods.
[0006] A hierarchical text classification method based on label information fusion includes the following steps:
[0007] 1) Data preprocessing: filter punctuation marks in the document, set the maximum text length according to the distribution of text length, and truncate the excess part.
[0008] 2) Construct a standard dataset, including constructing document-label pairs and hierarchical sequences of labels. For datasets with meaningless labels, extract label description keywords as labels for model training.
[0009] 3) Use the encoder of the pre-trained model BERT to encode the document text, obtain the corresponding encoding value, and then generate text embedding through the encoding value.
[0010] 4) Use the customized Graphormer as the structural encoder of the hierarchical text model to obtain the label embedding of the corresponding document.
[0011] 5) The tag embedding and document text embedding are fused by direct addition to obtain the text-tag hybrid embedding for final classification.
[0012] 6) When calculating the classification loss, the sigmoidF1 loss function is added to form a joint loss function together with the BCE loss function.
[0013] 7) Use the training set to train the model, use the results of the validation set to adjust the model parameters, and use the test set to test the performance of the generated model to obtain the optimal model.
[0014] The present invention uses BERT to model the semantic features of label text, so that different categories can be distinguished semantically. In the process of label hierarchy encoding, the semantic information of the label can be used to generate a label representation that incorporates the semantic information.
[0015] Specifically, the present invention uses a Graphormer graph encoder to model the hierarchical structure information of the label. Different from directly modeling the label hierarchy and label text, this method processes the input vector in the Graphormer encoder so that the semantic information of the label participates in the modeling task of the structure encoder, thereby generating a more informative label embedding.
[0016] Furthermore, appropriate embedding fusion technology is used to fuse label embedding and text embedding to participate in the final classification task. In addition, considering that the calculation of F1 score is not differentiable, the sigmoidF1 loss function is added when calculating the classification loss to further improve the classification performance.
[0017] Preferably, the implementation process of step 1) is: filtering the original document text through regular matching, and counting the length of the filtered text; setting the maximum text length according to the distribution of the text length, and truncating the text exceeding the maximum text length; calculating the maximum text length based on the statistical text length.
[0018] Preferably, the implementation process of step 2) is:
[0019] 2.1) Extract documents and tags from the original dataset and generate a standard dataset for training by matching them one by one; for datasets with meaningless tags, extract tag description keywords as tags for model training.
[0020] 2.2) Count all the labels, number them one by one according to the level, and generate a hierarchical sequence of labels for training the hierarchical text classification model.
[0021] Preferably, the implementation process of step 3) is:
[0022] 3.1) Using the pre-trained model BERT, in the task of encoding document text with the pre-trained model BERT, the input of the pre-trained model BERT encoder is a text sequence, wherein the text sequence includes one or more words, and each word in the output of the BERT encoder corresponds to an encoding value, and a CLS tag is added to the front of the encoded output of the document text, and a SEP tag is added to the end.
[0023] 3.2) Convert the encoded value of the text into an embedded value as the text embedding of the document as part of the classifier input.
[0024] 3.3) Use one-hot encoding to encode the label corresponding to the document. The value of the corresponding label in the label vector is set to 1, and the encoding values of other positions are set to 0 for model training.
[0025] Preferably, the implementation process of step 4) is:
[0026] 4.1) For the i-th label y i , its initialization represents f i It consists of label_emb randomly initialized according to the label number and name_emb after semantic encoding. Label_emb is a learnable embedding that takes a label as input and outputs a vector of size 768; name_emb is the average value of the label name after BERT token encoding, and the dimension is also 768.
[0027] 4.2) Use Graphormer graph encoder for modeling. Graphormer adds spatial encoding and edge encoding as bias items of the attention mechanism in the multi-head self-attention mechanism. Then the attention weight matrix A involved in the graph is G Perform Softmax processing, multiply by the value matrix, and then normalize with the residual connection layer, calculate the self-attention, and get L as the label feature for the next step.
[0028] Preferably, the implementation process of step 5) is: using appropriate embedding fusion technology to fuse the text embedding and the image embedding of the label to form the final label embedding. i , and compare it with the text representation h text Combined. Generate fused label text feature F i , which is then fed to the classifier.
[0029] F i =h text +L i
[0030] From the output vector of the classification result, select the i-th element and get the logit score l of label i i .
[0031] l i =(W c T F i +b c ) i
[0032] Among them, W c and b c are the weight and bias of the classifier respectively.
[0033] For the logit score l i After applying sigmoid(.), we get the predicted output y for label i i .
[0034] y i =sigmoid(l i )
[0035] Preferably, the implementation process of step 6) is as follows: As an approximate function of the F1 score, sigmoidF1 solves the non-differentiable problem of the F1 function, allowing it to be used in multi-label classification tasks. It is specially tailored for scenarios where sample prediction may include multiple labels of different numbers, and can optimize label prediction and quantity at the same time. The loss function sigmoidF1 is described as follows:
[0036]
[0037] The BCE loss function and the sigmoidF1 loss function form a joint loss function, and the final loss function is shown in the following formula.
[0038] L=L BCE +L sigmoidF1
[0039] Preferably, the implementation process of step 7) is:
[0040] 7.1) Use the training set to train the model. During the training process, use the validation set to evaluate the classification effect of each model. Select the model with the best evaluation index micro-F1 or macro-F1 in the validation set as the final training model;
[0041] 7.2) The test model uses the evaluation indicators micro-F1 and macro-F1 to evaluate the classification performance of the model.
[0042] Among them, micro-F1 calculates the total precision and recall of all classes, and then calculates the F1 value; macro-F1 first calculates the F1 value of each class, and then averages it. Specifically, the two methods of calculating F1 values are shown in the formula.
[0043]
[0044] Among them, TP represents the number of positive examples correctly predicted by the model as positive examples, FP represents the number of negative examples incorrectly predicted by the model as positive examples, TN represents the number of negative examples correctly predicted by the model as negative examples, and FN represents the number of positive examples incorrectly predicted by the model as negative examples.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] Graphormer is used to obtain the label feature L that incorporates structural information. The structural information is the graph composed of the label hierarchical information in hierarchical text classification. Graphormer adds spatial coding and edge coding as bias items of the attention mechanism in the multi-head self-attention mechanism, making the Transformer encoder more adaptable to the graph data structure.
[0047] In view of the shortcomings of existing methods in characterizing the complex semantic associations between documents and labels, BERT is used to model the semantic features of label text. By processing the input vector in the Graphormer encoder, the semantic information of the label is involved in the modeling task of the structure encoder, realizing the fusion of semantic and hierarchical features, thereby generating more informative label embeddings.
[0048] Hierarchical text classification is a multi-label classification task that involves the prediction of multiple labels and the estimation of the number of labels. The binary cross entropy loss function is used for model training, and the F1 score is used for model evaluation. The sigmoidF1 loss function is added during classification, and together with the BCE loss function, it forms a joint loss function as an approximate function of the F1 score, which solves the non-differentiable problem of the F1 function and allows it to be used in multi-label classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 This is a flow chart of a sentiment classification method for short texts with expressions based on ensemble learning. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0052] like Figure 1 As shown, a hierarchical text classification method based on label information fusion includes the following steps:
[0053] 1) Data preprocessing: filter punctuation marks in the document, set the maximum text length according to the distribution of text length, and truncate the excess part.
[0054] 2) Construct a standard dataset, including constructing document-label pairs and hierarchical sequences of labels. For datasets with meaningless labels, extract label description keywords as labels for model training.
[0055] 3) Use the encoder of the pre-trained model BERT to encode the document text, obtain the corresponding encoding value, and then generate text embedding through the encoding value.
[0056] 4) Use the customized Graphormer as the structural encoder of the hierarchical text model to obtain the label embedding of the corresponding document.
[0057] 5) The tag embedding and document text embedding are fused by direct addition to obtain the text-tag hybrid embedding for final classification.
[0058] 6) When calculating the classification loss, the sigmoidF1 loss function is added to form a joint loss function together with the BCE loss function.
[0059] 7) Use the training set to train the model, use the results of the validation set to adjust the model parameters, and use the test set to test the performance of the generated model to obtain the optimal model.
[0060] The implementation process of step 1) is: filter the original document text through regular matching, and count the length of the filtered text; set the maximum text length according to the distribution of text length, and truncate the text exceeding the maximum text length; calculate the maximum text length according to the length of the statistical text.
[0061] The implementation process of step 2) is:
[0062] 2.1) Extract documents and tags from the original dataset and generate a standard dataset for training by matching them one by one; for datasets with meaningless tags, extract tag description keywords as tags for model training.
[0063] 2.2) Count all the labels, number them one by one according to the level, and generate a hierarchical sequence of labels for training the hierarchical text classification model.
[0064] The implementation process of step 3) is:
[0065] 3.1) Using the pre-trained model BERT, in the task of encoding document text with the pre-trained model BERT, the input of the pre-trained model BERT encoder is a text sequence, wherein the text sequence includes one or more words, and each word in the output of the BERT encoder corresponds to an encoding value, and a CLS tag is added to the front of the encoded output of the document text, and a SEP tag is added to the end.
[0066] 3.2) Convert the encoded value of the text into an embedded value as the text embedding of the document as part of the classifier input.
[0067] 3.3) Use one-hot encoding to encode the label corresponding to the document. The value of the corresponding label in the label vector is set to 1, and the other positions are encoded as 0 for model training.
[0068] Step 4) The implementation process is:
[0069] 4.1) For the i-th label y i , its initialization represents f i It consists of label_emb randomly initialized according to the label number and name_emb after semantic encoding. Label_emb is a learnable embedding that takes a label as input and outputs a vector of size 768; name_emb is the average value of the label name after BERT token encoding, and the dimension is also 768.
[0070] 4.2) Use Graphormer graph encoder for modeling. Graphormer adds spatial encoding and edge encoding as bias items of the attention mechanism in the multi-head self-attention mechanism. Then the attention weight matrix A involved in the graph is GPerform Softmax processing, multiply by the value matrix, and then normalize with the residual connection layer, calculate the self-attention, and get L as the label feature for the next step.
[0071] The implementation process of step 5) is: using appropriate embedding fusion technology to fuse the text embedding and graph embedding of the label to form the final label embedding. i , and compare it with the text representation h text Combined. Generate fused label text feature F i , and then feed it to the classifier. From the output vector of the classification result, select the i-th element and get the logit score l of label i. i . For the logit score l i After applying sigmoid(.), we get the predicted output y for label i i .
[0072] The implementation process of step 6) is as follows: As an approximate function of the F1 score, sigmoidF1 solves the non-differentiable problem of the F1 function, allowing it to be used in multi-label classification tasks. It is specially tailored for scenarios where sample predictions may include multiple labels of different numbers, and can optimize label predictions and numbers at the same time. The BCE loss function and the sigmoidF1 loss function form a joint loss function.
[0073] The implementation process of step 7) is:
[0074] 7.1) Use the training set to train the model. During the training process, use the validation set to evaluate the classification effect of each model. Select the model with the best evaluation index micro-F1 or macro-F1 in the validation set as the final training model;
[0075] 7.2) The test model uses the evaluation indicators micro-F1 and macro-F1 to evaluate the classification performance of the model.
[0076] Micro-F1 calculates the total precision and recall of all classes, and then calculates the F1 value; macro-F1 first calculates the F1 value of each class, and then averages it.
[0077] Embodiment 1:
[0078] (1) Data preprocessing
[0079] In this method, the original document text is first filtered through regular matching, and the length of the filtered text is counted; the maximum text length is set according to the distribution of the text length, and the text exceeding the maximum text length is truncated; because this method uses BERT as the text encoder, the maximum text length cannot exceed 512. Under this premise, ensure that the number of truncated documents does not exceed 5% of the total number of documents, calculate the maximum text length based on the length of the statistical document text, and fill in the encoding result for the case where the selected text length is less than 512 for training.
[0080] (2) Construct a standard data set and generate a hierarchical sequence
[0081] Extract documents and labels from the original data set, and form document-label pairs in one-to-one correspondence to generate a standard data set for training. In order to utilize the semantic information of the category label, the premise is that the label words have real meaning. For datasets where the category label words have no real meaning, it is necessary to extract label description keywords separately as label words. Then, count all labels, number the labels one by one according to the level, and generate a hierarchical sequence of labels for training the hierarchical text classification model.
[0082] (3) Generation of text embedding
[0083] The text encoder of this method uses the encoder of the pre-trained model BERT to encode the document text, obtain the corresponding encoding value, and then generate text embedding through the encoding value. This embodiment uses the pre-trained model bert-base-uncased of huggingface as the text encoder, inputs a text sequence, and the text sequence includes one or more words. Each word in the output of the BERT encoder corresponds to an encoding value. The CLS tag is added to the front of the encoding output of the document text, and the SEP tag is added to the end. Then, the above encoding result is converted into an embedding value through the embedding layer of the BERT model as the text embedding of the document, as part of the classifier input.
[0084] (4) Generation of label embedding
[0085] This method uses the pre-trained model BERT to model the semantic features of the label text, and uses the Graphormer graph encoder to model the hierarchical structure information of the label. By processing the input vector in the Graphormer encoder, the semantic information of the label is involved in the modeling task of the structure encoder, thereby generating a more informative label embedding.
[0086] f i =label_emb(y i )+name_emb(y i )
[0087] For the i-th label y i , its initialization represents f i It consists of label_emb randomly initialized according to the label number and name_emb after semantic encoding. Among them, label_emb is a learnable embedding that takes a label as input and outputs a vector of size dh; name_emb is the average value of the label name after BERT token encoding. In addition, the embedding weights are shared between documents and labels to make the label features more instructive. All node features are stacked into a matrix F∈R k×dh , and then use the standard self-attention mechanism layer for feature transfer.
[0088] This method uses Graphormer to obtain the label feature L that integrates the structural information, and then the attention weight matrix A involved in the graph G After Softmax processing, multiply it by the value matrix and normalize it with the residual connection layer, calculate the self-attention, and get L as the label feature for the next step. Since the graph in the HTC problem is a tree structure, for node y i and j , there is only one path (e1,e2,...,e D ), so c ij It can represent the edge information between two nodes. is a learnable weight scalar for each edge.
[0089]
[0090] N=φ(y i ,y j )
[0091] L = LayerNorm(softmax(A G )V+F)
[0092] Where: c ij is the label y i and label y j The edge encoding between is a learnable scalar; φ(y i ,y j ) is the label y i and label y j The shortest path distance in the graph, It is a learnable scalar corresponding to the shortest path distance.
[0093] (5) Text-label embedding fusion
[0094] Use appropriate embedding fusion technology to fuse the text embedding and graph embedding of the label to form the final label embedding. i , and compare it with the text representation h text Combined. Generate fused label text feature F i , which is then fed to the classifier.
[0095] F i =h text +L i
[0096] From the output vector of the classification result, select the i-th element and get the logit score l of label i i .
[0097] l i =(W c T F i +b c ) i
[0098] Among them, W c and b c are the weight and bias of the classifier respectively.
[0099] For the logit score l i After applying sigmoid(.), we get the predicted output y for label i i .
[0100] y i =sigmoid(l i )
[0101] (6) Joint loss function
[0102] When calculating the classification loss, the sigmoidF1 loss function is added to form a joint loss function together with the BCE loss function. The loss function sigmoidF1 is described as follows:
[0103]
[0104] The BCE loss function and the sigmoidF1 loss function form a joint loss function, and the final loss function is shown in the following formula.
[0105] L=L BCE +L sigmoidF1
[0106] (7) Model training and testing
[0107] In order to train the model proposed in the present invention, the Adam algorithm was selected as the optimization algorithm and trained for 100 cycles. The optimization strategy adopted the early stopping strategy. Whether to stop training was determined based on the performance of the validation set to prevent overfitting. During the training process, hyperparameters such as the learning rate and batch size were adjusted to obtain the best performance.
[0108] The model is trained with the training set, where the data set includes all samples, each sample includes text and corresponding multiple labels, and the data set is divided into 8:2 ratios of training set plus validation set and test set, and 8:2 ratios of training set and validation set. The validation set is used to evaluate the effect of each training session during training, and the model with the best micro-F1 or macro-F1 performance in the validation set is selected as the final model.
[0109] After training, the test model uses micro-F1 and macro-F1 to evaluate the classification performance of the model.
[0110] Micro-F1 calculates the total precision and recall of all classes, and then calculates the F1 value; macro-F1 first calculates the F1 value of each class, and then averages it. The two methods of calculating F1 values are shown in the formula.
[0111]
[0112] Among them, TP represents the number of positive examples correctly predicted by the model as positive examples, FP represents the number of negative examples incorrectly predicted by the model as positive examples, TN represents the number of negative examples correctly predicted by the model as negative examples, and FN represents the number of positive examples incorrectly predicted by the model as negative examples.
[0113] The word semantics of category labels are very helpful for semantic distinction between different categories. This method considers the word semantics of category labels and fuses the semantic information of hierarchical labels with the structural information of hierarchical labels, thereby improving the performance of HTC. In addition, compared with other hybrid methods of label embedding and text embedding, the direct addition method reduces the loss of feature information and obtains better classification performance. In hierarchical text classification, this method adds the sigmoidF1 loss function and forms a joint loss function with the BCE loss function, which also improves the classification performance. As an approximate function of the F1 score, sigmoidF1 solves the non-differentiable problem of the F1 function, allowing it to be used in multi-label classification tasks. It is specially tailored for scenarios where sample predictions may include multiple labels of different numbers, and can optimize label predictions and quantities at the same time.
[0114] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them.
Claims
1. A hierarchical text classification method based on label information fusion, characterized in that: The following steps are involved: 1) Data preprocessing: filter punctuation marks in the document, set the maximum text length according to the distribution of text length, and truncate the excess part. 2) Construct a standard dataset, including constructing document-label pairs and hierarchical sequences of labels. For datasets with meaningless labels, extract label description keywords as labels for model training. 3) Use the encoder of the pre-trained model BERT to encode the document text, obtain the corresponding encoding value, and then generate text embedding through the encoding value. 4) Use the customized Graphormer as the structural encoder of the hierarchical text model to obtain the label embedding of the corresponding document. 5) The tag embedding and document text embedding are fused by direct addition to obtain the text-tag hybrid embedding for final classification. 6) When calculating the classification loss, the sigmoidF1 loss function is added to form a joint loss function together with the BCE loss function. 7) Use the training set to train the model, use the results of the validation set to adjust the model parameters, and use the test set to test the performance of the generated model to obtain the optimal model.
2. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: In step 1): filter the original document text by regular matching, and count the length of the filtered text; set the maximum text length according to the distribution of text length, and truncate the text exceeding the maximum text length; calculate the maximum text length according to the length of the statistical text.
3. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: The implementation process of step 2) is: 2.1) Extract documents and tags from the original dataset and generate a standard dataset for training by matching them one by one; for datasets with meaningless tags, extract tag description keywords as tags for model training. 2.2) Count all the labels, number them one by one according to the level, and generate a hierarchical sequence of labels for training the hierarchical text classification model.
4. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: The implementation process of step 3) is: 3.1) Using the pre-trained model BERT, in the task of encoding document text with the pre-trained model BERT, the input of the pre-trained model BERT encoder is a text sequence, wherein the text sequence includes one or more words, and each word in the output of the BERT encoder corresponds to an encoding value, and a CLS tag is added to the front of the encoded output of the document text, and a SEP tag is added to the end. 3.2) Convert the encoded value of the text into an embedded value as the text embedding of the document as part of the classifier input. 3.3) Use one-hot encoding to encode the label corresponding to the document. The value of the corresponding label in the label vector is set to 1, and the other positions are encoded as 0 for model training.
5. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: Step 4) The implementation process is: 4.1) For the i-th label y i , its initialization represents f i It consists of label_emb randomly initialized according to the label number and name_emb after semantic encoding. Label_emb is a learnable embedding that takes a label as input and outputs a vector of size 768; name_emb is the average value of the label name after BERT token encoding, and the dimension is also 768. 4.2) Use Graphormer graph encoder for modeling. Graphormer adds spatial encoding and edge encoding as bias items of the attention mechanism in the multi-head self-attention mechanism. Then the attention weight matrix A involved in the graph is G Perform Softmax processing, multiply by the value matrix, and then normalize with the residual connection layer, calculate the self-attention, and get L as the label feature for the next step.
6. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: The implementation process of step 5) is: using appropriate embedding fusion technology to fuse the text embedding and graph embedding of the label to form the final label embedding. i , and compare it with the text representation h text Combined. Generate fused label text feature F i , which is then fed to the classifier. F i =h text +L i From the output vector of the classification result, select the i-th element and get the logit score l of label i i . l i =(W c T F i +b c ) i Among them, W c and b c are the weight and bias of the classifier respectively. For the logit score l i After applying sigmoid(.), we get the predicted output y for label i i . y i =sigmoid(l i )。 7. A hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: The implementation process of step 6) is as follows: As an approximate function of the F1 score, sigmoidF1 solves the non-differentiable problem of the F1 function, allowing it to be used in multi-label classification tasks. It is specially tailored for scenarios where sample predictions may include multiple labels of different numbers, and can optimize label predictions and numbers at the same time. The loss function sigmoidF1 is described as follows: The BCE loss function and the sigmoidF1 loss function form a joint loss function, and the final loss function is shown in the following formula. L=L BCE +L sigmoidF1。 8. The hierarchical text classification method based on label information fusion as claimed in claim 1, characterized in that: The implementation process of step 7) is: 7.1) Use the training set to train the model. During the training process, use the validation set to evaluate the classification effect of each model. Select the model with the best evaluation index micro-F1 or macro-F1 in the validation set as the final training model; 7.2) The test model uses the evaluation indicators micro-F1 and macro-F1 to evaluate the classification performance of the model. Micro-F1 calculates the total precision and recall of all classes, and then calculates the F1 value; macro-F1 first calculates the F1 value of each class, and then averages it. The two methods of calculating F1 values are shown in the formula. Among them, TP represents the number of positive examples correctly predicted by the model as positive examples, FP represents the number of negative examples incorrectly predicted by the model as positive examples, TN represents the number of negative examples correctly predicted by the model as negative examples, and FN represents the number of positive examples incorrectly predicted by the model as negative examples.
Citation Information
Patent Citations
Construction method of government purchase item hierarchical classification model
CN113946678A
Public opinion text classification method and system based on multi-label embedding, terminal and medium
CN113987187A
Hierarchical text classification method based on pre-training generative model
CN115422349A