Hierarchical multi-label professional technical document classification method and system using label information

Through the hierarchical multi-label professional document classification method, the tag perception supervision comparison learning and hierarchical perception label attention module are used to solve the problem of neglecting the label hierarchy in the existing technology, the accuracy of document classification and the learning of tag hierarchy are realized, and the accuracy of document representation is improved.

CN116932763BActive Publication Date: 2025-08-12HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310973520.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-08-12
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

The standard label hierarchy and rich label semantic information are ignored in the prior art, resulting in inaccurate classification of professional technical documents.

Method used

The hierarchical multi-label professional document classification method is adopted, and the learning module is compared with the hierarchical perception of the tag attention module. Combined with the classification module, the text sequence representation of the text embedding attention is constructed, and the total loss function is constructed for model training, and the tag hierarchy is dynamically integrated.

Benefits of technology

Improve the accuracy of document classification, capture the semantic representation of text and learn label hierarchy information, and enhance the accuracy of document representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116932763B_ABST
    Figure CN116932763B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, storage medium and electronic device for classifying hierarchical multi-label professional and technical documents using label information, and relates to the technical field of document classification. The present invention fully utilizes label information for hierarchical multi-label professional and technical documents to accurately classify documents. This method can not only capture the semantic representation of the text, but also learn label hierarchy information. Among them, the supervised contrastive learning module based on label perception introduces the definition of label similarity in supervised contrastive learning; a similarity relationship is established between the anchor sample and the positive sample, making the document representation more accurate. The label attention module based on hierarchical perception captures the semantic relationship between the document and the corresponding hierarchical perception label description, with the purpose of dynamically integrating the label hierarchy into the classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document classification, and in particular to a hierarchical multi-label professional technical document classification method, system, storage medium and electronic device using label information. Background Art

[0002] Professional technical documents (hereinafter referred to as documents) provide a wealth of information for innovation management and engineering management. In recent years, the rapid increase in the types and quantity of documents has brought new challenges to the accurate classification of documents.

[0003] Existing research largely focuses on leveraging document semantic information (such as titles and abstracts) for classification tasks. For example, see the paper "Li S, Hu J, Cui Y, et al. DeepPatent: patent classification with convolutional neural networks and word embedding [J]. Scientometrics, 2018, 117:721-744." Li et al. proposed DeepPatent, a deep learning algorithm based on convolutional neural networks and word embedding, to mine textual information from titles and abstracts for automatic patent classification. However, the standard label hierarchy and rich label semantic information in existing classification systems are generally overlooked. Summary of the Invention

[0004] (1) Technical problems solved

[0005] In response to the deficiencies of the prior art, the present invention provides a hierarchical multi-label professional technical document classification method, system, storage medium and electronic device using tag information, which solves the technical problem of ignoring the standard tag hierarchy and rich tag semantic information.

[0006] (2) Technical solution

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0008] A hierarchical multi-label professional and technical document classification method using label information is based on a hierarchical multi-label professional and technical document classification model. The model includes a label-aware supervised contrastive learning module, a hierarchical label attention module, and a classification module. The classification method includes:

[0009] Obtain several document samples and divide them into training sets and test sets;

[0010] The anchor samples and corresponding positive and negative samples in the training set are used as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss;

[0011] The anchor samples and their hierarchical multi-labels are used as the input of the label attention module to obtain the text sequence representation with label embedding attention;

[0012] The text hidden representation and text sequence representation are used as inputs of the classification module to obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples; based on the first and second predicted probabilities, the document classification loss and text sequence representation classification loss are constructed respectively;

[0013] Construct a total loss based on label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and train the model until convergence;

[0014] The test set is used as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

[0015] Preferably, the method of using the anchor samples and the corresponding positive samples in the training set as inputs of the supervised contrastive learning module to obtain the hidden representation of the text includes:

[0016] Generate a first string sequence of anchor samples; use the first string sequence as input to the document encoder of the supervised contrastive learning module to obtain a first sequence representation;

[0017] The first sequence representation is used as the input of the projection network of the supervised contrastive learning module to obtain the text hidden representation.

[0018] Preferably, the label-aware supervised contrastive learning loss specifically refers to:

[0019]

[0020]

[0021]

[0022] Among them, L LSCLM represents the label-aware supervised contrastive learning loss; N represents the number of document samples in the current batch, and let i∈A≡{1,2,…,N} be the index of the anchor sample;

[0023] represents the label-aware supervised contrastive learning loss of the i-th anchor sample; L i represents the hierarchical multi-label set of the i-th anchor sample, L p represents the hierarchical multi-label set of the pth positive sample corresponding to the i-th anchor sample; P(i) represents the positive sample set corresponding to the i-th anchor sample; z i 、z p 、z aRepresents the i-th anchor sample and its positive sample, positive and negative samples respectively; sim represents the similarity function; exp represents the exponential function; τ represents the temperature coefficient;

[0024] J(X,Y)∈[0,1] represents the Jaccard similarity of any two label sets X and Y. The similarity of two sets is defined as the cardinality of their intersection |X∩Y| divided by the cardinality of their union |X∪Y|.

[0025] Preferably, the method of using the anchor sample and its hierarchical multi-label as input to the label attention module to obtain the label-embedded attention text sequence representation includes:

[0026] Generate a second string sequence of anchor samples; use the second string sequence as a document encoder of the label attention module to obtain a second sequence representation;

[0027] Generate a representation of the hierarchical perceptual path corresponding to the anchor sample; use the representation of the hierarchical perceptual path as the input of the document encoder of the label attention module to obtain label features;

[0028] According to the second sequence representation and label features, a label-embedded attentive text sequence representation is obtained based on the attention mechanism.

[0029] Preferably, the first and second predicted probabilities specifically refer to:

[0030] p ik =sigmoid(W c z i +b c ) k

[0031]

[0032] Among them, p ik 、p ik Represent the first and second predicted probabilities respectively; z i 、 Represents text sequence representation and label embedding attention text sequence representation respectively; sigmoid represents activation function; W c represents the weight matrix; b c represents the bias matrix; k represents the index of the label.

[0033] Preferably, the document classification loss and text sequence representation classification loss specifically refer to:

[0034]

[0035]

[0036] Among them, L c 、 They represent document classification loss and text sequence classification loss respectively; K represents the number of labels; y ik represents the true value; log represents the logarithmic function.

[0037] Preferably, the total loss specifically refers to:

[0038]

[0039] Where L represents the total loss and λ represents the hyperparameter.

[0040] A hierarchical multi-label professional and technical document classification system utilizing label information is based on a hierarchical multi-label professional and technical document classification model. The model includes a label-aware supervised contrastive learning module, a hierarchical label attention module, and a classification module. The classification system includes:

[0041] The partitioning module is used to obtain several document samples and divide them into training sets and test sets;

[0042] The first acquisition module is used to take the anchor samples and corresponding positive and negative samples in the training set as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss;

[0043] The second acquisition module is used to take the anchor sample and its hierarchical multi-label as the input of the label attention module to obtain the text sequence representation of the label embedding attention;

[0044] The prediction module is used to take the text hidden representation and text sequence representation as inputs of the classification module, obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples, and construct the document classification loss and text sequence representation classification loss based on the first and second predicted probabilities respectively;

[0045] A training module that constructs a total loss based on label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and trains the model until convergence;

[0046] The testing module is used to use the test set as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

[0047] A storage medium stores a computer program for classifying hierarchical multi-label professional and technical documents using label information, wherein the computer program enables a computer to execute the hierarchical multi-label professional and technical document classification method as described above.

[0048] An electronic device, comprising:

[0049] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs including a method for executing the hierarchical multi-label professional technical document classification method as described above.

[0050] (3) Beneficial effects

[0051] The present invention provides a hierarchical multi-label professional technical document classification method, system, storage medium, and electronic device using label information. Compared with the existing technology, it has the following advantages:

[0052] The present invention targets hierarchical multi-label professional and technical documents, fully utilizing label information to accurately classify them. This method not only captures the semantic representation of text, but also learns label hierarchy information. Specifically, a label-aware supervised contrastive learning module introduces the definition of label similarity in supervised contrastive learning; a similarity relationship is established between anchor samples and positive samples, making document representation more accurate. A hierarchical-aware label attention module captures the semantic relationship between documents and corresponding hierarchical-aware label descriptions, with the goal of dynamically integrating the label hierarchy into the classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 A block diagram of a hierarchical multi-label professional technical document classification method using label information provided by an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of the structure of a hierarchical multi-label professional technical document classification model provided by an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of a hierarchical perception path provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0058] The embodiments of the present application solve the technical problem of ignoring the standard label hierarchy and rich label semantic information by providing a hierarchical multi-label professional technical document classification method, system, storage medium and electronic device that utilize label information.

[0059] The technical solution in the embodiments of the present application is to solve the above technical problems, and the overall idea is as follows:

[0060] The embodiment of the present invention proposes a hierarchical multi-label document classification method that fully utilizes label information, which specifically includes three parts: a label-aware supervised contrastive learning module, a hierarchical label embedding attention module, and a classification module.

[0061] The label-aware supervised contrastive learning module introduces label set similarity in supervised contrastive learning. In supervised contrastive learning, it is assumed that the data features belonging to the same class (anchor samples and positive samples) are similar, while the data features belonging to different classes (anchor samples and negative samples) are different. Supervised contrastive learning aims to cluster samples belonging to the same class together in the semantic space and push samples of different classes apart in the semantic space. The embodiment of the present invention introduces the definition of label set similarity in supervised contrastive learning, so that anchor samples and positive samples with high label set similarity are closer in the semantic space to improve the representation of documents.

[0062] The hierarchical-aware tag attention module dynamically incorporates the hierarchical structure of tags into the classification system. This embodiment of the present invention correlates the representation of each document with the corresponding tag hierarchy representation. After sufficient training, this tag hierarchy information can be incorporated into the classification model to improve model performance. During the model testing phase, this module is discarded to avoid label exposure.

[0063] The classification module classifies documents and calculates the text representation classification loss based on the outputs of the label-aware supervised contrastive learning module and the hierarchy-aware label attention module.

[0064] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0065] Example:

[0066] like Figure 1As shown, the embodiment of the present invention provides a hierarchical multi-label professional technical document classification method using label information, based on Figure 2 The hierarchical multi-label professional technical document classification model shown in FIG. includes a label-aware supervised contrastive learning module, a hierarchical label attention module, and a classification module. The classification method includes:

[0067] S1. Obtain several document samples and divide them into training set and test set;

[0068] S2. Use the anchor samples and corresponding positive and negative samples in the training set as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss;

[0069] S3. Take the anchor sample and its hierarchical multi-label as the input of the label attention module to obtain the text sequence representation of the label embedding attention;

[0070] S4. Use the text hidden representation and text sequence representation as inputs of the classification module, respectively, to obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples; construct the document classification loss and text sequence representation classification loss based on the first and second predicted probabilities, respectively;

[0071] S5. Construct the total loss based on the label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and train the model until convergence;

[0072] S6. Use the test set as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

[0073] The embodiment of the present invention targets hierarchical multi-label professional technical documents, fully utilizes label information, and accurately classifies the documents. The method can not only capture the semantic representation of the text, but also learn the label hierarchy information.

[0074] Next, we will combine the instructions Figures 2-3 The following are the steps of the above technical solution:

[0075] In step S1, several document samples are obtained and divided into a training set and a test set.

[0076] In step S2, the anchor samples and corresponding positive and negative samples in the training set are used as input to the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss.

[0077] The label-aware supervised contrastive learning module in the embodiment of the present invention adopts a two-stage training method. Specifically, it includes:

[0078] In the first stage, anchor samples and corresponding positive samples are input to a document encoder to generate sample representations; these representations are input to a projection network, and the outputs are normalized on the unit hypersphere, allowing inner products to be used to measure distances in the projected space.

[0079] In the second stage, the projection network is discarded, the weight network of the document encoder is frozen, and the document encoder network is not fine-tuned.

[0080] Specifically, the S2 includes:

[0081] S21. Generate a first string sequence of anchor samples; use the first string sequence as input to the document encoder of the supervised contrastive learning module to obtain a first sequence representation.

[0082] In this step, the BERT pre-trained language model is selected as the document encoder to dynamically generate embedding vectors in the semantic space. A major advantage of BERT is its powerful language representation and feature extraction capabilities. It uses a "masked language model" to pre-train multiple bidirectional Transformer encoders.

[0083] Given the first string sequence of a document as input:

[0084] S i ={[CLS],w1,w2,...,w n-2 ,[SEP]}

[0085] Among them, [CLS] and [SEP] are two unique markers that specify the start and end of the sequence, w j Denotes the jth character token in the sequence. The first string sequence of the document is input into the document encoder.

[0086] For each first string sequence, the document encoder generates a first sequence representation:

[0087] H i ={r [CLS] ,r1,r,...,r n-2 ,r [SEP]}=BERT(S i )

[0088] in, d h Indicates the dimension of string embedding, [CLS] is used to represent the entire string sequence, expressed as h i =r [CLS] .

[0089] S22. Using the first sequence representation as the input of the projection network of the supervised contrastive learning module to obtain the text hidden representation.

[0090] Introducing a learnable nonlinear transformation between the representation layer and the contrastive learning loss can significantly improve the quality of the learned representation. The present embodiment uses a tiny projection network Proj(·) during training to map the sequence representation into a space that performs hierarchical supervised contrastive learning loss.

[0091] Afterwards, this step discards the projection network Proj(·) and uses the document encoder and sequence representation r i Perform the classification task. For example, a multilayer perceptron with one hidden layer is used to obtain the representation vector:

[0092] z i =Proj(·)=W (2) RELU(W (1) h i )

[0093] in, represents the weight matrix, z i Represents the text hidden representation, and ReLU() represents the linear rectification function.

[0094] S23, positive sample sampling.

[0095] Given a batch of N document samples, let i∈A≡{1,2,…,N} be the index of the anchor sample. First, samples with the same label as the anchor sample are taken as positive samples, and the rest of the samples in the batch are taken as negative samples. Then, another anchor sample is randomly selected, and the sampling strategy is repeated until all samples are sampled.

[0096] S24. Label-aware supervised contrastive learning loss.

[0097] The greater the similarity between the label sets of two document samples, the greater the semantic similarity of their texts. Based on the supervised contrastive learning loss, this step imposes a stronger penalty on positive sample pairs with high similarity, forcing them to be closer in the semantic space than positive sample pairs with low similarity.

[0098] This step uses the Jaccard similarity function, which defines the similarity between two sets as the cardinality of their intersection divided by the cardinality of their union. For two label sets X and Y, the similarity value is:

[0099]

[0100] Here, J(X,Y)∈[0,1] represents the Jaccard similarity of any two label sets X and Y, and the similarity of two sets is defined as the cardinality of their intersection |X∩Y| divided by the cardinality of their union |X∪Y|.

[0101] In a batch of documents, let i∈A≡{1,2,…,N} be the index of the anchor sample, and N represents the number of document samples in the current batch.

[0102] In label-aware supervised contrastive learning loss, the loss function is expressed as follows:

[0103]

[0104] in, represents the label-aware supervised contrastive learning loss of the i-th anchor sample; L i represents the hierarchical multi-label set of the i-th anchor sample, L p represents the hierarchical multi-label set of the pth positive sample corresponding to the i-th anchor sample; P(i) represents the positive sample set corresponding to the i-th anchor sample; z i 、z p 、z a They represent the i-th anchor sample and its positive sample, positive and negative samples respectively; sim represents the similarity function; exp represents the exponential function; τ represents the temperature coefficient.

[0105] The total loss is the average loss over all samples:

[0106]

[0107] Among them, L LSCLM represents the label-aware supervised contrastive learning loss,

[0108] In step S3, the anchor sample and its hierarchical multi-label are used as the input of the label attention module to obtain the text sequence representation with label embedding attention;

[0109] The predefined categories of documents contain important classification clues and can leverage label information to improve the performance of classification models. In a classification hierarchy, child nodes always inherit certain properties from their parent nodes (except the root node). The label embedding module can leverage not only the label information of leaf nodes, but also the label information provided by their corresponding parent and grandparent nodes to improve model performance.

[0110] An embodiment of the present invention proposes a hierarchical-aware label attention module, which embeds leaf node label information and hierarchical structure information into real-valued vectors with semantic meaning, and obtains a text sequence representation with label embedding attention by weighting the label embedding with each word vector in the document.

[0111] Specifically, the S3 includes:

[0112] S31. Label embedding.

[0113] like Figure 3As shown in the figure, assuming that a classification label of a document is A01B, then a top-down hierarchical perception path "A-A01-A01B" will be obtained. In the hierarchical structure of labels, each label node contains a unique ID number and a detailed description. In order to learn more about the hierarchical perception path, this step uses the ID number and corresponding description of each label node in the hierarchical perception path as the input of the hierarchical perception label attention module.

[0114] Generate the second string sequence of anchor samples {w1,w2,...,w n-2}; Use the second string sequence as the document encoder of the label attention module to obtain the second sequence representation

[0115] Generate the representation H(S) of the hierarchical perceptual path corresponding to the anchor sample i )={σ1,σ2,...,σ p}, where σ i represents the i-th string in the label text sequence (including the ID number and corresponding description of each label node), and p is the length of the label text; the representation of the hierarchical perception path is used as the input of the document encoder of the label attention module to obtain the label feature.

[0116] For example, H(S i ) is also embedded by BERT as:

[0117] {v1,v2,...,v p}=BERT(H(S i ))

[0118]

[0119] Among them, v i Represents the vector representation of each string in the label text sequence, represents the label feature, which is averaged over each word vector in the hierarchical perceptual path.

[0120] S32. According to the second sequence representation and label features, a text sequence representation with label embedding attention is obtained based on the attention mechanism.

[0121] In order to obtain the importance of each second string in the document on its corresponding hierarchical-aware label embedding, this step calculates the attention weight between the document string embedding and the label feature:

[0122]

[0123] Among them, q j , k represent the query vector and key vector in the attention mechanism respectively, kT represents the transpose of k, d h It is represented by the matrix q j , the dimension of the column vector in k, the weight matrix SoftMax() represents the normalized exponential function.

[0124] Then, the text sequence representation with label embedding attention can be obtained by taking the average weight of the second string:

[0125]

[0126] Among them, the text sequence representation of the label embedding attention Should contain label information, so the embodiment of the present invention uses The classification loss serves as a guide for the document encoder and the hierarchical-aware label attention module.

[0127] In step S4, the text hidden representation and the text sequence representation are respectively used as inputs of the classification module to obtain the first and second prediction probabilities of the classification results corresponding to the anchor samples; based on the first and second prediction probabilities, the document classification loss and the text sequence representation classification loss are respectively constructed.

[0128] This step inputs the text sequence representations generated by the label-aware supervised contrastive learning module and the layer-aware label attention module into the classification layer. The sequence representations are then fed into a linear layer, where the probability of text S on label k is calculated using a sigmoid function. Specifically, the first and second predicted probabilities are:

[0129] p ik =sigmoid(W c z i +b c ) k

[0130]

[0131] Among them, p ik 、p ik Represent the first and second predicted probabilities respectively; z i 、 Represent text hidden representation and text sequence representation respectively; sigmoid represents the activation function; represents the weight matrix; represents the bias matrix; k represents the index of the label.

[0132] The embodiment of the present invention selects a binary cross entropy loss function for the multi-label classification task. That is, the document classification loss and the text sequence representation classification loss are respectively:

[0133]

[0134]

[0135] Among them, L C 、 They represent document classification loss and text sequence classification loss respectively; K represents the number of labels; y ik represents the true value; log represents the logarithmic function.

[0136] In step S5, the total loss is constructed based on the label-aware supervised contrastive learning loss, the document classification loss, and the text sequence representation classification loss, and the model is trained until convergence.

[0137] The final loss function is a combination of document classification loss, label embedding-attentive text sequence representation classification loss, and label-aware supervised contrastive learning loss, specifically:

[0138]

[0139] Where L represents the total loss and λ represents the hyperparameter.

[0140] In step S6, the test set is used as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

[0141] It should be noted that, in fact, during the testing phase, the embodiment of the present invention only uses a document encoder with a classification layer to classify documents.

[0142] The embodiment of the present invention provides a hierarchical multi-label professional technical document classification system using label information. The system is based on a hierarchical multi-label professional technical document classification model. The model includes a label-aware supervised contrastive learning module, a hierarchical label attention module, and a classification module. The classification system includes:

[0143] The partitioning module is used to obtain several document samples and divide them into training sets and test sets;

[0144] The first acquisition module is used to take the anchor samples and corresponding positive and negative samples in the training set as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss;

[0145] The second acquisition module is used to take the anchor sample and its hierarchical multi-label as the input of the label attention module to obtain the text sequence representation of the label embedding attention;

[0146] The prediction module is used to take the text hidden representation and text sequence representation as inputs of the classification module, obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples, and construct the document classification loss and text sequence representation classification loss based on the first and second predicted probabilities respectively;

[0147] A training module that constructs a total loss based on label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and trains the model until convergence;

[0148] The testing module is used to use the test set as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

[0149] An embodiment of the present invention provides a storage medium storing a computer program for classifying hierarchical multi-label professional technical documents using tag information, wherein the computer program enables a computer to execute the hierarchical multi-label professional technical document classification method described above.

[0150] An embodiment of the present invention provides an electronic device, including:

[0151] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs including a method for executing the hierarchical multi-label professional technical document classification method as described above.

[0152] It can be understood that the hierarchical multi-label professional technical document classification system, storage medium and electronic device using tag information provided by the present invention correspond to the hierarchical multi-label professional technical document classification method using tag information provided by the present invention. The explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the hierarchical multi-label professional technical document classification method, and will not be repeated here.

[0153] In summary, compared with the existing technology, the present invention has the following beneficial effects:

[0154] The embodiments of the present invention fully utilize label information for hierarchical multi-label professional technical documents to accurately classify documents. This method can not only capture the semantic representation of the text, but also learn label hierarchy information. Specifically, the label-aware supervised contrastive learning module introduces the definition of label similarity in supervised contrastive learning; a similarity relationship is established between anchor samples and positive samples, making the document representation more accurate. The hierarchical-aware label attention module captures the semantic relationship between the document and the corresponding hierarchical-aware label description, with the aim of dynamically integrating the label hierarchy into the classification model.

[0155] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0156] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A hierarchical multi-label professional technical document classification method using label information, characterized in that: Based on a hierarchical multi-label professional technical document classification model, the model includes a label-aware supervised contrastive learning module, a hierarchical label attention module, and a classification module; the classification method includes: Obtain several document samples and divide them into training sets and test sets; The anchor samples and corresponding positive and negative samples in the training set are used as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss; The anchor samples and their hierarchical multi-labels are used as the input of the label attention module to obtain the text sequence representation with label embedding attention; The text hidden representation and text sequence representation are used as inputs of the classification module to obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples; based on the first and second predicted probabilities, the document classification loss and text sequence representation classification loss are constructed respectively; Construct a total loss based on label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and train the model until convergence; The test set is used as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results; The anchor samples and corresponding positive samples in the training set are used as inputs of the supervised contrastive learning module to obtain the text hidden representation, including: Generate a first string sequence of anchor samples; use the first string sequence as input to the document encoder of the supervised contrastive learning module to obtain a first sequence representation; The first sequence representation is used as the input of the projection network of the supervised contrastive learning module to obtain the text hidden representation; The label-aware supervised contrastive learning loss specifically refers to: Among them, L LSCLM represents the label-aware supervised contrastive learning loss; N represents the number of document samples in the current batch, and let i∈A≡{1,2,…,N} be the index of the anchor sample; represents the label-aware supervised contrastive learning loss of the i-th anchor sample; L i represents the hierarchical multi-label set of the i-th anchor sample, L p represents the hierarchical multi-label set of the pth positive sample corresponding to the i-th anchor sample; P(i) represents the positive sample set corresponding to the i-th anchor sample; z i 、z p 、z a Represents the i-th anchor sample and its positive sample, positive and negative samples respectively; sim represents the similarity function; exp represents the exponential function; τ represents the temperature coefficient; J(X,Y)∈[0,1] represents the Jaccard similarity of any two label sets X and Y. The similarity of two sets is defined as the cardinality of their intersection |X∩Y| divided by the cardinality of their union |X∪Y|.

2. The hierarchical multi-label professional technical document classification method according to claim 1, characterized in that: The anchor sample and its hierarchical multi-label are used as inputs of the label attention module to obtain the text sequence representation of the label embedding attention, including: Generate a second string sequence of anchor samples; use the second string sequence as a document encoder of the label attention module to obtain a second sequence representation; Generate a representation of the hierarchical perceptual path corresponding to the anchor sample; use the representation of the hierarchical perceptual path as the input of the document encoder of the label attention module to obtain label features; According to the second sequence representation and label features, a label-embedded attentive text sequence representation is obtained based on the attention mechanism.

3. The hierarchical multi-label professional technical document classification method according to claim 1, characterized in that: The first and second predicted probabilities specifically refer to: p ik =sigmoid(W c z i +b c ) k Among them, p ik 、p ik Represent the first and second predicted probabilities respectively; z i 、 Represents text sequence representation and label embedding attention text sequence representation respectively; sigmoid represents the activation function; W c represents the weight matrix; b c represents the bias matrix; k represents the index of the label.

4. The hierarchical multi-label professional technical document classification method according to claim 3, characterized in that: The document classification loss and text sequence classification loss are respectively: Among them, L C 、 They represent document classification loss and text sequence classification loss respectively; K represents the number of labels; y ik represents the true value; log represents the logarithmic function.

5. The hierarchical multi-label professional technical document classification method according to claim 4, characterized in that: The total loss specifically refers to: Where L represents the total loss and λ represents the hyperparameter.

6. A hierarchical multi-label professional technical document classification system using label information, characterized in that: Based on a hierarchical multi-label professional technical document classification model, the model includes a label-aware supervised contrastive learning module, a hierarchical-aware label attention module, and a classification module; The classification system is used to implement the hierarchical multi-label professional technical document classification method according to claim 1, comprising: The partitioning module is used to obtain several document samples and divide them into training sets and test sets; The first acquisition module is used to take the anchor samples and corresponding positive and negative samples in the training set as the input of the supervised contrastive learning module to obtain the text hidden representation and construct the label-aware supervised contrastive learning loss; The second acquisition module is used to take the anchor sample and its hierarchical multi-label as the input of the label attention module to obtain the text sequence representation of the label embedding attention; The prediction module is used to take the text hidden representation and text sequence representation as inputs of the classification module, obtain the first and second predicted probabilities of the classification results corresponding to the anchor samples, and construct the document classification loss and text sequence representation classification loss based on the first and second predicted probabilities respectively; A training module that constructs a total loss based on label-aware supervised contrastive learning loss, document classification loss, and text sequence representation classification loss and trains the model until convergence; The testing module is used to use the test set as the input of the converged model to obtain the hierarchical multi-label professional technical document classification results.

7. A storage medium, characterized in that: It stores a computer program for classifying hierarchical multi-label professional technical documents using label information, wherein the computer program enables a computer to execute the hierarchical multi-label professional technical document classification method according to any one of claims 1 to 5.

8. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing the hierarchical multi-label professional technical document classification method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-label text classification method and system based on comparative learning

    CN116450823A

  • Method for identifying vulnerabilities in computer program code and a system thereof

    US20230036159A1