Domain Long Text Classification Method and System Based on Knowledge Graph
By introducing knowledge graphs and graph neural networks into domain long text classification, combined with the BERT model, the problems of semantic information processing and domain knowledge in long text classification are solved, and the classification accuracy is significantly improved.
Patent Information
- Application Number
- CN202310624760.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-05-30
AI Technical Summary
The existing field long text classification technology is difficult to effectively process semantic information in long texts, and lacks expert domain knowledge, resulting in low classification accuracy.
Using a knowledge graph-based method, the text is encoded using the BERT model, combined with the GCN model and the knowledge graph, semantic information is extracted through entity relationships and dependency graph neural networks, and edge relationships and edge types are learned through the graph structure mask model, and knowledge features and data features are fused.
It improves the accuracy of field-long text classification, can express and utilize text semantic information more effectively, and enhances the model's learning ability on semantic dependence between words.
Smart Images

Figure CN116521882B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text classification, and particularly relates to a domain long text classification method and system based on a knowledge graph. Background Art
[0002] At present, artificial intelligence has developed rapidly and achieved great achievements in many fields, such as natural language processing, image processing, data mining, etc. Text mining is one of the research directions. According to the definition of Wikipedia, text mining, also known as text data mining or text analysis, is the process of obtaining high-quality information from text. Typical tasks include text classification, automatic question answering, sentiment analysis, machine translation, etc. Text classification is to divide data into predefined categories. The general process is as follows: 1. Preprocessing, such as word segmentation and removing stop words; 2. Text representation and feature selection; 3. Classifier construction; 4. The classifier classifies according to the features of the text; 5. Evaluation of the classification results. Due to the rapid development of artificial intelligence in recent years, text classification technology can already well determine the category of an unknown document, and the accuracy is also good. With the help of text classification, it is convenient to process massive information and save a large amount of information processing costs. It is widely used in social life such as information filtering, organization and management of information, digital libraries, spam filtering, etc.
[0003] Text classification is one of the most basic tasks in natural language processing and is the basis for many tasks such as recommendation tasks, question answering systems, and sentiment analysis. With the geometric growth of text data volume, the research on text classification tasks, which are the basis for many applications, has become increasingly important. In recent years, domain long text classification has received extensive attention from researchers and made great progress. However, there are the following three problems in the related research of existing domain long text classification technologies: 1) The text length is long. Due to the long text length, a large number of key information is scattered, and long sequence processing is likely to ignore the rich semantic information in the text hierarchical structure. 2) Lack of expert domain knowledge. The pre-trained model requires certain prior knowledge, and relying only on text similarity will cause problems of inaccurate classification. 3) The classification accuracy is not high. Since the text in the data set exists in various fields, and the lengths, features are different, the robustness of the existing model still needs to be improved, resulting in deviation of the classification results.
[0004] In summary, text classification is a fundamental technology in current natural language processing. Machine learning and deep learning have been extensively studied and made great progress in this task. However, these traditional methods can only process data in Euclidean space and cannot fully and effectively express the semantic information of text. Most existing long text classification studies focus on the "data" level, that is, extracting deep semantic representations of Chinese text through deep learning models to obtain context information. This approach is difficult to solve the problem of polysemy of words in different contexts, resulting in problems such as fuzzy semantic understanding and insufficient feature expression, which affect the classification accuracy. Summary of the Invention
[0005] To solve the problem of polysemy of words in different contexts in the prior art, the present invention proposes a method and system for classifying long domain texts based on a knowledge graph. At the data level, BERT is used to train dynamic word vectors to enrich semantic information; at the knowledge level, prior knowledge is introduced through the knowledge graph; by fusing knowledge features and data features, the accuracy of classifying long domain texts is improved.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a method for classifying long domain texts based on a knowledge graph, comprising the following steps:
[0008] Use the BERT model to encode the input text to obtain an initial vector containing rich semantic information, and use the initial vector corresponding to each word as a node of the GCN model. The GCN model includes two GCN modules, namely an entity relationship graph neural network and a dependency relationship graph neural network;
[0009] Use the trained entity relationship extraction model to extract entity information and relationship information between entities in the text to construct a knowledge graph; use a syntactic dependency tool to automatically process the text and generate a syntactic dependency tree, and construct a dependency relationship graph on the syntactic dependency tree;
[0010] Input the knowledge graph and the dependency relationship graph into the entity relationship graph neural network and the dependency relationship graph neural network respectively. In the GCN module, for each word, use its entity relationship type or dependency relationship type with relevant context words as context features for encoding; at the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with additional entity relationship type information, and the dependency relationship graph neural network outputs a word vector with additional dependency relationship type information, and fuse the initial word vector with the word vector with additional edge type information to obtain the final word vector;
[0011] Randomly mask the connection of the edges between two nodes using the graph structure masking model, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of the connection. Finally, obtain the entity relationship type vector and the dependency relationship type vector respectively, and splice the two vectors to obtain the edge type vector;
[0012] Adopt the cross-entropy loss function to optimize the model, and obtain the classification probability through the softmax function to achieve the classification of long domain texts.
[0013] Furthermore, for the BERT representation of long texts, use the sliding window method to intercept different parts of a long text through the sliding window, and then sum and average all the sentence representations as the final BERT vector.
[0014] Furthermore, splice the vectors of the last 4 layers of the BERT model, and the output vector of the BERT model is represented as h bert , and the calculation formula is as follows:
[0015] h bert = ReLU(W[h bert,-1 ; h bert,-2 ; h bert,-3 ; h bert,-4 ;]+b)
[0016] Among them, ; represents splicing, h bert,-1 、h bert,-2 、h bert,-3 、h bert,-4 respectively represent the vectors obtained from the last 4 layers of the BERT encoding, W is the trainable weight matrix, and b is the bias term.
[0017] Furthermore, the knowledge graph is a set of triples composed of nodes and edges, and the process of constructing the knowledge graph is as follows:
[0018] First, extract the keywords of the text through the TextRank algorithm, and at the same time cluster the text through the K-Means clustering algorithm. Through the keywords and the clustering results, obtain the preliminary domain key entity types, then refine and modify the entity types and define their entity relationships; finally, use a trained entity relationship extraction model to extract triples according to the defined entity types and entity relationship types.
[0019] Furthermore, the calculation formula for the output vector of the GCN module is as follows:
[0020] Set the initial feature matrix of the nodes where n doc is the number of text nodes, n entity is the number of extracted entities; for a text of length n, construct an adjacency matrix A=(a i,j ) n×n, when there is a syntactic dependency or entity relationship between the word x i and x j , then a i,j = 1, otherwise there is no relationship, then a i,j = 0;
[0021] For any word x i , the output representation of the l-th layer GCN is:
[0022]
[0023] where a i.j ∈ A, is the output of the word x j in the (l - 1)-th layer of the GCN, W (l) is a trainable matrix, b (l) is the bias of the l-th layer GCN, and σ is the activation function ReLU.
[0024] Furthermore, based on the attention mechanism, the entity relationship graph neural network outputs word vectors with entity relationship type information, including:
[0025] Use B = (r i,j ) n×n to represent the entity relationship type matrix, where r i,j is the entity relationship type between x i and x j , and map each type r i,j to its embedding Calculate the weight p (l) i,j between nodes i and j in the l-th layer of the GCN based on the attention mechanism, (l) i,j The calculation formula of p
[0026]
[0027] where, a i.j ∈ A, and are the intermediate vectors of x i and x j respectively, and The calculation formulas of are as follows:
[0028]
[0029]
[0030] where, and are the outputs of nodes i and j at the (l-1)-th layer of the GCN, denotes concatenation;
[0031] The output vector of the final entity-relationship graph neural network is calculated as follows:
[0032]
[0033] is a vector with entity-relationship type information added, and its calculation formula is:
[0034]
[0035] where, embeds the entity-relationship type and maps it to the same dimension as ; is the output of word x j at the (l-1)-th layer of the GCN.
[0036] Furthermore, based on the attention mechanism, the dependency graph neural network outputs word vectors with dependency type information, including:
[0037] Using C=(t i,j ) n×n to represent the dependency type matrix, where t i,j is the dependency type between x i and x j , maps each type t i,j to its embedding Based on the attention mechanism, calculate the weight q (l) i,j of the connection between nodes i and j in the l-th layer of the GCN, q (l) i,j is calculated as follows:
[0038]
[0039] where a i.j ∈A, and are the intermediate vectors of x i and x j respectively, and are calculated as follows:
[0040]
[0041]
[0042] where and They are the outputs of nodes \(i\) and \(j\) at the \((l - 1)\)-th layer of the GCN respectively, denotes concatenation;
[0043] Finally, the output of the dependency graph neural network has the following calculation formula:
[0044]
[0045] is a vector with dependency relation type information added, and the calculation formula is:
[0046]
[0047] where, embeds the dependency relation type and maps it to the same dimension as ; is the output of word \(x\) j at the \((l - 1)\)-th layer of the GCN;
[0048] Finally, the vectors generated by the two GCN modules are concatenated to obtain the output of the GCN model:
[0049]
[0050] Furthermore, the entity relation type vector and the dependency relation type vector are obtained through the graph structure masking model respectively, and the expressions are as follows:
[0051] h rel,edge,i = ReLU(W[h rel,gcn,i ; h rel,gcn,j +b)
[0052] h dep,edge,i = ReLU(W[h dep,gcn,i ; h dep,gcn,j +b)
[0053] where, \(h\) rel,edge,i is the entity relation type vector output by the graph structure masking training, \(h\) dep,edge,i is the dependency relation type vector output by the graph structure masking training, \(h\) rel,gcn,i ; h rel,gcn,j represents the entity relation edge between nodes \(i\) and \(j\), and \(h\) dep,gcn,i ; h dep,gcn,j represents the dependency relation edge between nodes \(i\) and \(j\);
[0054] The two parts of the masking training results are concatenated to obtain \(h\) edge :
[0055]
[0056] This embodiment also provides a domain long text classification system based on a knowledge graph, including an initial word vector obtaining module, a knowledge graph and dependency graph construction module, a final word vector obtaining module, an edge type vector obtaining module, and a model optimization module, where:
[0057] The initial word vector obtaining module is used to encode the input text using the BERT model to obtain an initial vector containing rich semantic information, and use the initial vector corresponding to each word as a node of the GCN model. This GCN model includes two GCN modules, namely an entity relationship graph neural network and a dependency relationship graph neural network;
[0058] The knowledge graph and dependency graph construction module is used to extract entity information and relationship information between entities in the text using a trained entity relationship extraction model to construct a knowledge graph; automatically process the text using a syntactic dependency tool and generate a syntactic dependency tree, and construct a dependency graph on the syntactic dependency tree;
[0059] The final word vector obtaining module is used to input the knowledge graph and the dependency graph into the entity relationship graph neural network and the dependency graph neural network respectively. In the GCN module, for each word, use its entity relationship type or dependency relationship type with related context words as context features for encoding; at the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with increased entity relationship type information, and the dependency graph neural network outputs a word vector with increased dependency relationship type information, and fuse the initial word vector with the word vector with increased edge type information to obtain the final word vector;
[0060] The edge type vector obtaining module is used to randomly mask the connection of the edge between two nodes using a graph structure masking model, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of the connection, and finally obtain an entity relationship type vector and a dependency relationship type vector respectively, and splice the two vectors to obtain an edge type vector;
[0061] The model optimization module is used to optimize the model using the cross-entropy loss function and obtain the classification probability through the softmax function to achieve domain long text classification.
[0062] Compared with the prior art, the present invention has the following advantages:
[0063] 1. The present invention is based on a number of cutting-edge technologies such as the BERT model, the GCN model, and the knowledge graph. First, the BERT model is used to encode the input text to obtain an initial vector containing rich semantic information, and the initial vector corresponding to each word is used as a node of the graph neural network. Second, the trained entity relation extraction model is used to extract entity information and relationship information between entities in the text, and together with the syntactic dependency information, they are used as the edges of the graph neural network. The document representation of BERT and the document vector of GCN are jointly used to calculate the loss and backpropagate. To further improve the model's learning ability for semantic dependencies between words, a graph structure masking model is used to learn edge relationships and edge types. Finally, the softmax function is used to obtain classification probabilities to achieve domain long text classification. The present invention improves the accuracy of domain long text classification by integrating knowledge features and data features.
[0064] 2. The present invention is based on entities and entity relationships in the knowledge graph, which jointly support the composition of the graph convolutional neural network and the iterative update of nodes in the graph, better combining knowledge graph information with the deep learning model, driving data with knowledge, and achieving improved performance in domain long text classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0066] Figure 1 is a flowchart of the domain long text classification method based on the knowledge graph according to the embodiment of the present invention;
[0067] Figure 2 is a structural block diagram of the domain long text classification system based on the knowledge graph according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0069] As Figure 1 shown, the domain long text classification method based on the knowledge graph in this embodiment includes the following steps:
[0070] Step S101: Encode the input text using the BERT model to obtain an initial vector containing rich semantic information. Take the initial vector corresponding to each word as a node of the GCN model, which includes two GCN modules, namely the entity relationship graph neural network and the dependency relationship graph neural network.
[0071] Step S102: Use the trained entity relationship extraction model to extract entity information and relationship information between entities in the text, and construct a knowledge graph; use a syntactic dependency tool to automatically process the text and generate a syntactic dependency tree (the text is processed using the open-source NLP tool Stanza from Stanford University to obtain the syntactic dependency tree), and construct a dependency relationship graph on the syntactic dependency tree.
[0072] Step S103: Input the knowledge graph and the dependency relationship graph into the entity relationship graph neural network and the dependency relationship graph neural network respectively. In the GCN module, for each word, encode the entity relationship type or dependency relationship type with its related context words as context features. At the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with additional entity relationship type information, and the dependency relationship graph neural network outputs a word vector with additional dependency relationship type information. Fuse the initial word vector in Step S101 with the word vector with additional edge type information to obtain the final word vector.
[0073] Step S104: The entity relationship graph neural network and the dependency relationship graph neural network use the graph structure masking model to randomly mask the connection between two nodes, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of the connection. Finally, obtain the entity relationship type vector and the dependency relationship type vector respectively, and splice the two vectors to obtain the edge type vector.
[0074] Step S105: Optimize the model using the cross-entropy loss function, and obtain the classification probability through the softmax function to achieve domain long text classification.
[0075] The vectorization of the long text is specifically as follows:
[0076] As an important part of modern NLP tasks, compared with embeddings learned from scratch, pre-trained word embeddings (such as word2vec and GloVe) can significantly improve the performance of NLP tasks. For each word in the vocabulary, although the meanings of these words may vary in different scenarios, these context-free models still produce only a single word embedding representation. Context models, such as OpenAI GPT, ELMo, and BERT, can generate different representations for the same word in different contexts according to the context (i.e., the surrounding words in the sentence). These context models usually contain more hidden layers and improve the model performance by training on a large amount of unlabeled data. When applied to specific domain tasks, only a small amount of labeled data is needed for fine-tuning. After the emergence of BERT, as a replacement for Word2Vec, it has significantly refreshed the accuracy in 11 directions in the NLP field, such as text classification, question answering, and language inference. Through large-scale pre-training, BERT can give words a semantically rich initial vector. Use BERT to train word vectors and use them as the nodes of GCN to avoid the problem that a word can only have one vector representation, thus learning dynamic word vectors.
[0077] The longest text in long text data can reach thousands of words, while the maximum length limit of BERT input is 512, including the CLS token. Therefore, special processing is required for the BERT representation of long texts. General processing can directly truncate the long text and only take a part of the beginning or the end, which may lead to the omission of key information. In this example, the sliding window method is adopted. Different parts of a long text are intercepted through the sliding window, and then the representations of all sentences are summed and averaged as the final BERT vector.
[0078] In addition to using the BERT output as the GCN node, referring to the prediction interpolation method of BertGCN, the document representation of BERT and the document vector of GCN are jointly used to calculate the loss and backpropagate. Since GCN has more parameters, the model is prone to overfitting or difficult to converge after being fused with BERT. To solve this problem, the vectors of the last 4 layers of BERT are concatenated to prevent the model from only using the output of the last layer of BERT and having too many parameter calculations. The output vector representation of the BERT model is h bert , and the calculation formula is as follows:
[0079] h bert = ReLU(W[h bert,-1 ; h bert,-2 ; h bert,-3 ; h bert,-4 ; ]+b) (1)
[0080] Among them, ; represents concatenation, h bert,-1 、h bert,-2, h bert,-3 , h bert,-4 respectively represent the vectors obtained from the last 4 layers of BERT encoding, W is a trainable weight matrix, and b is a bias term.
[0081] The process of constructing the knowledge graph is as follows:
[0082] Pre-trained language models (such as BERT, word2vec, and GloVe) lack common sense or domain-specific knowledge, which usually leads to unsatisfactory performance in text feature representation. To solve these two problems, a knowledge graph is constructed based on unclassified documents, taking its entities and entity relationships as the nodes and edges of the graph neural network to better achieve text classification. Referring to the top-down construction method of the domain knowledge graph, based on the classified data sources, ontology and schema information are extracted from high-quality data, and entities and relationships between entities are extracted to form domain prior knowledge in the form of a knowledge graph.
[0083] A knowledge graph is a set of triples composed of nodes and edges. The nodes are the entities existing in the sentences, and the edges are the relationships between entities. First, the keywords of the text are extracted through the TextRank algorithm, and at the same time, the text is clustered through the K-Means clustering algorithm. By manually observing the keywords and clustering results, the preliminary domain key entity types are obtained. Then, the entity types are refined and modified and their entity relationships are defined by using statistical methods combined with manual induction, referring to high-quality general graphs, and expert guidance. Finally, a trained entity relationship extraction model is used to extract triples according to the defined entity types and entity relationship types.
[0084] Graph neural network based on attention mechanism:
[0085] In the classical GCN, the connection relationships between words are not distinguished. If there is a connection between two nodes, the element in the adjacency matrix is 1, otherwise it is 0. Therefore, the GCN model cannot distinguish the importance of different connections, nor can it reflect the differences in entity relationships and dependency relationships between nodes. To fully learn and utilize the knowledge graph information, two GCN modules are used, one to process entity relationship type information and the other to process dependency relationship type information; and the original simple adjacency matrix of GCN is improved by using the attention mechanism to increase edge type features.
[0086] First, set the initial feature matrix of the node X = I ndoc+nentity , where n doc is the number of text nodes, n entity is the number of extracted entities, and use BERT to obtain node embeddings; for a text of length n, construct an adjacency matrix A = (a i,j ) n×n , when the word x i and x jIf there is a syntactic dependency or entity relationship between them, then a i,j = 1; conversely, if there is no relationship, then a i,j = 0; for any word x i , the output representation of the l-th layer of GCN is:
[0087]
[0088] where a i.j ∈ A, is the output of word x j in the (l - 1)-th layer of GCN, W (l) is a trainable matrix, b (l) is the bias of the l-th layer of GCN, and σ is the activation function ReLU.
[0089] Then use B = (r i,j ) n×n to represent the entity relationship type matrix, where r i,j is the entity relationship type between x i and x j . Map each type r i,j to its embedding Calculate the weight p (l) i,j between nodes i and j in the l-th layer of GCN based on the attention mechanism, and the calculation formula of p (l) i,j is as follows:
[0090]
[0091] where, a i.j ∈ A, and are the intermediate vectors of x i and x j respectively, and the calculation formulas of and are as follows:
[0092]
[0093]
[0094] where, and are the outputs of nodes i and j in the (l - 1)-th layer of GCN respectively, denotes concatenation;
[0095] The calculation formula of the final output vector of the entity relationship graph neural network is as follows:
[0096]
[0097] It is a vector with entity relationship type information added, and the calculation formula is:
[0098]
[0099] Among them, Embed the entity relationship type Map to the same dimension as and is the output of word x j at the (l - 1)-th layer of the GCN.
[0100] Finally, use C=(t i,j ) n×n to represent the dependency relationship type matrix, where t i,j is the dependency relationship type between x i and x j . Map each type t i,j to its embedding Calculate the weight q (l) i,j of the connection between nodes i and j in the l-th layer of the GCN based on the attention mechanism. The calculation formula of q (l) i,j is as follows:
[0101]
[0102] Among them, a i.j ∈A, and are the intermediate vectors of x i and x j respectively. The calculation formulas of and are as follows:
[0103]
[0104]
[0105] Among them and are the outputs of nodes i and j at the (l - 1)-th layer of the GCN respectively. denotes concatenation;
[0106] The calculation formula of the final output of the dependency relationship graph neural network is as follows:
[0107]
[0108] is a vector with additional dependency relation type information, and the calculation formula is:
[0109]
[0110] where embeds the dependency relation type maps to the same dimension as and is the output of word x j at the (l - 1)-th layer of the GCN;
[0111] Finally, the vectors generated by the two GCN modules are concatenated to obtain the output of the GCN model:
[0112]
[0113] The model's understanding of semantic dependencies is further enhanced through the graph structure masking model, specifically:
[0114] The Masked Language Model (MLM) is one of the training tasks of the BERT model. Different from general language models, it does not need to predict all texts like autoregressive models, but randomly masks certain words in the sentence and uses the context of the masked words to predict the word. This model is called the Autoencoder Language Model (Autoencoder LM). The BERT model trained through MLM has the ability to learn, establish better connections between words and texts, and output different dynamic word vectors in different context environments, which is also one of the reasons for improving the performance of the BERT model.
[0115] Based on the MLM idea, in order to further improve the model's learning ability for semantic dependencies between words and characters, the graph structure masking model is proposed. Similar to the masked language model, the graph structure masking model randomly masks the connection of edges between two nodes, allowing the graph structure masking model to predict whether there is a connection relationship between the two nodes and the type of connection. Finally, the entity relationship type vector and the dependency relationship type vector are obtained respectively:
[0116] h rel,edge,i = ReLU(W[h rel,gcn,i ; h rel,gcn,j +b) (14)
[0117] h dep,edge,i = ReLU(W[h dep,gcn,i ; h dep,gcn,j +b) (15)
[0118] where h rel,edge,i is the entity relationship type vector output by the graph structure masking training, and h dep,edge,iIt is the dependency relationship type vector of the graph structure mask training output, h rel,gcn,i ; h rel,gcn,j represents the entity relationship edge between nodes i and j, h dep,gcn,i ; h dep,gcn,j represents the dependency relationship edge between nodes i and j.
[0119] Concatenate the two parts of the mask training results to obtain h edge :
[0120]
[0121] The loss function is specifically:
[0122] The cross-entropy loss function is a commonly used loss function in deep learning. It has strong generalization ability and good convex optimization properties, which can help the model training converge, so that the prediction results of the model are close to the actual label values. Using the cross-entropy loss as the loss function, the BERT prediction probability distribution and the loss function L bert :
[0123]
[0124]
[0125] where y i is the true label of a certain sample belonging to class i, and M is the total number of labels.
[0126] The graph convolutional neural network prediction probability distribution and the loss function L gcn :
[0127]
[0128]
[0129] The linear interpolation calculation formula of BERT and GCN is as follows:
[0130] L cls = λ * L gcn + (1 - λ) * L bert (21)
[0131] where λ controls the trade-off between the two objectives. λ = 0 means only using the BERT model, and λ = 1 means only using the GCN module.
[0132] The graph structure mask module prediction probability distribution and the loss function L edge :
[0133]
[0134]
[0135] The loss function L of the final model is as follows:
[0136] L = (1 - α)L cls + αL edge (24)
[0137] Where α can adjust the weights of the graph structure masking module to further optimize the model performance.
[0138] Corresponding to the above domain long text classification method based on knowledge graph, as Figure 2 shown, this example also proposes a domain long text classification system based on knowledge graph, including an initial word vector obtaining module, a knowledge graph and dependency graph construction module, a final word vector obtaining module, an edge type vector obtaining module and a model optimization module, where:
[0139] The initial word vector obtaining module is used to encode the input text by using the BERT model to obtain an initial vector containing rich semantic information, and use the initial vector corresponding to each word as the node of the GCN model. This GCN model contains two GCN modules, namely the entity relationship graph neural network and the dependency relationship graph neural network.
[0140] The knowledge graph and dependency graph construction module is used to extract entity information and relationship information between entities in the text by using the trained entity relationship extraction model to construct a knowledge graph; use a syntactic dependency tool to automatically process the text and generate a syntactic dependency tree, and construct a dependency graph on the syntactic dependency tree.
[0141] The final word vector obtaining module is used to input the knowledge graph and the dependency graph into the entity relationship graph neural network and the dependency relationship graph neural network respectively. In the GCN module, for each word, its entity relationship type or dependency relationship type with relevant context words is used as the context feature for encoding; at the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with increased entity relationship type information, and the dependency relationship graph neural network outputs a word vector with increased dependency relationship type information, and the initial word vector is fused with the word vector with increased edge type information to obtain the final word vector.
[0142] The edge type vector obtaining module is used to randomly mask the connection between two nodes by using the graph structure masking model, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of the connection. Finally, an entity relationship type vector and a dependency relationship type vector are obtained respectively, and the two vectors are concatenated to obtain an edge type vector.
[0143] A model optimization module, which is used to optimize the model by using the cross-entropy loss function and obtain classification probabilities through the softmax function to achieve domain long text classification.
[0144] In the data aspect of the present invention, BERT is used to encode the text to obtain an initial vector containing rich semantic information; in the knowledge aspect, prior knowledge is introduced through a knowledge graph, and a trained entity relationship extraction model is used to extract entity information and the relationship information between entities in the text, and they are used as the edges of the graph neural network together with the syntactic dependency information. To further improve the model's learning ability for semantic dependencies between words and characters, a graph structure masking model is used to learn edge relationships and edge types. The present invention further improves the accuracy of long text classification by fusing knowledge features and data features.
[0145] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0146] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0147] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or multiple processes and / or one block or multiple blocks in the flow Figure 1 One process or multiple processes and / or Figure 1 One block or multiple blocks.
[0149] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0150] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for classifying long texts in a domain based on a knowledge graph, characterized in that, It includes the following steps: Step 1: Use the BERT model to encode the input text to obtain an initial vector containing rich semantic information. Take the initial vector corresponding to each word as the node of the GCN model. The GCN model includes two GCN modules, namely the entity relationship graph neural network and the dependency relationship graph neural network; Step 2: Use the trained entity relationship extraction model to extract entity information and relationship information between entities in the text, and construct a knowledge graph; Use a syntactic dependency tool to automatically process the text and generate a syntactic dependency tree, and construct a dependency relationship graph on the syntactic dependency tree; Step 3: Input the knowledge graph and the dependency relationship graph into the entity relationship graph neural network and the dependency relationship graph neural network respectively. In the GCN module, for each word, use its entity relationship type or dependency relationship type with related context words as context features for encoding; at the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with increased entity relationship type information, and the dependency relationship graph neural network outputs a word vector with increased dependency relationship type information. Fuse the initial word vector and the word vector with increased edge type information to obtain the final word vector; specifically, it includes: The calculation formula for the output vector of the GCN module is as follows: Set the initial feature matrix of nodes where n doc is the number of text nodes, and n entity is the number of extracted entities; for a text of length n, construct an adjacency matrix A=(a i,j ) n×n , when there is a syntactic dependency or entity relationship between words x i and x j , then a i,j =1, otherwise there is no relationship, then a i,j =0; There exists any word x i , then the output representation of the l-th layer GCN is: where a i.j ∈ A, is the output of the word x j at the (l - 1)-th layer of the GCN, W (l) is a trainable matrix, b (l) is the bias of the l-th layer of the GCN, and σ is the activation function ReLU; Based on the attention mechanism, the entity relationship graph neural network outputs a word vector with increased entity relationship type information, including: Use B=(r i,j ) n×n to represent the entity relationship type matrix, where r i,j is the entity relationship type between x i and x j . Map each type r i,j to its embedding Calculate the weight p (l) i,j of the connection between nodes i and j in the l-th layer of the GCN based on the attention mechanism. The calculation formula of p (l) i,j is as follows: where a i.j ∈A, and are the intermediate vectors of x i and x j respectively, and the calculation formulas of and are as follows: Among them, and are the outputs of nodes i and j at the (l-1)-th layer of the GCN respectively, denotes concatenation; Final entity relationship diagram neural network output vector The calculation formula is as follows: is a vector with entity relationship type information added, and the calculation formula is: Among them, Embed the entity relationship type Map to the same dimension as is the output of word x j at the output of the (l - 1)-th layer of the GCN; Based on the attention mechanism, the dependency relationship graph neural network outputs a word vector with increased dependency relationship type information, including: Use C=(t i,j ) n×n to represent the dependency relationship type matrix, where t i,j is the dependency relationship type between x i and x j . Map each type t i,j to its embedding Calculate the weight q (l) i,j of the connection between nodes i and j in the l-th layer of the GCN based on the attention mechanism. The calculation formula of q (l) i,j is as follows: where a i.j ∈ A, and are the intermediate vectors of x i and x j respectively, and the calculation formulas of and are as follows: where and are the outputs of nodes \(i\) and \(j\) at the \((l - 1)\)-th layer of the GCN respectively, denotes concatenation; Final Dependency Graph Neural Network Output The calculation formula is as follows: is a vector with dependency relationship type information added, and the calculation formula is: Among them, Embed the dependency type Map to the same dimension as is the output of word x j at the output of the (l-1)-th layer of the GCN; Finally, splice the vectors generated by the two GCN modules to obtain the output of the GCN model: Step 4: Use the graph structure masking model to randomly mask the connection between two nodes, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of connection. Finally, obtain the entity relationship type vector and the dependency relationship type vector respectively, and splice the two vectors to obtain the edge type vector; Step 5: Use the cross-entropy loss function to optimize the model, and obtain the classification probability through the softmax function to achieve domain long text classification.
2. The method for classifying long texts in a domain based on a knowledge graph according to claim 1, characterized in that, For the BERT representation of long text, use the sliding window method to intercept different parts of a long text through the sliding window, and then sum and average all the sentence representations as the final BERT vector.
3. The method for classifying long domain texts based on a knowledge graph according to claim 2, wherein Concatenate the vectors of the last 4 layers of the BERT model, and the output vector of the BERT model is denoted as h bert , and the calculation formula is as follows: h bert = ReLU(W[h bert,-1 ; h bert,-2 ; h bert,-3 ; h bert,-4 ; ] + b) Among them, ; represents concatenation, h bert,-1 , h bert,-2 , h bert,-3 , h bert,-4 respectively represent the vectors obtained from the last 4 layers of BERT encoding, W is a trainable weight matrix, and b is a bias term.
4. The method for classifying long domain texts based on a knowledge graph according to claim 1, wherein The knowledge graph is a triple composed of a set of nodes and edges. The process of constructing the knowledge graph is: First, extract the keywords of the text through the TextRank algorithm, and at the same time cluster the text through the K-Means clustering algorithm. Through the keywords and clustering results, obtain the preliminary domain key entity types, then refine and modify the entity types and define their entity relationships; finally, use a trained entity relationship extraction model to extract triples according to the defined entity types and entity relationship types.
5. The method for classifying long domain texts based on a knowledge graph according to claim 1, wherein Obtain the entity relationship type vector and the dependency relationship type vector respectively through the graph structure masking model. The expressions are as follows: h rel,edge,i = ReLU(W[h rel,gcn,i ; h rel,gcn,j + b) h dep,edge,i = ReLU(W[h dep,gcn,i ; h dep,gcn,j +b) Among them, h rel,edge,i is the entity relationship type vector output by the graph structure mask training, h dep,edge,i is the dependency relationship type vector output by the graph structure mask training, h rel,gcn,i ; h rel,gcn,j represents the entity relationship edge between nodes i and j, h dep,gcn,i ; h dep,gcn,j represents the dependency relationship edge between nodes i and j; Concatenate the training results of the two parts of the mask to obtain h edge :
6. A system for classifying long domain texts based on a knowledge graph, wherein A system for implementing the domain long text classification method based on a knowledge graph according to any one of claims 1-5, the system includes an initial word vector obtaining module, a knowledge graph and dependency graph construction module, a final word vector obtaining module, an edge type vector obtaining module, and a model optimization module, wherein: The initial word vector obtaining module is used to encode the input text by using a BERT model to obtain an initial vector containing rich semantic information, and use the initial vector corresponding to each word as a node of a GCN model. The GCN model includes two GCN modules, namely an entity relationship graph neural network and a dependency relationship graph neural network; The knowledge graph and dependency graph construction module is used to extract entity information and relationship information between entities in the text by using a trained entity relationship extraction model to construct a knowledge graph; use a syntactic dependency tool to automatically process the text and generate a syntactic dependency tree, and construct a dependency graph on the syntactic dependency tree; The final word vector obtaining module is used to input the knowledge graph and the dependency graph into the entity relationship graph neural network and the dependency relationship graph neural network respectively. In the GCN module, for each word, use its entity relationship type or dependency relationship type with related context words as context features for encoding; at the same time, based on the attention mechanism, the entity relationship graph neural network outputs a word vector with increased entity relationship type information, and the dependency relationship graph neural network outputs a word vector with increased dependency relationship type information, and fuse the initial word vector with the word vector with increased edge type information to obtain the final word vector; The edge type vector obtaining module is used to randomly mask the connection of the edges between two nodes by using a graph structure masking model, and let the graph structure masking model predict whether there is a connection relationship between the two nodes and the type of the connection, and finally obtain an entity relationship type vector and a dependency relationship type vector respectively, and splice the two vectors to obtain an edge type vector; The model optimization module is used to optimize the model by using a cross-entropy loss function, and obtain the classification probability through a softmax function to realize the domain long text classification.
7. A computer device, comprising a memory, a processor, and a computer program stored on the memory, wherein The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.