Text association analysis method and system based on LSTM neural network and knowledge graph
By combining LSTM neural network and knowledge graph, text theme vectors with semantic correlation are generated and semantic enhancement is performed, which solves the problem of difficult to understand the relationship between texts and semantic background in the prior art, and achieves more accurate and efficient text theme classification and analysis.
Patent Information
- Application Number
- CN202510267108.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-07
AI Technical Summary
It is difficult for the prior art to fully understand the relationships and semantic backgrounds between text themes in text theme analysis. There are limitations when LSTM is used alone, and the application of knowledge graphs in text theme analysis has failed to fully realize their potential.
Combining the LSTM neural network and knowledge graph, by collecting text topic data from the literature database, preprocessing, and using the LSTM network to generate vectors with semantic correlation, and constructing a text topic knowledge graph for semantic enhancement, and finally classifying and analyzing text topics based on the output of LSTM and knowledge graph.
It significantly improves the model's understanding of the interrelationships, domain background and cross-domain knowledge between texts, can process massive text data, completely eliminate manual intervention, automatically complete the identification and classification of text topics, greatly improve work efficiency, and explore potential cross-domain fields through the relationship between texts.
Smart Images

Figure CN120218076A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text analysis, and specifically relates to a text association method and system based on an LSTM neural network and a knowledge graph. Background Art
[0002] With the development of information technology and the Internet, academic resources have shown explosive growth. Various types of academic content, such as academic literature, research papers, patents, reports, etc., are continuously generated and accumulated. How to efficiently analyze and automatically classify text topics has become an important research topic in the fields of natural language processing and artificial intelligence.
[0003] In recent years, deep learning, especially the long short-term memory network (LSTM), has made remarkable breakthroughs in natural language processing tasks because it can effectively process sequential data and capture long-term and short-term dependencies. By introducing a gating mechanism, the LSTM network can effectively capture semantic information in long texts, overcoming the limitations of traditional methods in processing long texts. Especially in the analysis and classification of academic literature, it can better understand the context relationships in texts. However, when used alone, LSTM still has certain limitations. Especially in text domain analysis tasks, a single text representation method often cannot comprehensively understand the relationships and semantic backgrounds between texts.
[0004] As a technology that represents entities and their relationships through a graph structure, the knowledge graph has demonstrated its powerful capabilities in tasks such as semantic understanding, reasoning, and information retrieval. However, its application in text topic analysis is still relatively single, and its potential in text topic analysis has not been fully exploited. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention proposes a text association analysis method and system based on an LSTM neural network and a knowledge graph, which can more accurately extract text topics from scientific research literature. By combining deep learning and the knowledge graph, it can not only effectively improve the accuracy and efficiency of text topic classification, but also enhance the reasoning ability between texts and the recognition of cross-text relationships.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A text association analysis method based on an LSTM neural network and a knowledge graph, comprising:
[0008] Collecting text data of text topics from a literature database;
[0009] Preprocessing the collected text data of text topics;
[0010] Learn the pre - processed text topic text data through an LSTM network to generate vectors with semantic relevance for the text topic;
[0011] Construct a text topic knowledge graph, and perform semantic enhancement on the vectors with semantic relevance for the text topic based on the constructed text topic knowledge graph;
[0012] Classify and analyze the text topic based on the outputs of the LSTM neural network and the knowledge graph.
[0013] Specifically, the learning of the pre - processed text topic text data through the LSTM network to generate vectors with semantic relevance for the text topic includes:
[0014] Construct an LSTM network, including: a forget gate, an input gate, and an output gate;
[0015] Use a pre - trained word embedding model to convert the pre - processed text topic text data into text topic word vectors;
[0016] Train the constructed LSTM network, optimize the loss function, and obtain a trained LSTM network;
[0017] Input the text topic word vectors into the trained LSTM network, concatenate or perform weighted summation on the output vector of the multi - head attention mechanism, the output vector of the LSTM network at time t, and the output vector of semantic enhancement to obtain vectors with semantic relevance for the text topic.
[0018] Specifically, the construction of the text topic knowledge graph includes:
[0019] Construct a text topic knowledge graph KG, expressed as: KG = {E, G, J}, where E represents the entity set, G represents the relationship set, and J represents the graph structure;
[0020] Starting from any node in the text topic knowledge graph KG and the neighbor nodes of the any node, record all the nodes reached after k jumps determined by the reward function, and update the embedding representation of the any node at the k + 1 layer by combining the semantic relevance vectors of the neighbor nodes of the any node at the kth layer.
[0021] Specifically, the semantic enhancement of the vectors with semantic relevance for the text topic based on the constructed text topic knowledge graph includes:
[0022] Utilize the enhanced text topic knowledge graph and the vectors with semantic relevance for the text topic to establish a text topic context relationship model and construct a context relationship matrix between text topic entities.
[0023] Specifically, the analysis and classification of text topics based on the output of the LSTM neural network and knowledge graph include:
[0024] Using a support vector machine model, classify the text topic according to the output of the LSTM neural network and knowledge graph;
[0025] Analyze the classification results of the text topic, combine the entity and relationship information in the text topic knowledge graph, reason about the relationships between texts, and calculate the probability of cross-text topics.
[0026] Specifically, the preprocessing of the collected text topic text data includes:
[0027] Remove stop words from the text topic text data, standardize the terms, and perform part-of-speech reduction and word segmentation;
[0028] Based on text domain knowledge and corpus, construct a text feature word library containing common terms, keywords and their semantic associations in this text domain.
[0029] Specifically, the literature database includes: CNKI, Google Scholar and PubMed.
[0030] A text association analysis system based on the LSTM neural network and knowledge graph, used to implement the text association analysis method based on the LSTM neural network and knowledge graph, including: a data collection module, a data preprocessing module, a vector generation module, a semantic enhancement module and a classification and analysis module;
[0031] The data collection module is used to collect text data of text topics from the literature database;
[0032] The data preprocessing module is used to preprocess the collected text data of text topics;
[0033] The vector generation module is used to learn the preprocessed text data of text topics through the LSTM network and generate vectors with semantic relevance for text topics;
[0034] The semantic enhancement module is used to construct a text topic knowledge graph and semantically enhance the vectors with semantic relevance for text topics based on the constructed text topic knowledge graph;
[0035] The classification and analysis module is used to classify and analyze text topics based on the output of the LSTM neural network and knowledge graph.
[0036] Specifically, the vector generation module includes: an LSTM network construction unit, a training unit and a vector generation unit;
[0037] The LSTM network construction unit is used to construct an LSTM neural network model;
[0038] The training unit is used to train the constructed LSTM network and optimize the loss function;
[0039] The vector generation unit is used to convert the preprocessed text topic text data into text topic word vectors and input them into the trained LSTM network to generate vectors with semantic relevance for the text topic.
[0040] Specifically, the semantic enhancement module includes: a knowledge graph construction unit, an enhancement unit, and an association unit;
[0041] The knowledge graph construction unit is used to construct a text topic knowledge graph;
[0042] The enhancement unit is used to enhance the expression ability of the text topic knowledge graph by using a graph convolutional neural network;
[0043] The association unit is used to establish a text topic context relationship model and construct a context relationship matrix between text topic entities.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] 1. The present invention proposes a text association analysis method based on an LSTM neural network and a knowledge graph. By embedding entity and relationship information in the knowledge graph into the LSTM model, it can significantly improve the model's understanding of the mutual relationship, domain background, and cross-domain knowledge between texts and can handle complex situations of text intersection.
[0046] 2. The present invention proposes a text association analysis method based on an LSTM neural network and a knowledge graph, which can process a large amount of text data, has strong adaptability, and can completely eliminate manual intervention, automatically complete the recognition and classification of text topics, and greatly improve work efficiency.
[0047] 3. The present invention proposes a text association analysis method based on an LSTM neural network and a knowledge graph. The knowledge graph provides rich relationships between text domains. After combining with a deep learning model, it can perform higher-level reasoning. Through model analysis, not only can accurate classification be performed, but also potential cross-text domains can be mined through the relationships between texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of the method provided by the present invention;
[0049] Figure 2 is an architecture diagram of the LSTM network provided by the present invention;
[0050] Figure 3 An example diagram of the text topic knowledge graph provided by the present invention;
[0051] Figure 4 The system architecture diagram provided by the present invention. Specific implementation manners
[0052] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0053] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] It should be noted that if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the protection scope of the present application. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. In addition, the terms "first", "second", "third", etc. used in the present application do not limit the data and execution order, but only distinguish the same items or similar items with basically the same function and role.
[0055] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in this specification in the description of the present application are only for the purpose of describing specific embodiments and are not used to limit the present application. The term "and / or" used in this specification includes any and all combinations of one or more of the listed items.
[0056] Embodiment 1
[0057] The method of the present application is applied to the correlation analysis of related discipline topics.
[0058] Please refer to Figures 1-3 , an embodiment provided by the present invention: a text correlation analysis method based on an LSTM neural network and a knowledge graph, including the following specific steps:
[0059] Step S1: Collect text data of text topics from literature databases (such as CNKI, Google Scholar, PubMed, etc.);
[0060] Step S2: Preprocess the collected text topic text data;
[0061] The specific steps of Step S2 are as follows:
[0062] Step S201: Remove stop words from the text topic text data, standardize terms, and perform part-of-speech reduction and word segmentation;
[0063] Specifically, remove stop words, such as irrelevant words like "of", "and", "is", etc.; standardize terms, for example, "deep learning" and "deep neural network" can be unified as "deep learning";
[0064] Step S202: Based on text domain knowledge and a corpus, construct a text feature word library that contains common terms, keywords, and their semantic associations within this text domain.
[0065] Step S3: Use an LSTM network to learn the preprocessed text topic text data and generate vectors with semantic relevance for the text topic;
[0066] Step S301: Construct an LSTM network, including: a forget gate, an input gate, and an output gate;
[0067] The forget gate is used to control the proportion of information discarded, the input gate is used to control the addition of new information, and the output gate is used to determine the output of the hidden state at the current moment; through this mechanism, LSTM can learn complex context dependencies in the text, especially semantic information over long time spans;
[0068] Step S302: Use a pre-trained word embedding model to convert the preprocessed text topic text data into text topic word vectors;
[0069] The pre-trained word embedding model in Step S302 includes: Word2Vec, GloVe, or BERT, etc. Word2Vec and GloVe learn dense vector representations of words based on the context window of words and are trained by minimizing the conditional probability of context vocabulary and target vocabulary. BERT is a pre-trained model based on Transformer that can dynamically generate representations of vocabulary according to the context in a sentence;
[0070] Step S303: Train the constructed LSTM network, optimize the loss function, and obtain a trained LSTM network;
[0071] In this embodiment, the commonly used loss functions are the cross-entropy loss function (for classification tasks) or the mean squared error loss function (for regression tasks). When training long sequences, LSTM can effectively alleviate the gradient vanishing and gradient explosion problems of traditional RNNs, but gradient clipping still needs to be performed to avoid excessive parameter updates. To prevent overfitting, a Dropout layer can be introduced into the LSTM network to randomly discard some neurons to increase the generalization ability of the model;
[0072] Step S304: Input the text topic word vector into the trained LSTM network to obtain a vector with semantic relevance for the text topic. The specific formula is:
[0073]
[0074] where, represents the vector with semantic relevance for the text topic, MultH(Atten(Q, V, K)) represents the output vector of the multi-head attention mechanism, Q represents the query vector, that is, the query vector at the current moment, K represents the key vector, the key value vector used to match the query vector, V represents the value vector, that is, the value information associated with each key, LSTM(x t , h t-1 ) represents the output vector of the LSTM network at time t, x t represents the input vector, that is, the text topic word vector, h t-1 represents the hidden state of the LSTM network at time t-1, EeEn(h t ) represents the output vector of semantic enhancement, h t represents the hidden state of the LSTM network at time t, represents vector concatenation or weighted summation.
[0075] Explanation and principle of the above formula: The multi-head self-attention mechanism MultH(Atten(Q, V, K)) is one of the core mechanisms of the Transformer architecture. It can calculate attention weights in parallel in different subspaces to capture more context information. The specific formula is:
[0076]
[0077] Concat() represents the concatenation function used to concatenate multiple vectors, sl represents the number of attention heads, sm() represents the function to convert the similarity value into a weight distribution, represents the similarity between each query and all keys, T represents the transpose of the vector, and the query vector Q is generated from the input x t and / or the hidden state h t-1 at the previous moment. Usually Let \(R\) denote the set of real numbers, \(n\) denote the length of the input sequence, and \(d\) k denote the dimension of the query vector. The key vector \(K\) has the same dimension as the query vector. d v denote the dimension of the value vector;
[0078] The specific formula for \(EeEn(h\) t ) is: \(EeEn(h\) t ) = \(\tanh(W\) h \times h\) t + W\) e \times e\) t + b)\), where \(\tanh()\) represents the hyperbolic tangent function, \(e\) t represents the external semantic embedding at time \(t\), \(W\) h is the weight matrix used to map \(h\) t to the same dimension, \(W\) e is the weight matrix used to map \(e\) t to the same dimension, and \(b\) represents the bias term. The output vector of semantic enhancement is obtained by further processing the hidden state of the LSTM and using external semantic information to enhance the semantic perception ability of the model. Specifically, by introducing semantic embeddings (such as Word2Vec, GloVe, or BERT embeddings) to strengthen the semantic understanding of the model, the enhanced output vector will contain more semantic information, which can better help the model capture complex text relationships;
[0079] The final output is obtained by concatenating or weighted summing multiple parts. The concatenation or weighted summation of these parts aims to combine local information (LSTM), global information (multi-head self-attention), and external semantic information, thereby generating a more refined and rich representation.
[0080] Step S4: Construct a text topic knowledge graph and perform semantic enhancement on vectors that have semantic relevance to the text topic based on the constructed text topic knowledge graph;
[0081] The specific steps of Step S4 are as follows:
[0082] Step S401: Construct a text topic knowledge graph \(KG\), denoted as: \(KG=\{E, G, J\}\), where \(E\) represents the entity set, containing entities in all domains, \(G\) represents the relationship set, which are various relationships between entities, and \(J\) represents the graph structure, which is the connection between entities;
[0083] In this embodiment, a knowledge graph is a graph structure composed of entities (such as people, places, things, concepts, etc.) and relationships (such as "belong to", "contain", "with...", etc.). Each node represents an entity, and the edge represents a certain relationship between entities. The main role of the knowledge graph is to provide the ability to reason about semantic relationships between entities, and it can supplement implicit knowledge and information in text data. For example, in a knowledge graph of a text domain, the nodes can be text terms such as "deep learning", "convolutional neural network", "computer vision", etc., and the relationships include "is", "contains", "belongs to", etc. "Convolutional neural network" - belongs to - "deep learning", "computer vision" - contains - "convolutional neural network";
[0084] Step S402: For any node in the text topic knowledge graph, enhance the expression ability of the node. The specific formula is:
[0085]
[0086] Where, represents the embedding representation of node v at the k + 1 layer, σ(·) represents the activation function, represents the embedding representation of node u at the k layer, N(v) represents the set of neighbor nodes of node v, N(v) k represents all the nodes reached by k jumps determined by the reward function starting from node v, N(u) k represents all the nodes reached by k jumps determined by the reward function starting from node u, W k represents the weight matrix, which is used for linear transformation of node information;
[0087] The principle of the above formula: The embedding representation of node v at the k + 1 layer is achieved by aggregating the information of the k-hop neighbor nodes of node v. The goal of multi-hop neighbor aggregation is to utilize the information of the k-hop neighbors around node v to gradually enhance the context information of the node representation. The information of the k-hop neighbor nodes of node v is achieved through node jumps, that is, starting from node v, the nodes obtained by the first jump are the nodes connected to node v, the nodes obtained by the second jump are the nodes connected to the nodes obtained by the first jump, and when reaching the kth jump, the nodes obtained are the nodes connected to the nodes obtained by the (k - 1)th jump. The jump direction of the node and the update method of the node embedding representation are determined by the reward function. Aggregating all the nodes obtains the embedding representation N(u) kSimilarly, by selecting an action through the policy network, that is, selecting which neighbor node to jump to and how to update the embedding representation of the current node at the k+1 layer, and then adjusting the parameters of the policy network according to the reward signal feedback from the environment, the optimal jumping path and embedding update policy can be learned to maximize the expressive ability of the enhanced node;
[0088] N(v) k By considering the multi-hop neighbors of nodes, the model can obtain complex relationship information between nodes, not limited to the relationship of one-hop neighbors. The aggregation of this multi-hop information can help capture more complex graph structure features. Especially when dealing with graph structures with long-range dependencies or complex relationships, this aggregation mechanism is particularly important. By normalizing the weights of each neighbor's information, the influence of some nodes with too many connections on the aggregation result can be avoided, ensuring the fairness of each neighbor's information during aggregation;
[0089] The formula describes a process of multi-hop information aggregation, aiming to gradually introduce the high-order neighbor information of nodes (i.e., neighbors after multiple jumps) into the node representation. Among many nodes, through the reward function, the jumping direction of the node is determined. There are some nodes with relatively small values that need to be screened to select the optimal nodes, thereby enhancing the expressive ability of the node representation.
[0090] Step S403: Use the enhanced text topic knowledge graph and vectors with semantic relevance to the text topic to establish a text topic context relationship model and construct a context relationship matrix between text topic entities. The specific formula is:
[0091]
[0092] where S represents the context relationship matrix between text topic entities, and rel(s p ,s q ) represents the semantic relationship between entity s p and entity s q .
[0093] The principle of the above formula: Matrix S is a square matrix, and its dimension is equal to the number of entities recognized in the text. Each element rel(s p ,s q ) in the matrix represents the semantic relationship between entity s p and entity s q in the text. This relationship is usually the actual association between entities extracted from the knowledge graph or the similarity measured in some way. For each pair of entities, rel(s p ,s q ) represents their semantic relationship, such as "belongs to", "", "is located at", etc.;
[0094] Central representation methods of matrix S: 1) Binary relation. If there is a certain relationship between entities, for example, an "is" relationship, then rel(s p , s q ) = 1; otherwise, it is 0. 2) Text-based similarity: In text, some entities may not have a direct relationship definition, but their relationship can be measured based on semantic similarity. For example, word vectors can be used to calculate the cosine similarity between entities in the text. 3) Weighted graph relationship. If there are multiple relationships in a knowledge graph, weights can be assigned to each relationship. For example, if there are two relationships, "belong to" and "contain", between s p and s q , and the importance of these two relationships in the knowledge graph is 0.7 and 0.3 respectively, then the elements in the relationship matrix can be calculated as a weighted sum;
[0095] Exemplarily, assume there are the following three entities: s1 = deep learning, s2 = convolutional neural network, and s3 = computer vision, and the following relationships: deep learning contains convolutional neural network, and convolutional neural network is part of computer vision. Then the relationship matrix can be expressed as: In this example, rel(s1, s2) = 1, indicating an inclusion relationship between s1 and s2, and rel(s2, s3) = 0.8, indicating a strong association between them.
[0096] Step S5: Analyze and classify the text theme based on the outputs of the LSTM neural network and the knowledge graph.
[0097] The specific steps of step S5 are as follows:
[0098] Step S501: Use a support vector machine (SVM) model to classify the text theme according to the outputs of the LSTM neural network and the knowledge graph;
[0099] Step S502: Analyze the classification results of the text theme, and combine the entity and relationship information in the text theme knowledge graph to infer the relationships between texts. The specific formula is:
[0100]
[0101] where P(y cross |Z) represents the probability of cross-text themes, that is, the comprehensive probability of the cross-text theme to which it belongs when given the text theme semantic vector Z. Y represents the set of text themes, containing all possible text themes. P(y i |Z) represents the classification probability of text theme y i , and P(y j |Z) represents the classification probability of text theme y jThe classification probability, rel(y i |y j ) represents the semantic relationship between the text topic y i and the text topic y j . denotes the summation over all text topic pairs (y i , y j ). By considering the relationships between all possible text pairs, the probability of the cross - text topic is inferred.
[0102] Principle of the above formula: The core idea of this formula is to utilize the relationships between text topics to infer the cross - text topics involved in a text. Traditional text classification methods usually only consider that a text belongs to a single text category. However, in actual scientific research and academic texts, many contents involve the intersection of multiple texts. For example, computer vision involves multiple text fields such as computer science and medicine. In this process, the formula comprehensively infers the relationships between multiple texts that a text may involve by combining the classification probability of each text topic and the relationship measure between texts, thereby obtaining the probability of the cross - text topic. This method can effectively handle the situation of text intersection and provide a richer understanding of texts than traditional single - text classification methods.
[0103] For example: Through reasoning, the relationship between "deep learning" and "computer vision" can be identified, and then the cross - domain information can be inferred. For example, "medical image analysis" belongs to the cross - domain of "computer vision" and "deep learning".
[0104] It should be noted that Figure 3 this is only a partial display example in the text topic knowledge graph.
[0105] Example 2
[0106] Please refer to Figure 4 , another example provided by the present invention: A text association analysis system based on an LSTM neural network and a knowledge graph, including: a data collection module, a data pre - processing module, a vector generation module, a semantic enhancement module, and a classification and analysis module;
[0107] The data collection module is used to collect text data of text topics from a literature database;
[0108] The data pre - processing module is used to pre - process the collected text data of text topics;
[0109] The vector generation module is used to learn the pre - processed text data of text topics through an LSTM network to generate vectors with semantic relevance for text topics;
[0110] The semantic enhancement module is used to construct a text topic knowledge graph and semantically enhance vectors that have semantic relevance to the text topic based on the constructed text topic knowledge graph;
[0111] The classification and analysis module is used to classify and analyze the text topic based on the outputs of the LSTM neural network and the knowledge graph.
[0112] The vector generation module includes: an LSTM network construction unit, a training unit, and a vector generation unit;
[0113] The LSTM network construction unit is used to construct an LSTM neural network model;
[0114] The training unit is used to train the constructed LSTM network and optimize the loss function;
[0115] The vector generation unit is used to convert the preprocessed text topic text data into text topic word vectors and input them into the trained LSTM network to generate vectors that have semantic relevance to the text topic.
[0116] The semantic enhancement module includes: a knowledge graph construction unit, an enhancement unit, and an association unit;
[0117] The knowledge graph construction unit is used to construct a text topic knowledge graph;
[0118] The enhancement unit is used to enhance the expressive power of the text topic knowledge graph using a graph convolutional neural network;
[0119] The association unit is used to establish a text topic context relationship model and construct a context relationship matrix between text topic entities.
[0120] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0121] As described in the above specific embodiments, the purpose, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above is only the specific embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A text association analysis method based on LSTM neural network and knowledge graph, characterized in that: include: Collect text data from literature databases; Preprocess the collected text data; The preprocessed text data is learned through the LSTM network to generate vectors with semantic relevance of text topics; Construct a text topic knowledge graph, and perform semantic enhancement on vectors with semantic relevance to text topics based on the constructed text topic knowledge graph; Based on the output of LSTM neural network and knowledge graph, text topics are classified and analyzed.
2. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 1, characterized in that: The method of learning the preprocessed text data through the LSTM network to generate a vector with semantic relevance of the text topic includes: Build an LSTM network, including forget gate, input gate and output gate; Use the pre-trained word embedding model to convert the pre-processed text topic data into text topic word vectors; Train the constructed LSTM network, optimize the loss function, and obtain a trained LSTM network; The text topic word vector is input into the trained LSTM network, and the output vector of the multi-head attention mechanism, the output vector of the LSTM network at time t and the output vector of semantic enhancement are concatenated or weighted summed to obtain a vector with semantic relevance of the text topic.
3. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 2, characterized in that: The construction of the text subject knowledge graph includes: Construct a text topic knowledge graph KG, expressed as: KG = {E, G, J}, where E represents the entity set, G represents the relationship set, and J represents the graph structure; Starting from any node and its neighbor nodes in the text topic knowledge graph KG, record all nodes reached by k jumps determined by the reward function, combine the semantic relevance vectors of the neighbor nodes of the arbitrary node at the kth layer, and update the embedding representation of the arbitrary node at the k+1th layer.
4. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 3, characterized in that: The method of performing semantic enhancement on vectors having semantic relevance to text topics based on constructing a text topic knowledge graph includes: The enhanced text topic knowledge graph and the vectors with semantic relevance of text topics are used to establish a text topic contextual relationship model and construct a contextual relationship matrix between text topic entities.
5. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 4, characterized in that: The text topics are analyzed and classified based on the output of the LSTM neural network and the knowledge graph, including: Using the support vector machine model, the text topics are classified according to the output of the LSTM neural network and the knowledge graph; The classification results of text topics are analyzed, and the entity and relationship information in the text topic knowledge graph are combined to infer the relationship between texts and calculate the probability of cross-text topics.
6. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 5, characterized in that: The preprocessing of the collected text subject text data includes: Remove stop words from text topic text data, standardize terms, and perform part-of-speech restoration and word segmentation; Based on the text domain knowledge and corpus, a text feature vocabulary is constructed, which contains common terms, keywords and their semantic associations in the text domain.
7. The text association analysis method based on LSTM neural network and knowledge graph as claimed in claim 6, characterized in that: The literature databases include: CNKI, Google Scholar and PubMed.
8. A text association analysis system based on LSTM neural network and knowledge graph, used to implement the text association analysis method based on LSTM neural network and knowledge graph described in any one of claims 1 to 7, characterized in that: include: Data collection module, data preprocessing module, vector generation module, semantic enhancement module and classification and analysis module; The data collection module is used to collect text data of text topics from a literature database; The data preprocessing module is used to preprocess the collected text subject text data; The vector generation module is used to learn the preprocessed text topic text data through the LSTM network to generate vectors with semantic relevance of the text topic; The semantic enhancement module is used to construct a text topic knowledge graph, and to perform semantic enhancement on vectors with semantic relevance to text topics based on the constructed text topic knowledge graph; The classification and analysis module is used to classify and analyze text topics based on the output of the LSTM neural network and the knowledge graph.
9. The text association analysis system based on LSTM neural network and knowledge graph as claimed in claim 8, characterized in that: The vector generation module includes: an LSTM network construction unit, a training unit and a vector generation unit; The LSTM network construction unit is used to construct an LSTM neural network model; The training unit is used to train the constructed LSTM network and optimize the loss function; The vector generating unit is used to convert the preprocessed text subject text data into text subject word vectors, and input them into the trained LSTM network to generate vectors with semantic relevance of the text subject.
10. The text association analysis system based on LSTM neural network and knowledge graph as claimed in claim 9, characterized in that: The semantic enhancement module includes: a knowledge graph construction unit, an enhancement unit and an association unit; The knowledge graph construction unit is used to construct a text topic knowledge graph; The enhancement unit is used to enhance the expression ability of the text topic knowledge graph by using a graph convolutional neural network; The association unit is used to establish a text subject context relationship model and construct a context relationship matrix between text subject entities.
Citation Information
Patent Citations
A neural network text classification method based on a multi-knowledge map
CN108984745A
Text classification method based on bidirectional long-short-term memory model and knowledge graph
CN115391532A
Domain long text classification method and system based on knowledge graph
CN116521882A
Method and device for text-enhanced knowledge graph joint representation learning
US20220147836A1