Intelligent question answering system based on knowledge graph
Through the intelligent question-answering system based on knowledge graph, through data collection, preprocessing and named entity recognition, semantic fusion and intelligent question-answering modules, the shortcomings of the existing policy question-answering system in accurate answers, information retrieval efficiency and personalized services are solved, and efficient, accurate and personalized policy interpretation is achieved.
Patent Information
- Application Number
- CN202510925462.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-05
- Publication Date
- 2025-10-03
AI Technical Summary
Existing policy question-and-answer systems have difficulty providing accurate answers when faced with complex or cross-domain issues. They have low information retrieval efficiency, lack personalized services, cannot effectively understand the semantics of policy texts, and have difficulty responding to diverse question-and-answer needs.
An intelligent question-answering system based on knowledge graph is adopted. Through knowledge base data collection, preprocessing, named entity recognition, semantic fusion and intelligent question-answering modules, Sentence-BERT and Word2vec models are used for entity matching and answer generation. The Attention mechanism is combined to integrate information and provide personalized policy interpretation.
The policy question-and-answer system has improved its answer accuracy and information retrieval efficiency, enabling it to handle diverse questions, provide personalized services, and enhance user experience.
Smart Images

Figure CN120745835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and more specifically to an intelligent question-answering system based on a knowledge graph. Background Art
[0002] With the advancement of information disclosure, more and more policy documents are being made available to the public, leading to a growing demand for access to and interpretation of policy information. However, policy documents are often complex and contain numerous technical terms and intricate clauses, making it difficult for the public to understand and apply these policies. Existing policy question-and-answer systems primarily rely on preset rules, text retrieval, and keyword matching to provide answers, but they still suffer from the following drawbacks: (1) Low answer accuracy: Traditional question-answering systems usually rely on preset rules and keyword matching to generate answers, lacking a deep understanding of text semantics. This approach easily leads to interpretation bias, and the resulting answers often do not meet the user's actual needs. Especially in the face of newly introduced policies or frequent policy changes, the adaptability of traditional systems is particularly insufficient. It is difficult to reason based on the semantic relevance of policy content, and thus cannot generate accurate answers. (2) Low information retrieval efficiency: When processing massive amounts of policy documents, existing systems often have difficulty locating relevant content quickly and accurately. Users often need to try different keyword combinations multiple times to obtain the required information, which reduces efficiency. Traditional retrieval methods are difficult to provide comprehensive answers to diverse question-and-answer needs. In particular, when users ask questions involving cross-domain knowledge, the system's performance is particularly insufficient, greatly reducing the user experience. (3) Difficulty in responding to diverse question-answering needs: Traditional question-answering systems often exhibit significant limitations in handling diverse user questions. In particular, when faced with complex questions that require cross-domain knowledge, the system often fails to provide accurate or effective answers. This is because traditional systems are typically based on simple retrieval and matching mechanisms, making it difficult to integrate and understand knowledge across multiple fields. Therefore, when users raise questions that span multiple policy areas or require comprehensive analysis, the system often struggles to meet their needs, resulting in low accuracy and relevance in answers. (4) Lack of personalized services: Existing question-and-answer systems are usually unable to provide personalized policy interpretations and suggestions based on the user's background information (such as region, industry or personal needs); due to the lack of consideration of the user's personalized needs, the system can only provide general and non-targeted content when generating answers; this makes it difficult for users to obtain guidance closely related to their own situation during use, and it is impossible to fully utilize the potential advantages of the intelligent question-and-answer system, which reduces the actual application effect of the system and user satisfaction. Summary of the Invention
[0003] In order to overcome the above-mentioned defects of the prior art, the present invention provides an intelligent question-answering system based on knowledge graph to solve the problems existing in the above-mentioned background technology.
[0004] The present invention provides the following technical solution: an intelligent question-answering system based on a knowledge graph, comprising: Knowledge base data acquisition module: collects target data and forms a data set; Data preprocessing module: preprocesses the data set and stores it in the database to facilitate subsequent knowledge graph construction; Knowledge graph construction module: uses a pre-trained named entity recognition model to extract entities and relationships from the database; Semantic Fusion Module: Based on the Sentence-BERT model, it matches and fuses similar entities from different sources in the knowledge graph, calculates cosine similarity, and automatically merges synonymous and similar entity nodes; Intelligent question answering module: uses a pre-trained named entity recognition model to parse and identify natural language questions input by users; Association and Enhancement Module: This module processes further questions raised by users regarding preliminary answers. It uses a graph traversal algorithm to automatically retrieve other entities and attributes associated with the preliminary answers through the node relationships in the knowledge graph, and uses a model based on the Attention mechanism to integrate the answer information to generate the final answer.
[0005] Preferably, the target data are policy documents, laws and regulations, expert interpretation articles and media reports from official websites; the target data can be collected by adopting an efficient crawling toolkit to automatically collect policy and regulatory information from websites and news platforms on a regular basis; the data preprocessing includes data cleaning and text data processing, and the text data processing includes word segmentation, removal of stop words and word form restoration.
[0006] Preferably, the named entity recognition model is a BERT-NER model, which can process long text input and use context information to identify entities and relationships in policy texts; the entities and relationships are converted into nodes and edges in a graph, the entities are policy texts, and the relationships are relationships between entities; the extracted entities and relationships are represented as triples using knowledge graph symbols, and the triples are subject, predicate, and object; the triples represent the relationship structure between policy entities.
[0007] Preferably, the preliminary answer is the answer generated by the intelligent question and answer module; the intelligent question and answer module vectorizes the user's question, identifies the key entities and intentions therein, maps the parsed question to the corresponding node in the knowledge graph, and matches the entities in the question with the entities in the knowledge graph by embedding the Word2vec model. If there is a matching node in the knowledge graph, a semantic query is performed in the knowledge graph according to the intention of the user's question.
[0008] Preferably, the semantic fusion module uses the Sentence-BERT model to generate a semantic vector representation of the entity; For each sentence or entity collected in the database, we map it into a fixed-length vector through the Sentence-BERT model. The formula is expressed as follows: ;in, Representing an entity The semantic vector of is the representation of the Sentence-BERT model used to generate vectors, Represents an entity; The semantic similarity between two entities is quantified by calculating the cosine similarity between different entity vectors; the formula is expressed as: ;in, and Represent the vector representation of the two entities respectively, represents the vector dot product, and represents the norm of a vector; express and The cosine similarity of .
[0009] Preferably, the intelligent question answering module uses a pre-trained named entity recognition model to parse the user's natural language questions, capture the keywords and intentions in the questions, and convert them into high-dimensional embedding vectors; The formula for problem vectorization is expressed as: ; in, and are the embedding vectors of keywords and intents respectively, To train the model, The question text entered by the user.
[0010] Preferably, the Word2vec model accurately matches the entities in the question with the entities in the knowledge graph, and the formula is expressed as: ; in, is the predicted center word vector, is the word vector of the word in the context, c is the size of the context window, and j represents the relative position of the center word.
[0011] Preferably, the Attention mechanism can dynamically adjust the importance of different answer information to form an optimal answer representation, which is expressed as follows: ;in, represents the Attention weight, Indicates the Answer information, ; Indicates the best answer.
[0012] Technical effects and advantages of the present invention: The present invention is provided with a semantic fusion module, an intelligent question-answering module, and an association and enhancement module, which is conducive to structuring the information in the policy text into a graph through entity and relationship extraction, storing it in a graph database, and realizing the association representation of policy information. It matches and fuses similar entities in multi-source data to reduce the redundancy of the knowledge graph; parses the natural language questions input by the user, maps them to the knowledge graph and generates accurate answers; retrieves related information and processes the user's supplementary questions to generate more detailed answers; effectively improves the answer accuracy and ability to handle diverse questions of the policy question-answering system, enhances the user experience, and finally integrates the answer information based on the model of the Attention mechanism, so that it can cope with diverse question-answering needs, while improving the answer accuracy and information retrieval efficiency, and providing users with personalized services. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a structural diagram of the knowledge graph-based intelligent question-answering system of the present invention.
[0014] Figure 2 Schematic diagram of feature extraction of the present invention.
[0015] Figure 3 Schematic diagram of the entity naming recognition process of the present invention.
[0016] Figure 4 Schematic diagram of the Sentence-BERT model of the present invention. DETAILED DESCRIPTION
[0017] The technical solutions of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the present invention. In addition, the forms of the various structures described in the following embodiments are merely examples. The knowledge graph-based intelligent question-answering system involved in the present invention is not limited to the various structures described in the following embodiments. All other implementations obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0018] like Figure 1 As shown, the present invention provides an intelligent question-answering system based on a knowledge graph, comprising: Knowledge base data collection module: collects target data and forms a data set; the target data includes policy documents, regulations, expert interpretation articles, and media reports from official websites. This target data can be collected by using an efficient web scraping toolkit to automatically collect policy and regulatory information from websites or news platforms on a regular basis; Data preprocessing module: Preprocesses the data set and stores it in the database to facilitate subsequent knowledge graph construction. Useable policy document data is obtained through data preprocessing, which includes data cleaning and text data processing. Text data processing includes word segmentation, stop word removal, and lemmatization to ensure that the text input in the target data conforms to the model's expected input format. Regular expressions and natural language processing techniques are used to standardize the text format, address issues such as redundant information, noise, or inconsistent formats in the text data, and delete irrelevant content in the text data. Knowledge graph construction module: Entities and relationships are extracted from the database using a pre-trained named entity recognition model, namely the BERT-NER model. This named entity recognition model can process long text inputs and use contextual information to identify entities and relationships in policy texts. The entities and relationships are converted into nodes and edges in a graph, where the entities are policy texts and the relationships are relationships between entities. The extracted entities and relationships are represented as triples using knowledge graph symbols, where the triples are subject, predicate, and object. The triples represent the relationship structure between policy entities and are stored in the knowledge graph database. The knowledge graph database provides efficient query and traversal operations, enabling the system to quickly retrieve relevant information in subsequent intelligent question-answering. Semantic Fusion Module: This module uses the Sentence-BERT model, an embedding-based entity alignment technology, to match and fuse similar entities from different sources in the knowledge graph. Through cosine similarity calculation and weighted self-learning tuning, it automatically merges synonymous or similar entity nodes, reducing graph redundancy and providing the best matching valid corpus for subsequent search queries, thereby generating satisfactory answers for users. Its purpose is to solve the problem of duplicate entities in multi-source data that may be encountered during the construction process. Intelligent Question Answering Module: This module uses a pre-trained named entity recognition model to parse and identify natural language questions entered by users, thereby extracting information to obtain the keywords and capturing the semantic features of the questions. It vectorizes the user's questions, identifies key entities and intent, maps the parsed questions to corresponding nodes in the knowledge graph, and uses entity linking technology and the Word2vec embedding model to match entities in the question with entities in the knowledge graph. If there is a matching node in the knowledge graph, a semantic query is performed in the knowledge graph based on the user's intent. The most relevant answer or a series of candidate answers are found and input into the large model as contextual knowledge base prompts, ultimately outputting a complete, fluent, and accurate answer that conforms to natural language expression. Association and enhancement module: processes further refined or supplementary questions raised by users regarding preliminary answers; uses graph traversal algorithms to automatically retrieve other entities and attributes associated with preliminary answers through node relationships in the knowledge graph; the preliminary answers are the answers generated by the intelligent question-answering module; uses vector-based similarity search to match possible similar questions, further expanding the range of possible answers; generates new question prompt information based on the user's specific supplementary questions and the retrieved related node information. This information is processed through a large model, and then a model based on the Attention mechanism is used to integrate answer information from different sources. Information fusion technology is used to ensure the accuracy and relevance of the final answer, thereby forming a more detailed or in-depth answer.
[0019] In this embodiment, it is important to specify that the knowledge base data acquisition module uses Scrapy, an efficient and scalable web data crawling framework, to regularly and automatically collect policy documents, regulatory provisions, expert interpretation articles, and related media reports from data sources such as official websites, news portals, and policy release platforms. By inheriting from the scrapy.Spider base class, custom data crawling rules are written for the corresponding sites to ensure that the acquired policy text data is comprehensive and authoritative. The collected data is stored in a locally deployed, high-performance PostgreSQL database with initialized parameters for subsequent processing.
[0020] In this embodiment, it should be specifically explained that the data preprocessing module uses the PdfReader.extract_text() method of the pypdf toolkit to extract all the text in the pdf file for the collected policy documents, and uses TesseractOCR to perform OCR text extraction on the pages that cannot be extracted from the pdf. For the text extracted from data sources such as policy documents, media reports, and interpretation articles, the cut() method reported in jieba is used to perform word segmentation operations to split continuous text strings into independent words or phrases. The text after word segmentation is subjected to morphological restoration. Lemmatic restoration restores words to their basic form to ensure that different forms of the same vocabulary can be consistently recognized and processed. The replace() method of the re toolkit is used to clean up irrelevant characters, symbols, HTML tags, and extra spaces in the text. Text cleaning can remove potential noise and ensure the standardization of data.
[0021] In this embodiment, it should be specifically explained that the knowledge graph construction module uses the pre-trained BERT-NER model for Chinese for named entity recognition. BertTokenizer and BertModel are imported from the transformers toolkit to load the pre-trained model. For the cleaned text in the database, BertTokenizer is first used to divide the paragraph sentences into tokens that the model can accept. These tokens are then input into the model in the form of tensors. The output obtained is that each token corresponds to an entity category: Tokens of the categories B_PER, I_PER, B_ORG, I_ORG, B_LOC, and I_LOC are selected from the model output. These categories represent key entities in the policy text, such as names of people, organizations, and places. By selecting tokens from these categories, important entity information can be extracted from the text, providing the necessary foundation for constructing a knowledge graph. Once these entities are identified and extracted, they are further processed to form nodes in the knowledge graph. Furthermore, using the contextual understanding capabilities of the BERT model, relationships between these entities can be extracted from the text, identifying implicit or explicit policy associations. The extracted entities and relationships are represented as triples (e.g., "subject-predicate-object") using the Resource Description Framework (RDF) format and stored in the Neo4j graph database using the py2neo toolkit's Graph.create(Relationship) method, providing efficient data support for subsequent semantic queries and intelligent question-answering.
[0022] In this embodiment, it should be specifically noted that the semantic fusion module uses the Sentence-BERT model to generate semantic vector representations of entities. Sentence-BERT is a sentence embedding model specifically used to generate semantic similarity. It can capture semantic information in text and convert it into a high-dimensional vector. For each sentence or entity collected in the database, we use the Sentence-BERT model to map it to a fixed-length vector, which is expressed as follows: ;in, Representing an entity The semantic vector of is the representation of the Sentence-BERT model used to generate vectors, Represents an entity; By calculating the cosine similarity between different entity vectors, we can quantify the semantic similarity between two entities; the formula is expressed as: ;in, and Represent the vector representation of the two entities respectively, represents the vector dot product, and represents the norm of a vector; express and The cosine similarity of By calculating cosine similarity, the system can determine the degree of similarity between entities. When the similarity value exceeds a preset threshold, the two entities are considered semantically similar. For semantically similar entities, we merge them and use the Graph.create(Relationship) method to update the nodes and relationships of similar entities.
[0023] In this embodiment, it should be specifically noted that the intelligent question-answering module uses a pre-trained named entity recognition model to parse the user's natural language questions, capture the keywords and intent in the questions, and convert them into high-dimensional embedding vectors; The formula for problem vectorization is expressed as: ; in, and are the embedding vectors of keywords and intents respectively, To train the model, Question text entered by the user; The Word2vec model is used to accurately match the entities in the question with the entities in the knowledge graph. The CBOW method of the Word2vec model predicts the central word through the context words. The formula is expressed as: ; in, is the predicted center word vector, is the word vector of the word in the context, c is the size of the context window, and j represents the relative position of the center word. Similarity is then calculated to accurately match the entity in the question with the entity node in the knowledge graph. The SPARQL query language is then used to search the knowledge graph for triples (entity-relationship-entity) relevant to the user's question. The query results serve as the context for the question. After obtaining the context, the system inputs the question and the retrieved semantic information into a generative model based on GPT. GPT models excel at natural language generation, generating fluent and coherent answers based on the context and question.
[0024] In this embodiment, it is important to specify that the association and enhancement module uses graph traversal algorithms, such as depth-first search or breadth-first search, to automatically retrieve other entities and attributes associated with the preliminary answer in the knowledge graph for the user's refined supplementary question. These algorithms can effectively discover potential information and connections related to the user's question. Other entities associated with the preliminary answer serve as alternative context to enrich the system's understanding and answering capabilities, providing a broader perspective and deeper background information, resulting in a more comprehensive answer to the user's question. Using vector-based similarity search to match possible similar questions further expands the range of possible answers. By vectorizing the user's question and the questions and answers in the knowledge base, a metric such as cosine similarity is used to compare the similarity between these vectors. If a question with a high degree of similarity to the user's question is found, its corresponding answer is used as an alternative answer, further enhancing the system's answering ability; The data obtained throughout the entire process is then embedded and merged with the preliminary answers. By using the Attention mechanism, we can effectively integrate information from different sources, making the final generated answer more comprehensive and accurate. The Attention mechanism dynamically adjusts the importance of different answer information, highlighting the most relevant parts and forming an optimal answer representation. The formula is as follows: ;in, represents the Attention weight, Indicates the Answer information, ; Indicates the optimal answer; it represents the importance of different contextual information when generating the final answer. Through a corpus integration strategy based on the Attention mechanism, the system can filter and integrate information from multiple sources, outputting the best matching contexts, prioritizing content that is most relevant to the user's question. Information fusion technology is then used to ensure the accuracy and relevance of the final answer, resulting in a more detailed or in-depth answer.
[0025] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0026] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An intelligent question-answering system based on knowledge graph, characterized by: include: Knowledge base data acquisition module: collects target data and forms a data set; Data preprocessing module: preprocesses the data set and stores it in the database to facilitate subsequent knowledge graph construction; Knowledge graph construction module: uses a pre-trained named entity recognition model to extract entities and relationships from the database; Semantic Fusion Module: Based on the Sentence-BERT model, it matches and fuses similar entities from different sources in the knowledge graph, calculates cosine similarity, and automatically merges synonymous and similar entity nodes; Intelligent question answering module: uses a pre-trained named entity recognition model to parse and identify natural language questions input by users; Association and Enhancement Module: This module processes further questions raised by users regarding preliminary answers. It uses a graph traversal algorithm to automatically retrieve other entities and attributes associated with the preliminary answers through the node relationships in the knowledge graph, and uses a model based on the Attention mechanism to integrate the answer information to generate the final answer.
2. The intelligent question-answering system based on knowledge graph according to claim 1, characterized in that: The target data refers to policy documents, laws and regulations, expert interpretation articles and media reports from official websites; the target data can be collected by using an efficient crawling toolkit to automatically collect policy and regulatory information from websites and news platforms on a regular basis; the data preprocessing includes data cleaning and text data processing, and the text data processing includes word segmentation, removal of stop words and word form restoration.
3. The intelligent question-answering system based on knowledge graph according to claim 2, characterized in that: The named entity recognition model is the BERT-NER model, which can process long text input and use contextual information to identify entities and relationships in policy texts; the entities and relationships are converted into nodes and edges in a graph, the entities are policy texts, and the relationships are relationships between entities; the extracted entities and relationships are represented as triples using knowledge graph symbols, and the triples are subject, predicate, and object; the triples represent the relationship structure between policy entities.
4. The intelligent question-answering system based on knowledge graph according to claim 3, characterized in that: The preliminary answer is the answer generated by the intelligent question and answer module; the further questions are further detailed questions and supplementary questions; the intelligent question and answer module vectorizes the user's question, identifies the key entities and intentions therein, maps the parsed question to the corresponding node in the knowledge graph, and matches the entities in the question with the entities in the knowledge graph by embedding the Word2vec model. If there is a matching node in the knowledge graph, a semantic query is performed in the knowledge graph according to the intention of the user's question.
5. The intelligent question-answering system based on knowledge graph according to claim 4, characterized in that: The semantic fusion module uses the Sentence-BERT model to generate semantic vector representations of entities; For each sentence or entity collected in the database, we map it into a fixed-length vector through the Sentence-BERT model. The formula is expressed as follows: ;in, Representing an entity The semantic vector of is the representation of the Sentence-BERT model used to generate vectors, Represents an entity; The semantic similarity between two entities is quantified by calculating the cosine similarity between different entity vectors; the formula is expressed as: ;in, and Represent the vector representation of the two entities respectively, represents the vector dot product, and represents the norm of a vector; express and The cosine similarity of .
6. The intelligent question-answering system based on knowledge graph according to claim 5, characterized in that: The intelligent question-answering module uses a pre-trained named entity recognition model to parse the user's natural language questions, capture the keywords and intent in the questions, and convert them into high-dimensional embedding vectors; The formula for problem vectorization is expressed as: ; in, and are the embedding vectors of keywords and intents respectively, To train the model, The question text entered by the user.
7. The intelligent question-answering system based on knowledge graph according to claim 6, characterized in that: The Word2vec model accurately matches the entities in the question with the entities in the knowledge graph, and the formula is expressed as: ; in, is the predicted center word vector, is the word vector of the word in the context, c is the size of the context window, and j represents the relative position of the center word.
8. The intelligent question-answering system based on knowledge graph according to claim 7, characterized in that: The Attention mechanism can dynamically adjust the importance of different answer information to form an optimal answer representation. The formula is: ;in, represents the Attention weight, Indicates the Answer information, ; Indicates the best answer.
Citation Information
Cited By
Education question and answer method based on multi-modal knowledge graph and multiple agents and related device
CN121880514A