Intelligent question and answer method based on knowledge graph

By constructing a knowledge graph of multi-source heterogeneous data and combining large language models for dynamic path search and optimization, the problem of insufficient logical reasoning and knowledge graph coverage of large language models is solved, and efficient and accurate intelligent Q&A is achieved, improving the Q&A performance and user experience in complex scenarios.

CN120297415APending Publication Date: 2025-07-11NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER
View PDF 0 Cites 42 Cited by

Patent Information

Application Number
CN202510379654.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Large language models are prone to problems with inaccurate or insufficient answers in logical reasoning and multi-hop reasoning scenarios, and there are problems with insufficient coverage of knowledge graphs and untimely data updates, which affects the performance of the intelligent question-and-answer system.

Method used

By building a knowledge graph of multi-source heterogeneous data, combining large language models and attention mechanisms, single-hop and multi-hop queries are performed, dynamic path searches are performed, and the knowledge graph is optimized to form a closed-loop system to achieve efficient and accurate intelligent question-and-answer.

Benefits of technology

It has improved the Q&A performance and user experience of the intelligent Q&A system in complex scenarios, and has the ability to continuously optimize. It is suitable for search engines, customer service systems, medical diagnosis and financial analysis and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297415A_ABST
    Figure CN120297415A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent question and answer method and system based on a knowledge graph, and relates to the technical field of knowledge graphs, the method comprises the following steps: S1, integrating multi-source heterogeneous data to construct an initial knowledge graph; s2, analyzing the questions of the user, executing single-hop query and judging whether a result meets requirements or not; if not, entering S3, performing multi-hop dynamic search based on an adaptive path extension algorithm, and generating a reasoning path and evidence; s4, calling a large language model to combine with an attention mechanism to perform semantic matching on a result or a path, and screening an optimal answer; and S5, correcting the knowledge graph and performing closed-loop feedback to the query process to realize continuous optimization. Through cooperation of the knowledge graph and the large language model and combination of dynamic path search and semantic matching, the problems that a traditional method is insufficient in knowledge coverage, low in reasoning efficiency and poor in answer interpretability are solved, the method has the advantages of being efficient, accurate and self-optimized, and the performance and user experience in a complex scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graphs, and particularly to an intelligent question-answering method based on a knowledge graph. Background Art

[0002] With the remarkable achievements of large language models (LLMs) in natural language processing tasks such as dialogue question-answering, machine translation, and information extraction, they have shown good performance in logical reasoning and language generation. However, in practical applications, due to data noise in the model training process and the lack of effective reference to external entity relationships and knowledge, LLMs are prone to the "hallucination" phenomenon, generating answers that do not conform to facts or are inaccurate. In addition, in multi-hop reasoning or long-chain reasoning scenarios, the model may make incorrect or insufficient inferences due to missing context or incomplete associations.

[0003] To alleviate the above problems, the knowledge graph, as a structured and semantic external knowledge source, provides accurate and interpretable factual support for LLMs and other intelligent question-answering systems. By introducing the knowledge graph into the question-answering process and combining single-hop or multi-hop queries, the accuracy and extensibility of the model's answers can be improved to a certain extent. However, if the knowledge graph itself has problems such as insufficient coverage, untimely data updates, and a single query strategy, it will also affect the performance of the intelligent question-answering system in complex scenarios. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an intelligent question-answering method and system based on a knowledge graph, which realizes efficient, accurate, and interpretable intelligent question-answering through the deep cooperation between the knowledge graph and the language model, and at the same time has the closed-loop ability of continuous optimization. It can be widely applied to fields such as search engines, customer service systems, medical diagnosis, and financial analysis, significantly improving the question-answering performance and user experience in complex scenarios.

[0005] To achieve the above purpose, the present invention provides the following solutions:

[0006] An intelligent question-answering method based on a knowledge graph, comprising:

[0007] S1. Obtain multi-source heterogeneous data and the natural language question statement input by the user, and construct an initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text;

[0008] S2, performing word segmentation, named entity recognition and semantic analysis on the question sentence, extracting the target entity and preliminary relationship requirements, and executing a single-hop query based on the initial knowledge graph and the preset single-hop query request to generate the first round of query results. If the first round of query results meet the answer requirements, jump to step S4; otherwise, go to step S3;

[0009] S3, based on the adaptive path extension algorithm, performing dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the preliminary relationship requirements and the first round query results, to obtain an extended reasoning path and associated evidence; the reasoning path includes a multi-hop entity relationship link;

[0010] S4, calling a preset large language model, and combining the attention mechanism to perform semantic matching on the first round query results or the reasoning path and the associated evidence, to obtain a list of candidate answers and corresponding semantic similarity scores, and determining the best answer among the candidate answers according to the semantic similarity scores;

[0011] S5. Modify the entities, relationships and attributes in the initial knowledge graph to obtain an updated knowledge graph, and apply the updated knowledge graph to steps S2 and S3 to form a continuously optimized closed-loop system.

[0012] Preferably, multi-source heterogeneous data and natural language question statements input by users are obtained, and an initial knowledge graph is constructed based on the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text, including:

[0013] Performing field parsing on the structured data, extracting entity names, attributes and relationships, and obtaining a classified structured data table;

[0014] Segment the unstructured text into sentences and words, and identify potential entities and relationships through natural language processing technology to obtain a candidate set of entity relationships in the unstructured text;

[0015] Deleting redundant fields of the structured data table, standardizing the naming format and merging synonymous entities to obtain a cleaned structured data table;

[0016] Filter out entity relationship pairs whose confidence is less than a preset threshold in the unstructured text entity relationship candidate set, and eliminate ambiguous expressions to obtain cleaned and standardized unstructured entity relationship pairs;

[0017] The cleaned structured data table and the entities, attributes and relationships in the normalized unstructured entity-relationship pairs are uniformly converted into standardized semantic triples;

[0018] Identify the conflicting triples in the standardized semantic triples through pre-set conflict detection rules, and verify the conflicting triples to obtain the verified semantic triples; the conflict detection rules include: entity uniqueness rule, entity attribute consistency rule, relationship uniqueness rule, relationship logic constraint rule, attribute value type rule, attribute value range rule, context consistency rule, timeline consistency rule, multi-source data consistency rule, data source priority rule, incremental data conflict rule, and version consistency rule; the entity uniqueness rule is used to detect whether there are entities with the same name but different meanings in the knowledge graph. If a conflict is detected, distinguish and assign a unique identifier through context or attributes; the entity attribute consistency rule is used to detect whether the attribute values of the same entity conform to logical constraints. If a conflict is detected, mark it as data to be corrected; the relationship uniqueness rule is used to detect whether there are multiple conflicting objects for the same pair of entities under a specific relationship. If a conflict is detected, select the correct value according to the pre-set priority; the relationship logic constraint rule is used to detect whether the relationship conforms to common sense or domain rules. If a conflict is detected, mark it as an invalid relationship; the attribute value type rule is used to detect whether the attribute value conforms to the predefined data type. If a conflict is detected, mark it as abnormal data; the attribute value range rule is used to detect whether the attribute value is within a reasonable range. If a conflict is detected, mark it as out-of-range data; the context consistency rule is used to detect whether the attribute values of entities or relationships are consistent with the context information. If a conflict is detected, mark it as context-contradictory data; the timeline consistency rule is used to detect whether the time attributes of entities or relationships conform to the timeline logic. If a conflict is detected, mark it as timeline-abnormal data; the multi-source data consistency rule is used to detect whether the information of the same entity or relationship obtained from different data sources is consistent. If a conflict is detected, mark it as multi-source contradictory data; the data source priority rule is used to select the correct value according to the pre-set data source credibility priority when multi-source data conflicts; the incremental data conflict rule is used to detect whether the newly added data is consistent with the historical data. If a conflict is detected, mark it as incremental contradictory data; the version consistency rule is used to detect whether the core entities and relationships are consistent between different versions of the knowledge graph. If a conflict is detected, mark it as version-abnormal data;

[0019] Import the verified semantic triples into the graph database to construct the initial knowledge graph; the initial knowledge graph includes entity nodes, relationship edges, and attribute labels.

[0020] Preferably, perform word segmentation, named entity recognition, and semantic analysis on the question statement, extract the target entity and the preliminary relationship requirements, and execute a single-hop query based on the initial knowledge graph and the pre-set single-hop query request to generate the first-round query result. If the first-round query result meets the answering requirements, jump to step S4; otherwise, enter step S3, including:

[0021] Filter punctuation marks, unify case, and process special characters for the question statement to generate a standardized text;

[0022] Perform word segmentation on the standardized text to generate a corresponding sequence of words;

[0023] Identify target entities in the sequence of words through a pre-trained named entity recognition model, and combine semantic role labeling technology to extract semantic relationships between target entities, obtaining a list of target entities and preliminary relationship requirements;

[0024] Based on the preliminary relationship requirements, map the target entities to nodes in the initial knowledge graph to generate a single-hop query request; the single-hop query request uses a graph query language;

[0025] Execute a single-hop query in the initial knowledge graph to obtain entities, relationships, and attributes directly associated with the target entities, obtaining a query result;

[0026] Evaluate the confidence of the query result and determine the query result as the first-round query result; the evaluation metrics include: result coverage rate, relationship matching degree, and evidence sufficiency; if the result coverage rate ≥ 90% and the evidence sufficiency meets the standard, it is determined that "the first-round query result meets the answering requirements".

[0027] Preferably, based on the adaptive path extension algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entities, the preliminary relationship requirements, and the first-round query result, obtaining an extended inference path and associated evidence, including:

[0028] Use the entities and relationships in the first-round query result as initial nodes to construct a candidate path set;

[0029] Filter the relationship types directly associated with the target entities according to the preliminary relationship requirements to generate an initial candidate path set;

[0030] Use a path scoring function to calculate the comprehensive score of each path, and sort the candidate paths from high to low according to the comprehensive score, and preferentially expand the high-scoring paths to obtain a sorted candidate path queue;

[0031] Adjust the maximum expansion depth of the sorted candidate path queue according to the current path score and a preset threshold;

[0032] At each hop of expansion, only retain the relationship types related to the semantics of the user's question, filter irrelevant branches, perform a graph traversal query on each path to obtain the next-hop entities and relationships, obtaining an extended multi-hop path set and new entity relationship evidence;

[0033] Perform redundant pruning, contradiction detection, and evidence fusion on the expanded multi-hop path set to obtain the optimized inference path and associated evidence.

[0034] Preferably, the path scoring function Score(P) is:

[0035] Score(P) = α·Rel_Confidence + β·Sem_Similarity + γ·Path_Length_Penalty;

[0036] where Rel_Confidence is the average confidence of the relationships in the path, and the average confidence is the statistical weight from the initial knowledge graph; Sem_Similarity is the cosine similarity between the path semantics and the user's question; Path_Length_Penalty is the path length penalty factor, Path_Length_Penalty = e -0.5n , n is the number of hops; α, β, and γ are all dynamic weights, and the initial values are set to 0.6, 0.3, and 0.1 respectively.

[0037] Preferably, the formula for the maximum expansion depth MaxDepth is:

[0038] Preferably, when the path length > 3, increase γ to 0.2 to suppress overly long paths; if the user's question contains explicit relationship keywords, increase α to 0.7.

[0039] Preferably, call a pre-set large language model and, in combination with the attention mechanism, perform semantic matching on the first-round query results or the inference path and the associated evidence to obtain a list of candidate answers and the corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores, including:

[0040] Encode the first-round query results into a structured text sequence; or encode the inference path and the associated evidence into a structured text sequence;

[0041] Call the pre-trained large language model GPT-4 and use the unidirectional attention mechanism to perform semantic matching on the structured text sequence to obtain a list of candidate answers and the corresponding semantic similarity scores;

[0042] Determine the candidate answer in the candidate answer list with the highest semantic similarity score as the optimal answer.

[0043] Preferably, the calculation formula for the semantic similarity score Sim(Q, E) is:

[0044]

[0045] Among them, Q is the user's question, the natural language question input by the user, which is decomposed into n semantic units after encoding; E is the structured text sequence, which is decomposed into m semantic units after encoding; q is the i-th semantic unit of the question, and e j is the j-th semantic unit of the evidence; φ(q i ) is the question field strength function, φ(q i ) = ‖q i ‖·exp(-λ·Entropy(q i ))), where Entropy(q i ) is the information entropy of the semantic unit, and λ is the attenuation coefficient; ψ(e j ) is the evidence field strength function, ψ(e j ) = ||e j ||·exp(-λ·Entropy(e j )); Dist(q i , e j ) is the semantic unit distance, where Σ is the covariance matrix, pre-computed through a large-scale corpus; γ' is the distance attenuation exponent, and γ' is 2; Z is the normalization factor, used to scale the score to the [0, 1] interval.

[0046] An intelligent question-answering system based on a knowledge graph, comprising:

[0047] A knowledge graph construction unit, configured to obtain multi-source heterogeneous data and a question statement of the natural language input by the user, and construct an initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text;

[0048] A first-round query unit, configured to perform word segmentation, named entity recognition, and semantic analysis on the question statement, extract the target entity and the preliminary relationship requirement, and perform a single-hop query based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answer requirement, it jumps to the answer screening unit; otherwise, it enters the second-round query unit;

[0049] A second-round query unit, based on the adaptive path extension algorithm, performs dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the preliminary relationship requirement, and the first-round query result to obtain an extended inference path and associated evidence; the inference path includes a multi-hop entity relationship link;

[0050] An answer screening unit, configured to call a preset large language model, and perform semantic matching on the first-round query result or the reasoning path and the associated evidence in combination with the attention mechanism, obtain a list of candidate answers and corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores;

[0051] A knowledge graph update unit, configured to correct entities, relationships, and attributes in the initial knowledge graph, obtain an updated knowledge graph, and apply the updated knowledge graph to the first-round query unit and the second-round query unit to form a continuously optimized closed-loop system.

[0052] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:

[0053] The present invention provides an intelligent question-answering method and system based on a knowledge graph. The method includes: S1. Obtain multi-source heterogeneous data and a natural language question statement input by a user, and construct an initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text; S2. Perform word segmentation, named entity recognition, and semantic analysis on the question statement, extract a target entity and a preliminary relationship requirement, and perform a single-hop query based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answering requirement, jump to step S4; otherwise, enter step S3; S3. Based on an adaptive path expansion algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the preliminary relationship requirement, and the first-round query result to obtain an extended reasoning path and associated evidence; the reasoning path includes a multi-hop entity relationship link; S4. Call a preset large language model, and perform semantic matching on the first-round query result or the reasoning path and the associated evidence in combination with the attention mechanism, obtain a list of candidate answers and corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores; S5. Correct entities, relationships, and attributes in the initial knowledge graph, obtain an updated knowledge graph, and apply the updated knowledge graph to step S2 and step S3 to form a continuously optimized closed-loop system. Through the deep cooperation between the knowledge graph and the language model, the present invention realizes efficient, accurate, and interpretable intelligent question answering, and at the same time has the closed-loop ability of continuous optimization, and can be widely applied to fields such as search engines, customer service systems, medical diagnoses, and financial analysis, significantly improving the question-answering performance and user experience in complex scenarios. Description of the Drawings

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0055] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention;

[0056] Figure 2 It is the schematic diagram of the technical route provided by the embodiment of the present invention;

[0057] Figure 3 It is the schematic diagram of the system structure provided by the embodiment of the present invention. Detailed implementation manners

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0059] The object of the present invention is to provide an intelligent question-answering method and system based on a knowledge graph. Through the deep cooperation between the knowledge graph and the language model, efficient, accurate, and interpretable intelligent question answering is achieved. At the same time, it has the closed-loop ability of continuous optimization and can be widely applied to fields such as search engines, customer service systems, medical diagnosis, and financial analysis, significantly improving the question-answering performance and user experience in complex scenarios.

[0060] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail in conjunction with the accompanying drawings and specific implementation manners.

[0061] As Figure 1 and Figure 2 shown, the present invention provides an intelligent question-answering method based on a knowledge graph, including:

[0062] S1. Obtain multi-source heterogeneous data and the natural language question statement input by the user, and construct an initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text;

[0063] S2. Tokenize, perform named entity recognition, and semantic analysis on the question statement, extract the target entity and the initial relationship requirements, and perform a single-hop query based on the initial knowledge graph and the preset single-hop query request to generate the first-round query results. If the first-round query results meet the answering requirements, jump to step S4; otherwise, enter step S3;

[0064] S3. Based on the adaptive path extension algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the initial relationship requirements, and the first-round query results to obtain the extended inference path and associated evidence; the inference path includes multi-hop entity relationship links;

[0065] S4. Invoke the preset large language model, and perform semantic matching on the first-round query results or the inference path and associated evidence by combining the attention mechanism to obtain the candidate answer list and the corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores;

[0066] S5. Modify the entities, relationships, and attributes in the initial knowledge graph to obtain the updated knowledge graph, and apply the updated knowledge graph to steps S2 and S3 to form a continuously optimized closed-loop system.

[0067] Preferably, obtain multi-source heterogeneous data and the question statement in natural language input by the user, and construct the initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text, including:

[0068] Perform field parsing on the structured data, extract the entity names, attributes, and relationships to obtain the classified structured data table;

[0069] Perform sentence splitting and tokenization on the unstructured text, and identify potential entities and relationships through natural language processing techniques to obtain the unstructured text entity relationship candidate set;

[0070] Delete the redundant fields in the structured data table, and perform standardized naming format and merge synonymous entities to obtain the cleaned structured data table;

[0071] Filter the entity relationship pairs in the unstructured text entity relationship candidate set with confidence less than the preset threshold, and eliminate ambiguous expressions to obtain the cleaned and standardized unstructured entity relationship pairs;

[0072] Uniformly convert the entities, attributes, and relationships in the cleaned structured data table and the standardized unstructured entity relationship pairs into standardized semantic triples;

[0073] Identify the conflicting triples in the standardized semantic triples through pre-set conflict detection rules, and verify the conflicting triples to obtain the verified semantic triples; the conflict detection rules include: entity uniqueness rule, entity attribute consistency rule, relationship uniqueness rule, relationship logic constraint rule, attribute value type rule, attribute value range rule, context consistency rule, timeline consistency rule, multi-source data consistency rule, data source priority rule, incremental data conflict rule, and version consistency rule; the entity uniqueness rule is used to detect whether there are entities with the same name but different meanings in the knowledge graph. If a conflict is detected, distinguish and assign a unique identifier through context or attributes; the entity attribute consistency rule is used to detect whether the attribute values of the same entity conform to logical constraints. If a conflict is detected, mark it as data to be corrected; the relationship uniqueness rule is used to detect whether there are multiple conflicting objects for the same pair of entities under a specific relationship. If a conflict is detected, select the correct value according to the pre-set priority; the relationship logic constraint rule is used to detect whether the relationship conforms to common sense or domain rules. If a conflict is detected, mark it as an invalid relationship; the attribute value type rule is used to detect whether the attribute value conforms to the predefined data type. If a conflict is detected, mark it as abnormal data; the attribute value range rule is used to detect whether the attribute value is within a reasonable range. If a conflict is detected, mark it as out-of-range data; the context consistency rule is used to detect whether the attribute values of entities or relationships are consistent with the context information. If a conflict is detected, mark it as context-contradictory data; the timeline consistency rule is used to detect whether the time attributes of entities or relationships conform to the timeline logic. If a conflict is detected, mark it as timeline-abnormal data; the multi-source data consistency rule is used to detect whether the information of the same entity or relationship obtained from different data sources is consistent. If a conflict is detected, mark it as multi-source contradictory data; the data source priority rule is used to select the correct value according to the pre-set data source credibility priority when multi-source data conflicts; the incremental data conflict rule is used to detect whether the newly added data is consistent with the historical data. If a conflict is detected, mark it as incremental contradictory data; the version consistency rule is used to detect whether the core entities and relationships are consistent between different versions of the knowledge graph. If a conflict is detected, mark it as version-abnormal data;

[0074] Import the verified semantic triples into a graph database to construct the initial knowledge graph; the initial knowledge graph includes entity nodes, relationship edges, and attribute labels.

[0075] Specifically, the sub-steps of step S1 "Knowledge Graph Construction and Initialization" in this embodiment are as follows:

[0076] S1.1 Multi-source heterogeneous data acquisition and classification: acquire multi-source heterogeneous data from structured databases, API interfaces, unstructured texts crawled from web pages, and user logs; perform field parsing on structured data (such as database tables, CSV files) to extract entity names, attributes, and relationships; perform sentence splitting and word segmentation on unstructured texts (such as documents, web page content), and identify potential entities and relationships through natural language processing technologies (NER, relationship extraction models); finally obtain a classified structured data table and an unstructured text entity relationship candidate set.

[0077] S1.2 Data cleaning and normalization: perform operations on the structured data output by S1.1, such as deleting redundant fields, standardizing naming formats (such as unifying the date format to "YYYY-MM-DD"), and merging synonymous entities (such as mapping "Beijing" and "Beijing City" to the same entity); perform operations on unstructured data, such as filtering low-confidence entity relationship pairs (confidence < preset threshold) and eliminating ambiguous expressions (such as distinguishing "Apple (company)" and "Apple (fruit)" through a context disambiguation model), to obtain a cleaned structured data table and a normalized unstructured entity relationship pair.

[0078] S1.3 Semantic triple generation and conflict detection: uniformly convert entities, attributes, and relationships in structured data and unstructured data into standardized semantic triples (subject-relationship-object); identify conflict triples through conflict detection rules (such as "the same subject cannot have multiple contradictory objects under the same relationship") to obtain a preliminary semantic triple set and a conflict triple list.

[0079] S1.4 Semi-automatic manual verification and knowledge graph construction: based on the preliminary semantic triple set and conflict triple list output by S1.3, through a visualization verification tool, display conflict triples and context evidence for manual confirmation of correction schemes (such as deleting incorrect triples or supplementing constraint conditions); import the verified semantic triples into a graph database (such as Neo4j) to construct an initial knowledge graph, including entity nodes, relationship edges, and attribute labels, to obtain an initial knowledge graph (including entities, relationships, attributes, and topological structures).

[0080] Optionally, the conflict detection rules in this embodiment are defined as:

[0081] ① Entity conflict detection rules

[0082] (1) Entity uniqueness rule: the same entity should have a unique identifier in the knowledge graph to avoid duplication or ambiguity.

[0083] Example: When it is detected that "Apple (company)" and "Apple (fruit)" have the same name but are different entities, they need to be distinguished by context or attributes and assigned a unique ID. When it is detected that "Beijing City" and "Beijing" are the same entity, they need to be merged into a single node.

[0084] (2) Entity attribute consistency rule: The attribute values of the same entity should conform to logical constraints to avoid contradictions.

[0085] Example: When it is detected that the "date of birth of a person" is "2023", if the age attribute of this person is "50 years old", it is determined as a conflict. When it is detected that the "population of a certain city" is "10 million", if the city attribute is "small town", it is determined as a conflict.

[0086] ② Relationship conflict detection rule

[0087] (1) Relationship uniqueness rule

[0088] Rule description: The same pair of entities should have a unique object under a specific relationship to avoid contradictory relationships.

[0089] Example: When it is detected that the "CEO of Company A" is "Zhang San" and "Li Si", it is determined as a conflict (a company can only have one CEO). When it is detected that the "spouse of a person" is "Wang Wu" and "Zhao Liu", it is determined as a conflict (under monogamy, a person can only have one spouse).

[0090] (2) Relationship logical constraint rule: The relationship should conform to logical constraints to avoid violating common sense or domain rules.

[0091] Example: When it is detected that the "father of a person" is "oneself", it is determined as a conflict (violating biological logic). When it is detected that the "capital of a certain city" is "another city", it is determined as a conflict (one city cannot be the capital of another city).

[0092] ③ Attribute conflict detection rule

[0093] (1) Attribute value type rule: The attribute value should conform to the predefined data type.

[0094] Example: When it is detected that the "age of a person" is "abc", it is determined as a conflict (age should be a numerical type). When it is detected that the "establishment date of a certain company" is "a future date", it is determined as a conflict (the establishment date should be a past or current date).

[0095] (2) Attribute value range rule: The attribute value should be within a reasonable range.

[0096] Example: When it is detected that the "height of a person" is "300 cm", it is determined as a conflict (exceeding the human height range). When it is detected that the "population of a certain city" is "-1000", it is determined as a conflict (the population cannot be negative).

[0097] ④ Context conflict detection rule

[0098] (1) Context Consistency Rule: The attribute values of entities or relationships should be consistent with the context information.

[0099] Example: When it is detected that the "occupation of a person" is "doctor", but their educational background is "no medical education", it is determined to be in conflict. When it is detected that the "headquarters of a company" is "Beijing", but its main business area is "limited to Europe", it is determined to be in conflict.

[0100] (2) Timeline Consistency Rule: The time attributes of entities or relationships should conform to the timeline logic.

[0101] Example: When it is detected that the "graduation time of a person" is earlier than the "enrollment time", it is determined to be in conflict. When it is detected that the "occurrence time of an event" is a "future date", but its description is a "historical event", it is determined to be in conflict.

[0102] ⑤ Multi-source Data Conflict Detection Rule

[0103] (1) Multi-source Data Consistency Rule: The information of the same entity or relationship obtained from different data sources should be consistent.

[0104] Example: When it is detected that the "date of birth of a person" is "1990" in data source A and "1995" in data source B, it is determined to be in conflict. When it is detected that the "market value of a company" is "10 billion" in data source A and "20 billion" in data source B, it is determined to be in conflict.

[0105] (2) Data Source Priority Rule: When there is a conflict in multi-source data, the correct value should be selected according to the credibility priority of the data sources.

[0106] Example: It is set that the "official database" has a higher priority than "social media". When there is a conflict between the two, the official data is preferred.

[0107] ⑥ Dynamic Update Conflict Detection Rule

[0108] (1) Incremental Data Conflict Rule: The newly added data should be consistent with the historical data to avoid introducing conflicts.

[0109] Example: When it is detected that the "nationality of a person" in the newly added data is inconsistent with the historical data, it is determined to be in conflict. When it is detected that the "CEO of a company" in the newly added data is inconsistent with the historical data, it is determined to be in conflict.

[0110] (2) Version Consistency Rule: The core entities and relationships should be consistent between different versions of the knowledge graph.

[0111] Example: When it is detected that a core entity in the historical version is deleted in the new version, it is determined to be in conflict. When it is detected that a key relationship in the historical version is modified in the new version, it is determined to be in conflict.

[0112] In summary, the conflict detection rules are the core of knowledge graph quality control, covering multiple dimensions such as entities, relationships, attributes, context, multi-source data, and dynamic updates. Through the above rules, inconsistencies in the knowledge graph can be effectively identified and resolved to ensure its accuracy, integrity, and timeliness.

[0113] Preferably, the question statement is tokenized, named entity recognition and semantic analysis are performed, the target entities and preliminary relationship requirements are extracted, and a single-hop query is executed based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answering requirements, jump to step S4; otherwise, enter step S3, including:

[0114] Filter punctuation marks, unify case, and process special characters in the question statement to generate a normalized text;

[0115] Perform tokenization on the normalized text to generate a corresponding sequence of words;

[0116] Identify the target entities in the sequence of words through a pre-trained named entity recognition model, and combine semantic role labeling technology to extract the semantic relationships between the target entities to obtain a list of target entities and preliminary relationship requirements;

[0117] Based on the preliminary relationship requirements, map the target entities to the nodes in the initial knowledge graph to generate a single-hop query request; the single-hop query request uses a graph query language;

[0118] Execute a single-hop query in the initial knowledge graph to obtain the entities, relationships, and attributes directly associated with the target entities to obtain a query result;

[0119] Evaluate the confidence of the query result and determine the query result as the first-round query result; the evaluation indicators include: result coverage rate, relationship matching degree, and sufficiency of evidence; if the result coverage rate ≥ 90% and the sufficiency of evidence meets the standard, it is determined that "the first-round query result meets the answering requirements".

[0120] Specifically, the sub-steps of step S2 "Question analysis and single-hop query" in this embodiment specifically include:

[0121] S2.1 Natural language question preprocessing, filter punctuation marks, unify case, and process special characters in the natural language question statement input by the user to generate a normalized text; perform tokenization on the text to generate a sequence of words (Token sequence), for example, using a tokenization model based on a dictionary or BERT, and finally obtain the preprocessed normalized text and tokenization result.

[0122] S2.2 Named Entity Recognition and Semantic Role Labeling. Based on the normalized text and word segmentation results output by S2.1, use pre-trained named entity recognition models (such as BiLSTM-CRF, BERT-NER) to identify the target entities in the question (such as person names, locations, times, organizations); combine semantic role labeling (SRL) technology to extract the semantic relationships between entities (such as "someone's occupation", "participants in an event") to obtain a list of target entities and preliminary relationship requirements (for example, entity A - relationship R - entity B).

[0123] S2.3 Single-hop Query Request Generation. According to the list of target entities and preliminary relationship requirements output by S2.2, map the target entities to the nodes in the knowledge graph to generate a single-hop query request; the single-hop query request uses a graph query language (such as Cypher, SPARQL), and the format is "MATCH( entity )-[ relationship ]->( target ) WHERE entity = 'X' RETURN target";

[0124] S2.4 Single-hop Query Execution and Result Evaluation. Based on the structured single-hop query statement and the initial knowledge graph generated by S2.3, execute a single-hop query in a graph database (such as Neo4j, Amazon Neptune) to obtain the entities, relationships, and attributes directly associated with the target entity; perform a confidence evaluation on the query results, and the evaluation indicators include: result coverage rate, whether all target entities in the question are covered; relationship matching degree, the semantic consistency between the query result and the preliminary relationship requirements; evidence sufficiency, whether the result contains an evidence chain sufficient to support the answer generation; if the result coverage rate ≥ the preset threshold (such as 90%) and the evidence sufficiency meets the standard, it is determined as "meeting the answer requirements", and finally obtain the first-round query results (including entities, relationships, attributes, and evidence) and evaluation conclusions.

[0125] S2.5 Process Jump Decision. If the evaluation conclusion is "meeting the answer requirements", input the first-round query results into step S4 for answer generation; otherwise, input the target entity, preliminary relationship requirements, and the first-round query results into step S3 for dynamic path search dependent claims (step S2 refinement)

[0126] Furthermore, in step S2.2, the named entity recognition model is optimized through domain adaptation training, specifically:

[0127] On the basis of pre-training in a general corpus, use the labeled data in the target domain (such as medical, financial) to fine-tune the model;

[0128] Through the entity disambiguation module, combine the entity attributes in the knowledge graph to eliminate the ambiguity of homonymous entities (such as distinguishing "Apple (company)" from "Apple (fruit)").

[0129] In the step S2.3, the generation of the single-hop query request further includes:

[0130] Semantically expand the preliminary relationship requirements, and through a thesaurus or ontology mapping, convert the relationship expressed by the user into a standardized relationship type in the knowledge graph;

[0131] Example: Map "works at" in the user's question to the "employer" relationship in the knowledge graph.

[0132] In the step S2.4, the confidence evaluation is achieved in the following manner:

[0133] Use a pre-trained semantic matching model (such as Sentence-BERT) to calculate the semantic similarity between the user's question and the query result; if the semantic similarity ≥ threshold and the result coverage meets the standard, it is determined to meet the answering requirement.

[0134] In the step S2.5, the process jump decision further includes:

[0135] When the first-round query results partially meet the requirements, directly pass the high-confidence part of the results to step S4, and at the same time trigger step S3 to supplement the query for the low-confidence part; Example: If the confidence of the query result of "someone's occupation" is high, directly generate an answer; if the result of "participants in an event" is insufficient, trigger a multi-hop query.

[0136] Preferably, based on the adaptive path extension algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the preliminary relationship requirements, and the first-round query results, to obtain an extended inference path and associated evidence, including:

[0137] Take the entities and relationships in the first-round query results as the initial nodes to construct a candidate path set;

[0138] According to the preliminary relationship requirements, screen the relationship types directly associated with the target entity to generate an initial candidate path set;

[0139] Use a path scoring function to calculate the comprehensive score of each path, and sort the candidate paths from high to low according to the comprehensive score, and preferentially expand the high-scoring paths to obtain a sorted candidate path queue;

[0140] According to the current path score and the preset threshold, adjust the maximum expansion depth of the sorted candidate path queue;

[0141] At each hop expansion, only retain the relationship types semantically relevant to the user's question, filter out irrelevant branches, perform a graph traversal query on each path to obtain the next-hop entity and relationship, and obtain an extended multi-hop path set and new entity relationship evidence;

[0142] Perform redundant pruning, contradiction detection, and evidence fusion on the expanded multi-hop path set to obtain the optimized inference path and associated evidence. Exemplarily, the evidence fusion in this embodiment associates the entity relationship links in the multi-hop path with the first-round query results to construct a complete inference evidence chain.

[0143] Optionally, in this embodiment, the relationship types directly associated with the target entity are screened according to the preliminary relationship requirements to generate an initial candidate path set, which specifically includes:

[0144] Through a predefined relationship ontology library, map natural language relationship words to standardized relationship types in the knowledge graph (e.g., map "CEO" to "executive_leader"). If there are polysemous words (e.g., "apple" corresponding to "company" or "fruit"), disambiguate them in combination with the context entity attributes (e.g., the industry label of "Apple Inc.");

[0145] Execute a single-hop query in the graph database to retrieve all relationship edges directly associated with the target entity, and screen the matching relationship edges according to the standardized relationship type list (e.g., only retain edges of executive_leader and founder types);

[0146] Convert each relationship edge into a single-hop path in the format of entity-relationship-target entity (e.g., Path1: Apple Inc. → executive_leader → Tim Cook); if the preliminary relationship requirements are not fully matched (e.g., the user's question implies a multi-hop relationship), extend to the one-hop relationship of adjacent entities, and filter out low-quality paths through a path confidence threshold (e.g., paths with a relationship weight lower than 0.5), and finally obtain the initial candidate path set and associated evidence (e.g., path list, relationship weight, entity attributes);

[0147] Based on the semantic context of the user's question, calculate the matching degree of each path with the question through a semantic similarity model (e.g., Sentence-BERT), and dynamically adjust the path priority in combination with the statistical weights of the relationships in the knowledge graph (e.g., occurrence frequency, source credibility), sort by the comprehensive score, generate a priority queue for subsequent expansion, and finally obtain a sorted list of high-value candidate paths.

[0148] Preferably, the path scoring function Score(P) is:

[0149] Score(P) = α·Rel_Confidence + β·Sem_Similarity + γ·Path_Length_Penalty;

[0150] Among them, Rel_Confidence is the average confidence of the relationships in the path, and the average confidence is the statistical weight from the initial knowledge graph; Sem_Similarity is the cosine similarity between the path semantics and the user's question; Path_Length_Penalty is the path length penalty factor, Path_Length_Penalty = e -0.5n , where n is the number of hops; α, β, and γ are all dynamic weights, and their initial values are set to 0.6, 0.3, and 0.1 respectively.

[0151] Preferably, the formula for the maximum expansion depth MaxDepth is: Exemplarily, in the adaptive path expansion, the radius expansion strategy further includes: through the relationship attribute filtering mechanism of the knowledge graph, only retain the relationships that match the constraint conditions such as time and location in the user's question; Example: If the question limits "events after 2024", then filter the relationship edges in the knowledge graph whose time attributes are earlier than 2024.

[0152] Preferably, when the path length > 3, increase γ to 0.2 to suppress overly long paths; if the user's question contains clear relationship keywords (such as "causal relationship"), then increase α to 0.7.

[0153] Specifically, the contradiction detection is based on the logical rule base of the knowledge graph (such as "doctors need to have a medical degree") to detect the contradictory relationships in the path; if a contradiction is detected, the path score is reduced by 50%. Redundancy pruning is to merge duplicate paths (such as A→B→C and A→D→C but with equivalent semantics), and retain the one with the highest score.

[0154] Preferably, call a preset large language model, and combine the attention mechanism to perform semantic matching on the first-round query results or the inference path and the associated evidence to obtain a list of candidate answers and the corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores, including:

[0155] Encode the first-round query results into a structured text sequence; or encode the inference path and the associated evidence into a structured text sequence;

[0156] Call the pre-trained large language model GPT-4, and use the unidirectional attention mechanism to perform semantic matching on the structured text sequence to obtain a list of candidate answers and the corresponding semantic similarity scores;

[0157] Determine the candidate answer in the candidate answer list corresponding to the highest semantic similarity score as the optimal answer.

[0158] Specifically, the sub-steps of "answer generation and screening" in this embodiment specifically include:

[0159] S4.1 Input data integration. If it is determined in step S2 that the answering requirements are met, the input is the result of the first-round query (entities, relationships, and evidence); if multi-hop query is required, the input is the reasoning path and associated evidence output in step S3. Encode the input data into a structured text sequence (such as "Entity A - Relationship - Entity B"), retaining the core entities and relationship links. Perform abstract extraction on long-text evidence, only retaining the content relevant to the user's question, and finally obtain the refined structured evidence text.

[0160] S4.2 Direct semantic matching and candidate answer generation. Based on the structured evidence text output in S4.1 and the user's original question, call a pre-trained large language model (such as BERT, GPT-3.5), and calculate the semantic similarity between the user's question and each piece of evidence through a unidirectional attention mechanism: Sort according to the similarity score, and select the Top-K (default K = 3) evidence to generate candidate answers. If the evidence contains a clear answer (such as "The attribute value of Entity A is X"), directly extract the answer; if reasoning is required, generate an answer through a prompt template (such as "According to evidence [E], the inferred answer is [X]"), obtaining a list of candidate answers and their corresponding semantic similarity scores.

[0161] S4.3 Confidence screening and answer output. If the highest similarity score ≥ the preset threshold (such as 0.8), directly select this answer; if the scores are similar (the difference between the highest and the second highest is < 5%), then verify with the knowledge graph, check whether the entity relationships in the answer exist in the graph, and preferentially select the answer passed by the graph verification, outputting the final answer and the associated evidence abstract.

[0162] Specifically, the unidirectional attention mechanism in this embodiment is optimized in the following way: assign higher weights to the keywords (such as entities, time, location) in the user's question to enhance the pertinence of semantic matching; Example: If the question is "The occupation of someone", then calculate the weighted value for the "occupation" relationship field in the evidence.

[0163] Optionally, candidate answer generation further includes: If the directly extracted answer conflicts with the generated answer, preferentially select the directly extracted result; Example: If the evidence is clearly "The birthplace of someone is Beijing", then reject the generated answer "May be Shanghai".

[0164] Furthermore, the knowledge graph verification is implemented in the following way:

[0165] Execute a real-time graph query on the entity relationship pairs in the answer. If it exists, the confidence is increased by 20%; if it does not exist, it is marked as "low confidence" and triggers manual review.

[0166] Preferably, the calculation formula for the semantic similarity score Sim(Q, E) is:

[0167]

[0168] Among them, Q is the user's question, the natural language question input by the user, which is decomposed into n semantic units after encoding; E is the structured text sequence, which is decomposed into m semantic units after encoding; q is the i-th semantic unit of the question, and e j is the j-th semantic unit of the evidence; φ(q i ) is the question field strength function, φ(q i ) = ‖q i ‖·exp(-λ·Entropy(q i ))), where Entropy(q i ) is the information entropy of the semantic unit, and λ is the attenuation coefficient; ψ(e j ) is the evidence field strength function, ψ(e j ) = ||e j ||·exp(-λ·Entropy(e j )); Dist(q i , e j ) is the semantic unit distance, where Σ is the covariance matrix, pre-computed through a large-scale corpus; γ' is the distance attenuation exponent, and γ' is 2; Z is the normalization factor, used to scale the score to the [0, 1] interval.

[0169] Specifically, in this embodiment, each semantic unit is regarded as a "field source" in the field space, and its intensity is jointly determined by the vector norm and the information entropy (the lower the information entropy, the higher the semantic certainty and the greater the field strength). By the ratio of the field strength product to the distance attenuation, the interaction of semantic units is simulated, and the physical meaning is clear. In this embodiment, the covariance matrix Σ is introduced, and the correlation of different semantic units is statistically calculated through the pre-trained corpus, so as to solve the problem that the traditional cosine similarity ignores the feature correlation. The normalization factor Z dynamically adapts to the input text length and semantic density, avoiding the long text from dominating the scoring result.

[0170] Optionally, step S5 of this embodiment specifically includes:

[0171] S5.1 By real-time monitoring of the unanswered queries in the Q&A process (such as the high-frequency questions of users not covered by the knowledge graph) and the inference paths of the low-confidence answers in step S4, locate the missing entities or contradictory relationships in the knowledge graph. At the same time, based on the predefined logic rule library (such as "the occupation of a doctor needs to be associated with a medical degree") and the graph consistency verification algorithm (such as subgraph isomorphism detection), automatically identify attribute conflicts (such as the "establishment time" and "dissolution time" of an entity overlapping) or relationship contradictions (such as "a person is the CEO of Company A" and "a person is a full-time employee of Company B").

[0172] S5.2 For the detected knowledge gaps, supplement the missing entities and relationships by automatically crawling external authoritative data sources (such as corporate annual reports, academic databases); for conflicting data, adopt a semi-automatic manual verification mechanism, mark the conflict points through visualization tools and provide correction suggestions (such as deleting incorrect edges or adding constraint conditions). In addition, through federated multi-source alignment technology, disambiguate entities and map relationships between the external knowledge base (such as GeoNames, Wikidata) and the existing knowledge graph to ensure data consistency and semantic compatibility.

[0173] S5.3 The updated knowledge graph is incorporated into the system incrementally (instead of full reconstruction) and ensures seamless switching through a version control mechanism. In the subsequent question-and-answer process (steps S2 / S3), the new knowledge graph is preferentially used to execute queries, and at the same time, the improvement effects of the new knowledge graph on the answer confidence and user satisfaction are recorded, forming a closed loop of "detection - correction - verification". For example, if a certain knowledge graph update improves the answer confidence of the query "CEO of a certain company" from 70% to 95%, then trigger the priority maintenance rule for this type of knowledge to achieve system self-optimization.

[0174] Corresponding to the above method, this embodiment also provides an intelligent question-and-answer system based on a knowledge graph, as Figure 3 shown, including:

[0175] A knowledge graph construction unit, configured to obtain multi-source heterogeneous data and a natural language question statement input by a user, and construct an initial knowledge graph according to the multi-source heterogeneous data; the multi-source heterogeneous data includes structured data and unstructured text;

[0176] A first-round query unit, configured to perform word segmentation, named entity recognition, and semantic analysis on the question statement, extract target entities and preliminary relationship requirements, and perform a single-hop query based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answer requirement, jump to the answer screening unit; otherwise, enter the second-round query unit;

[0177] A second-round query unit, based on an adaptive path extension algorithm, performs dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entities, the preliminary relationship requirements, and the first-round query result to obtain an extended inference path and associated evidence; the inference path includes a multi-hop entity relationship link;

[0178] An answer screening unit, configured to call a preset large language model, and perform semantic matching on the first-round query result or the inference path and the associated evidence in combination with an attention mechanism to obtain a list of candidate answers and corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores;

[0179] A knowledge graph update unit, which is used to correct entities, relationships and attributes in the initial knowledge graph to obtain an updated knowledge graph, and apply the updated knowledge graph to the first-round query unit and the second-round query unit to form a continuously optimized closed-loop system.

[0180] The beneficial effects of the present invention are as follows:

[0181] (1) The knowledge of the present invention has a wide coverage and strong timeliness. By fusing structured data (such as databases, APIs) and unstructured text (such as web pages, documents), an initial knowledge graph with a high coverage rate is constructed, breaking through the limitations of a single data source. In step S5, based on the knowledge gaps or conflict information in the user's question, the knowledge graph is corrected in real time (such as adding new entities, correcting attributes) to ensure the timeliness of the knowledge base.

[0182] (2) The complex reasoning of the present invention has high efficiency. The single-hop query in step S2 quickly responds to simple questions, and the dynamic path search in step S3 efficiently processes multi-hop reasoning through adaptive algorithms (such as depth control, pruning strategies), avoiding resource waste of traditional fixed strategies; based on a scoring formula of semantic similarity and logical coherence (such as the field strength superposition model), high-value paths are preferentially expanded to reduce invalid searches.

[0183] (3) The answer accuracy and interpretability of the present invention are both improved. Step S4 combines a large language model (such as GPT-4) with an attention mechanism to accurately calculate the semantic association between the user's question and the evidence, and screen out high-confidence answers; it supports the visual display of the reasoning path (such as an interactive graph), and users can trace the logical link of the answer generation to enhance trust.

[0184] (4) The system of the present invention has strong adaptability. Step S5 re-enters the updated knowledge graph into the question-answering process (S2-S3) to form a closed loop of "data-driven optimization - performance improvement - feedback iteration" to continuously improve the system performance; in links such as semantic matching and path expansion, parameters (such as attention weights, retrieval depth) are dynamically adjusted according to the input type (single-hop result or multi-hop path) to adapt to different scenario requirements.

[0185] (5) The present invention has high domain scalability and compatibility. Each step (such as knowledge construction, query, update) is decoupled and independent, and can be customized and extended for fields such as medical and finance (such as adding domain conflict detection rules); it supports the federated query and multi-source alignment of external knowledge bases (such as professional databases, industry graphs), providing in-depth support for vertical fields.

[0186] (6) The resource consumption of the present invention is controllable. In dynamic path search, through contradiction detection, path merging, and iteration count limitation, excessive consumption of computing resources is avoided; through structured text encoding and summary extraction, the computational load of the large language model is reduced to adapt to edge device deployment.

[0187] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0188] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An intelligent question-answering method based on a knowledge graph, characterized in that, Including: S1. Obtain multi-source heterogeneous data and a query statement in natural language input by a user, and construct an initial knowledge graph based on the multi-source heterogeneous data; The multi-source heterogeneous data includes structured data and unstructured text; S2. Perform word segmentation, named entity recognition, and semantic analysis on the query statement, extract target entities and preliminary relationship requirements, and perform a single-hop query based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answering requirement, jump to step S4; Otherwise, enter step S3; S3. Based on the adaptive path expansion algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entities, the preliminary relationship requirements, and the first-round query result to obtain an extended inference path and associated evidence; the inference path includes a multi-hop entity relationship link; S4. Invoke a preset large language model, and perform semantic matching on the first-round query result or the inference path and the associated evidence in combination with the attention mechanism to obtain a candidate answer list and corresponding semantic similarity scores, and determine the optimal answer in the candidate answers according to the semantic similarity scores; S5. Correct entities, relationships, and attributes in the initial knowledge graph to obtain an updated knowledge graph, and apply the updated knowledge graph to steps S2 and S3 to form a continuously optimized closed-loop system.

2. The intelligent question-answering method based on a knowledge graph according to claim 1, wherein Obtain multi-source heterogeneous data and a query statement in natural language input by a user, and construct an initial knowledge graph based on the multi-source heterogeneous data; The multi-source heterogeneous data includes structured data and unstructured text, including: Perform field parsing on the structured data, extract entity names, attributes, and relationships to obtain a classified structured data table; Perform sentence splitting and word segmentation on the unstructured text, and identify potential entities and relationships through natural language processing techniques to obtain an unstructured text entity relationship candidate set; Delete redundant fields in the structured data table, and perform standardized naming format and merge synonymous entities to obtain a cleaned structured data table; Filter entity relationship pairs in the unstructured text entity relationship candidate set with a confidence level less than a preset threshold, and eliminate ambiguous expressions to obtain a cleaned and normalized unstructured entity relationship pair; Uniformly convert entities, attributes, and relationships in the cleaned structured data table and the normalized unstructured entity relationship pair into standardized semantic triples; Identify the conflicting triples in the standardized semantic triples through pre-set conflict detection rules, and verify the conflicting triples to obtain the verified semantic triples; the conflict detection rules include: entity uniqueness rule, entity attribute consistency rule, relationship uniqueness rule, relationship logic constraint rule, attribute value type rule, attribute value range rule, context consistency rule, timeline consistency rule, multi-source data consistency rule, data source priority rule, incremental data conflict rule, and version consistency rule; the entity uniqueness rule is used to detect whether there are entities with the same name but different meanings in the knowledge graph. If a conflict is detected, unique identifiers are distinguished and assigned through context or attributes; the entity attribute consistency rule is used to detect whether the attribute values of the same entity conform to logical constraints. If a conflict is detected, it is marked as data to be corrected; the relationship uniqueness rule is used to detect whether there are multiple conflicting objects for the same pair of entities under a specific relationship. If a conflict is detected, the correct value is selected according to the pre-set priority; the relationship logic constraint rule is used to detect whether the relationship conforms to common sense or domain rules. If a conflict is detected, it is marked as an invalid relationship; the attribute value type rule is used to detect whether the attribute value conforms to the predefined data type. If a conflict is detected, it is marked as abnormal data; the attribute value range rule is used to detect whether the attribute value is within a reasonable range. If a conflict is detected, it is marked as out-of-range data; the context consistency rule is used to detect whether the attribute values of entities or relationships are consistent with the context information. If a conflict is detected, it is marked as context-contradictory data; the timeline consistency rule is used to detect whether the time attributes of entities or relationships conform to the timeline logic. If a conflict is detected, it is marked as timeline-abnormal data; the multi-source data consistency rule is used to detect whether the information of the same entity or relationship obtained from different data sources is consistent. If a conflict is detected, it is marked as multi-source contradictory data; the data source priority rule is used to select the correct value according to the pre-set data source credibility priority when multi-source data conflicts; the incremental data conflict rule is used to detect whether the newly added data is consistent with the historical data. If a conflict is detected, it is marked as incremental contradictory data; the version consistency rule is used to detect whether the core entities and relationships are consistent between different versions of the knowledge graph. If a conflict is detected, it is marked as version-abnormal data; Import the verified semantic triples into the graph database to construct the initial knowledge graph; the initial knowledge graph includes entity nodes, relationship edges, and attribute labels.

3. The intelligent question-answering method based on a knowledge graph according to claim 1, wherein, Perform word segmentation, named entity recognition, and semantic analysis on the question statement, extract the target entity and the preliminary relationship requirements, and execute a single-hop query based on the initial knowledge graph and the pre-set single-hop query request to generate the first-round query result. If the first-round query result meets the answering requirements, jump to step S4; Otherwise, enter step S3, including: Filter punctuation marks, unify case, and process special characters in the question statement to generate a normalized text; Perform word segmentation on the normalized text to generate the corresponding vocabulary sequence; Identify the target entities in the lexical sequence through a pre-trained named entity recognition model, and combine semantic role labeling technology to extract the semantic relationships between the target entities, obtaining a list of target entities and preliminary relationship requirements; Based on the preliminary relationship requirements, map the target entities to the nodes in the initial knowledge graph to generate a single-hop query request; the single-hop query request uses a graph query language; Execute a single-hop query in the initial knowledge graph to obtain the entities, relationships, and attributes directly associated with the target entities, obtaining a query result; Evaluate the confidence of the query result and determine the query result as the first-round query result; the evaluation indicators include: result coverage rate, relationship matching degree, and evidence sufficiency; if the result coverage rate ≥ 90% and the evidence sufficiency meets the standard, it is determined that "the first-round query result meets the answering requirements".

4. The intelligent question-answering method based on a knowledge graph according to claim 1, wherein Based on the adaptive path extension algorithm, perform dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entities, the preliminary relationship requirements, and the first-round query result, obtaining an extended inference path and associated evidence, including: Use the entities and relationships in the first-round query result as the initial nodes to construct a candidate path set; Filter the relationship types directly associated with the target entities according to the preliminary relationship requirements to generate an initial candidate path set; Use a path scoring function to calculate the comprehensive score of each path, and sort the candidate paths from high to low according to the comprehensive score, and preferentially expand the high-scoring paths to obtain a sorted candidate path queue; Adjust the maximum expansion depth of the sorted candidate path queue according to the current path score and a preset threshold; At each hop expansion, only retain the relationship types relevant to the semantics of the user's question, filter out irrelevant branches, perform a graph traversal query on each path to obtain the next-hop entities and relationships, obtaining an extended multi-hop path set and new entity relationship evidence; Perform redundant pruning, contradiction detection, and evidence fusion on the extended multi-hop path set to obtain the optimized inference path and associated evidence.

5. The intelligent question answering method based on a knowledge graph according to claim 4, wherein The path scoring function Score(P) is: Score(P) = α·Rel_Confidence + β·Sem_Similarity + γ·Path_Length_Penalty; Among them, Rel_Confidence is the average confidence of the relationships in the path, and the average confidence is the statistical weight from the initial knowledge graph; Sem_Similarity is the cosine similarity between the path semantics and the user's question; Path_Length_Penalty is the path length penalty factor, Path_Length_Penalty = e -0.5n , where n is the number of hops; α, β, and γ are all dynamic weights, and their initial values are set to 0.6, 0.3, and 0.1 respectively.

6. The intelligent question answering method based on a knowledge graph according to claim 5, wherein The formula for the maximum expansion depth MaxDepth is:

7. The intelligent question answering method based on a knowledge graph according to claim 6, characterized in that When the path length > 3, increase γ to 0.2 to suppress overly long paths; if the user's question contains clear relationship keywords, increase α to 0.

7.

8. The intelligent question-answering method based on a knowledge graph according to claim 1, wherein Call a preset large language model, and combine the attention mechanism to perform semantic matching on the first-round query result or the inference path and the associated evidence to obtain a list of candidate answers and the corresponding semantic similarity scores, and determine the optimal answer in the candidate answers according to the semantic similarity scores, including: Encode the first-round query result into a structured text sequence; or encode the inference path and the associated evidence into a structured text sequence; Call the pre-trained large language model GPT-4, and use the unidirectional attention mechanism to perform semantic matching on the structured text sequence to obtain a list of candidate answers and corresponding semantic similarity scores; Determine the candidate answer in the candidate answer list corresponding to the highest semantic similarity score as the optimal answer.

9. The intelligent question answering method based on a knowledge graph according to claim 8, wherein The calculation formula of the semantic similarity score Sim(Q,E) is: Among them, Q is the user's question, the natural language question input by the user, which is decomposed into n semantic units after encoding; E is the structured text sequence, which is decomposed into m semantic units after encoding; q is the i-th semantic unit of the question, and e j is the j-th semantic unit of the evidence; φ(q i ) is the question field strength function, φ(q i ) = ||q i ||·exp(-λ·Entropy(q i ))), where Entropy(q i ) is the information entropy of the semantic unit, and λ is the attenuation coefficient; ψ(e j ) is the evidence field strength function, ψ(e j ) = ||e j ||·exp(-λ·Entropy(e j )); Dist(q i , e j ) is the semantic unit distance, where Σ is the covariance matrix, pre-computed through a large-scale corpus; γ' is the distance attenuation exponent, and γ' is 2; Z is the normalization factor, used to scale the score to the [0, 1] interval.

10. An intelligent question answering system based on a knowledge graph, characterized in that, including: A knowledge graph construction unit for obtaining multi-source heterogeneous data and a natural language question statement input by the user, and constructing an initial knowledge graph based on the multi-source heterogeneous data; The multi-source heterogeneous data includes structured data and unstructured text; The first-round query unit is used to tokenize, name entity recognition and semantic analysis of the question statement, extract the target entity and preliminary relationship requirements, and perform a single-hop query based on the initial knowledge graph and a preset single-hop query request to generate a first-round query result. If the first-round query result meets the answering requirements, jump to the answer screening unit; Otherwise, enter the second-round query unit; The second-round query unit, based on the adaptive path extension algorithm, performs dynamic path search and multi-hop iterative query in the initial knowledge graph according to the target entity, the preliminary relationship requirements and the first-round query result to obtain an extended inference path and associated evidence; the inference path includes a multi-hop entity relationship link; The answer screening unit is used to call a preset large language model, and perform semantic matching on the first-round query result or the inference path and the associated evidence in combination with the attention mechanism to obtain a list of candidate answers and corresponding semantic similarity scores, and determine the optimal answer among the candidate answers according to the semantic similarity scores; The knowledge graph update unit is used to correct the entities, relationships and attributes in the initial knowledge graph to obtain an updated knowledge graph, and apply the updated knowledge graph to the first-round query unit and the second-round query unit to form a continuously optimized closed-loop system.

Citation Information

Cited By

  • Multi-modal data dynamic reasoning system and method based on cognitive map

    CN120471179A

  • Knowledge base management system and enhancement method

    CN120525039A

  • Time sequence knowledge graph federal collaborative optimization method, system and device and storage medium

    CN120611070A

  • Temporal knowledge graph federated collaborative optimization method, system, device and storage medium

    CN120611070B

  • Demand-driven data resource dynamic identification and discovery method

    CN120671853A