A knowledge graph enhanced question-answer driven event analysis method, system, computer and medium

CN122777641APending Publication Date: 2026-09-18SUZHOU AEROSPACE INFORMATION RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512055506.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]综上所述,现有技术普遍存在以下不足:其一,缺乏一种能够主动响应用户问题、驱动事件要素动态建模的机制;其二,单独依赖知识图谱或大语言模型,均难以同时满足动态性、准确性和可解释性的要求

Benefits of technology

[0044] (1) In the information extraction stage, a hybrid extraction mechanism is proposed, which combines the large language model with the rule base to achieve high-precision extraction of entity, attribute and event elements, while improving the stability of the extraction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777641A_ABST
    Figure CN122777641A_ABST
Patent Text Reader

Abstract

The application discloses a kind of knowledge graph enhanced question-answer driven event analysis methods, including data acquisition, entity and attribute extraction, relationship and event extraction, knowledge graph construction and question-answer and prediction steps.This method is dynamically event modeling by question-answer driving, combined with knowledge graph and retrieval enhancement (RAG) mechanism, realizes the fine-grained analysis and trend prediction of event elements, outputs answer and evidence chain, improves accuracy, interpretability and reliability, and is suitable for news event research and judgment, public security and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and natural language processing technology, specifically relating to a knowledge graph-enhanced question-answering driven event analysis method, system, computer, and medium. Background Technology

[0002] With the rapid development of news media and social media platforms, the sources of news information are becoming increasingly diverse, and the speed of information updates is constantly accelerating, leading to the challenge of large-scale, multi-source, heterogeneous data in event analysis. Most existing technologies are based on keyword retrieval, topic clustering, or shallow text classification methods to organize and present information from news reports. While these methods can achieve basic information aggregation, they often only provide fragmented results when revealing deep semantic connections across reports and time periods, making it difficult to support dynamic modeling and trend prediction of event elements.

[0003] Some existing solutions attempt to improve event modeling using knowledge graphs. Knowledge graphs utilize a triple structure of "entity-relationship-attribute" to provide structured storage and retrieval support for people, objects, and their relationships in news events. However, these methods primarily rely on static extraction and offline construction, which can easily lead to insufficient coverage and update lag, making it difficult to meet the needs of real-time evolution of news events. Furthermore, existing knowledge graphs mostly focus on static relationships, lacking a dynamic depiction of event sequence and evolution.

[0004] Existing technologies also incorporate large language models, leveraging their powerful semantic understanding and generation capabilities for event element identification and answer generation. However, practice shows that directly relying on large models to generate results can easily lead to unclear logical chains and inconsistent results. Furthermore, when processing long texts or multiple reports, there are issues with the loss of contextual information and the omission of key information. These shortcomings limit the application of large models in the dynamic analysis and prediction of complex news events.

[0005] In summary, existing technologies generally suffer from the following shortcomings: First, they lack a mechanism that can proactively respond to user questions and drive dynamic modeling of event elements; second, relying solely on knowledge graphs or large language models is insufficient to simultaneously meet the requirements of dynamism, accuracy, and interpretability. Therefore, a dynamic method that proactively drives event analysis and combines knowledge graphs and large language models is needed. Summary of the Invention

[0006] The purpose of this invention is to provide a knowledge graph-enhanced question-answering driven event analysis method, system, computer, and medium.

[0007] The technical solution to achieve the purpose of this invention is: a knowledge graph-enhanced question-answering driven event analysis method, comprising the following steps:

[0008] Step S1: Data Collection and Cleaning: Raw text data of news events is obtained from multiple sources, including news media, social media platforms, and online forums. After noise filtering, format standardization, and structured sentence segmentation, a cleaned news event text database is constructed. ;

[0009] Step S2, Entity and Attribute Extraction: A hybrid strategy combining large language models, rule bases, and dictionary matching is employed to extract entities and attributes from the text database. Extract people, organizations, objects, geography, events, time, and numerical entities from the data;

[0010] Step S3, Unified Entity Representation: For each entity Perform type determination and extract a set of attributes including role, geographical location, quantity, time, status, and parameters. Based on the alias merging mechanism, different representational entities pointing to the same real-world object are mapped to a unified, normalized entity. ;

[0011] Step S4, Relationship and Event Chain Construction: Based on Normalized Entities Based on the entity's attributes, semantic relationships between entities are extracted through three mechanisms: syntactic dependency analysis, event rule template matching, and large language model completion of latent logic. , generate The event triples, formally represented, constitute the event triple set R, where the relation It belongs to the set of relations P, which includes action, state, cause and effect, time, and space.

[0012] Step S5: Knowledge graph construction and updating with schema constraints: Predefine a knowledge graph schema S, which specifies the legal entity type-relationship-entity type combinations; perform a validity check on each event triple, store triples that meet the check conditions in the graph database to form a news event knowledge graph KG, and support dynamic updates;

[0013] Step S6, Question-Answer Driven Event Analysis: Receive the user's natural language question Q. Through three mechanisms—semantic decomposition, event element mapping, and trend prediction extension based on a large language model—question Q is parsed into a set of multiple information demand sub-questions. ;

[0014] Step S7, Answer Prediction Generation: Employing a retrieval-enhanced generation mechanism, each sub-question q is predicted and generated. j With text library The set of fragments D obtained after sentence segmentation is vectorized and similarity is compared to select the top K most relevant candidate fragments. This is then concatenated with relevant information retrieved from the knowledge graph KG to form a fusion context C; based on this fusion context C and the knowledge graph KG, a conditional probability model is used to calculate and select... The candidate answer A with the highest probability is output; at the same time, the set of text fragments actually used in the process of generating answer A is also output. As a chain of evidence.

[0015] Furthermore, in step S3, the alias merging mechanism includes character format normalization, semantic similarity calculation based on vector representation, alias dictionary matching, and attribute consistency verification. The specific methods are as follows:

[0016] Character and format normalization: unify the case and full-width / half-width characters in entity representations, and remove modifiers;

[0017] Semantic similarity determination: Calculating entity similarity and vector representation and cosine similarity If the similarity is greater than a preset threshold, it is determined to be a candidate for the same entity;

[0018] Alias ​​dictionary matching steps: Standardize the mapping of entity representations using a pre-built alias dictionary;

[0019] Attribute consistency check: Compare the attribute sets of candidate entities. and If the similarity exceeds a preset threshold, the entities are ultimately identified as the same entity, and a mapping relationship is established. .

[0020] Furthermore, in step S4, syntactic dependency analysis is used to identify explicit semantic connections in the subject-verb-object and prepositional phrase structures in the text; event rule template matching uses a predefined relation template library to identify specific event behaviors; and large language model completion is used to identify complex relationships in the text that are not explicitly stated but logically implicit.

[0021] Furthermore, in step S5, a validity check is performed on each event triple, and the judgment formula is as follows:

[0022]

[0023] This condition corresponds to the Schema validation function. Its definition is:

[0024]

[0025] A value of 1 indicates a ternary combination method, which is written into the knowledge graph; a value of 0 indicates an invalid method, which is not written.

[0026] Furthermore, in step S7, the enhanced generation mechanism is retrieved, specifically including:

[0027] S7-1, Vectorized Representation: Using the same semantic encoder function f_enc() to represent subproblems q j and text fragment d i Encode into vectors respectively and ;

[0028] S7-2, Similarity Calculation and Ranking: Calculating Subproblem Vectors With each text fragment vector cosine similarity The fragments in the fragment set D are sorted from high to low based on their similarity scores;

[0029] S7-3, Sorting and Deduplication: Select the K most similar segments to form the initial candidate set C. K ; For C K Execute the redundancy removal function f_dedup(), whereby redundancy removal includes: if any two segments and vector similarity Exceeding the threshold or the overlap of its entity set Exceeding the threshold If a segment is found to be semantically duplicated, one of the segments is removed, resulting in a deduplicated candidate segment set. ;

[0030] S7-4, Context Concatenation: The candidate fragment set is concatenated using the concatenation function f_Concat(). The information is then fused with the entity, relation, and attribute information related to the sub-problem retrieved from the knowledge graph (KG) to form the final fusion context C.

[0031] Furthermore, in step S7, the answer generation and evidence chain output are performed using the following methods:

[0032] Let C be the fusion context obtained through retrieval enhancement and generation processing, KG be the news event knowledge graph, and Q be the user's original question;

[0033] The answer generation process is represented as follows: Where A is the generated optimal answer;

[0034] Chain of evidence The generation process is represented as: ,in For indicator functions, when fragment Used to generate answers The value is 1 when the condition is met, and 0 otherwise; the evidence chain set Output the answer A*.

[0035] A knowledge graph-enhanced question-answering driven event analysis system, used to implement the aforementioned knowledge graph-enhanced question-answering driven event analysis method, includes:

[0036] The data acquisition and cleaning module is used to perform step S1;

[0037] The entity, attribute extraction and unified representation module is used to execute steps S2 and S3;

[0038] The relationship and event chain construction module is used to execute step S4;

[0039] The knowledge graph construction and update module with schema constraints is used to execute step S5;

[0040] The question-and-answer driven event analysis and prediction module is used to execute steps S6 and S7.

[0041] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the knowledge graph-enhanced question-answering driven event analysis method.

[0042] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the knowledge graph-enhanced question-answering driven event analysis method.

[0043] Compared with the prior art, the significant advantages of this invention are:

[0044] (1) In the information extraction stage, a hybrid extraction mechanism is proposed, which combines the large language model with the rule base to achieve high-precision extraction of entity, attribute and event elements, while improving the stability of the extraction results.

[0045] (2) Regarding the question-and-answer mechanism, an event-based predictive question-and-answer framework was designed. This framework can not only answer the questions entered by users, but also generate key questions related to news events in advance, and make inferences about the possible development direction of events, thereby providing users with forward-looking references.

[0046] (3) In the answer generation stage, a search-enhanced generation (RAG) mechanism was introduced to achieve dual-source fusion of text database and knowledge graph. Through steps such as text segmentation, sub-question vectorization comparison, context splicing, sorting and deduplication, the accuracy and controllability of the answer generation process were ensured.

[0047] (4) At the output level, a schema constraint and verification mechanism are introduced to ensure the consistency and traceability of the output, thereby improving the interpretability and reliability of the system in news event analysis and prediction scenarios. Attached Figure Description

[0048] Figure 1 This is a flowchart of the question-and-answer driven event analysis with knowledge graph enhancement according to the present invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] This invention constructs a dedicated knowledge graph for news events, achieving high-precision identification and coverage of elements such as people, objects, and devices, thus overcoming the limitations of existing technologies in dynamic event analysis. Simultaneously, this invention introduces a controllable question-and-answer prediction mechanism, ensuring stability and consistency in the process of question generation and answer output, reducing logical inconsistencies and result fluctuations. Furthermore, this invention supports evidence chain tracing, guaranteeing the interpretability and credibility of the output results, thereby fulfilling application needs such as news event analysis, trend analysis, and decision support.

[0051] A knowledge graph-enhanced question-answering driven event analysis method includes the following steps.

[0052] Step S1: Collect and clean data from multiple sources such as news media, social platforms and forums through the data acquisition module to build a news event text library.

[0053] Let the original text set be:

[0054]

[0055] Cleaning and preprocessing functions The text undergoes noise filtering, formatting standardization, and structured sentence segmentation, including but not limited to: removing HTML tags, URLs, script characters, and emojis; handling encoding differences and punctuation conventions; filtering duplicate content; performing language recognition and invalid text removal; and segmenting sentences according to syntactic boundaries. The resulting cleaned text set is as follows:

[0056]

[0057] in, For the original text collection, This is the cleaned text collection.

[0058] Step S2.1: Clean the text The input entity extraction module identifies core entities involved in the text, including people entities (individuals, groups, public figures, organization members), organizations entities (government agencies, enterprises, media, public security departments, etc.), objects entities (goods, resources, weapons, vehicles, equipment, etc.), geographical entities (countries, cities, administrative regions, landmarks, locations), event entities (news events, emergencies, actions, activities), time entities (dates, time periods, points in time), and numerical entities (amounts, statistical values, quantity expressions).

[0059] Named entity recognition function A hybrid strategy of large language model extraction, rule base correction, and dictionary matching is adopted. The LLM provides semantic-level recognition capabilities, while the rule base provides industry-specific keyword matching; the combination of the two improves recall and precision. The entity set is represented as:

[0060]

[0061] in, It includes core entities such as people, objects, and equipment.

[0062] Step S3: Enter the type and attribute extraction module, and extract each entity e. i Perform type determination and attribute extraction. Attributes include, but are not limited to: role, geographical location, quantity, time, status, and important parameters. Attributes and type are determined using an attribute extraction function. get:

[0063]

[0064] in, Representing entities The set of attributes, This represents the normalized entity after merging aliases.

[0065] To avoid multiple representations of the same entity in the knowledge graph, this invention employs an alias merging and normalization mechanism. Normalized entities. Generate through the following steps:

[0066] First, unify the characters and formatting, standardize uppercase and lowercase, full-width and half-width characters, and remove modifiers;

[0067] Next, semantic similarity is determined, calculated based on vector representation:

[0068]

[0069] If the value is greater than the threshold, it is considered to be the same entity.

[0070] Alias dictionary matching is performed for reuse, such as "New York" = "New York".

[0071] Finally, attribute consistency check is performed, and whether the entities are the same entity is confirmed by comparing the similarity of attribute sets. The entity mapping relationship is expressed as:

[0072]

[0073] Step S4: Enter the relation and event extraction module, identify the semantic relations and event logic between entities, and generate a triple-based event chain.

[0074]

[0075] Wherein, is a relation set, including but not limited to action relations (e.g., attack, arrest, meet, inform, announce), state relations (e.g., located at, belong to, contain, affected by), causal relations (e.g., lead to, cause, trigger, originate from), temporal relations (e.g., occurred on, started on, ended on) and spatial relations (e.g., located at, come from, lead to).

[0076] represents the extracted event triples. Relation extraction adopts three types of mechanisms: the first is syntactic dependency analysis to identify semantic connections in subject-predicate-object and preposition-object relations; the second is event rule template matching, which uses a relation template library to identify event behaviors; additionally, large language models complete complex relations to identify potential implicit event logics, such as "related to", "allegedly involved in" etc.

[0077] Step S5: Input the triples into the knowledge graph construction module, perform standardized storage and management under Schema constraints, and form a dynamically updatable news event knowledge graph.

[0078] Schema S defines legal entity type-relation-entity type combinations, such as: (person, arrest, person), (organization, announce, event), (object, belong to, organization), (event, occur at, location). Schema validity check shall be performed for each triple before construction. The judgment formula is as follows:

[0079]

[0080] This judgment formula corresponds to the Schema check function , which is defined as:

[0081]

[0082] Wherein, is the Schema constraint set of the knowledge graph.

[0083] Step S6: When the user inputs a natural language question Q, the question answering and prediction module parses it and generates several information demand sub-questions. This process employs three mechanisms: semantic decomposition, event element decomposition, and predictive expansion, to perform fine-grained structuring of the user input, ensuring that the information demand is fully covered and clearly expressed.

[0084] First, based on dependency parsing, referential resolution, and question type identification, the system breaks down complex natural language questions into multiple semantic atomic questions. Then, according to the schema of the event knowledge graph, the system maps the questions to relevant event elements, including subject, behavior, time, location, and result, and automatically generates further extended requirements. Finally, the system uses a large language model to predict the trends of news events and supplements possible inference questions to improve the overall completeness of the retrieval.

[0085] Ultimately, the result of the problem analysis can be expressed as:

[0086]

[0087] in, This is a set of sub-problems related to the requirements.

[0088] Step S2.2: Generate a set of subproblems {q} j The cleaned news text library T′ is input into the RAG retrieval module along with the cleaned news text library T′ as input for subsequent vectorization, similarity comparison, and candidate segment selection.

[0089]

[0090] The RAG module will perform vectorization, fragment sorting and deduplication, and knowledge graph fusion on each sub-question to obtain the final contextual semantic material used for answer generation.

[0091] Step S7 (RAG sub-process):

[0092] S7-1 Vectorization: The RAG module first vectorizes the requirement sub-problems and the sentence-level or semantic unit-level fragments after text library segmentation. The text library T′, after sentence segmentation and semantic fragmentation, yields a set of fragments:

[0093]

[0094] Pair problem q j and text fragment d i Apply the same semantic encoder function f enc(), which can be a Transformer-based semantic embedding model, such as BERT, RoBERTa, Sentence-BERT, or a domain-specific custom embedding model. The vectorized result is:

[0095]

[0096] in, The vector representation of the demand subproblem Vector representation of a text segment For encoder functions (such as BERT, Word2Vec, etc.).

[0097] S7-2 Sorting and Deduplication: The system is based on the vector of demand subproblems. Text fragment vector The similarity between segments is ranked to determine the candidate segments most relevant to the sub-problem. Cosine similarity is used to calculate the similarity.

[0098]

[0099] Based on similarity ranking, the top K most relevant segments are selected from the segment set D to form a candidate set C. K The Top-K screening function is defined as follows:

[0100]

[0101] To avoid semantic duplication of candidate segments, the system executes a redundancy removal mechanism f dedup Redundancy determination includes: vector similarity deduplication; if the vector similarity between two segments exceeds a threshold... Entity overlap deduplication: If fragments contain entities with the same height, they are considered semantic duplicates, as follows:

[0102]

[0103] And content overlap deduplication, that is, identifying synonymous expressions or repeated descriptions through rules. After this judgment process, an optimized set of candidate segments is obtained.

[0104] S7-3 Context Concatenation: To improve the accuracy and interpretability of answer generation, the system merges the filtered text fragments with relevant triples in the knowledge graph to construct the final context material. Define the context concatenation function. :

[0105]

[0106] in, The final context combines candidate fragments with knowledge graph information.

[0107] Step S8: The question-answering and prediction module generates an answer based on the knowledge graph and RAG results, and outputs the corresponding evidence chain. The evidence chain provides traceable source information to ensure the reliability and transparency of the results. Finally, the answer is returned to the user.

[0108] In the answer generation stage of this invention, the system determines the final answer based on a formula. Specifically, the question input by the user is denoted as... The context material output by the RAG module is denoted as The knowledge graph of news events is denoted as For all possible candidate answers Calculate its conditional probability And select the answer that maximizes the conditional probability as the final output, that is:

[0109]

[0110] For questions entered by the user;

[0111] Contextual material after RAG processing;

[0112] For knowledge graphs;

[0113] This is the optimal answer.

[0114] To ensure the interpretability and reliability of the results, this invention generates a chain of evidence for tracing the source while outputting the answer. Specifically, let... Given a set of candidate text fragments obtained through Top-K filtering, if a certain fragment... In generating answers If it is actually used in the process, then it is included in the chain of evidence. The formal representation is as follows:

[0115]

[0116] in, For indicator functions, when fragment Used to generate answers The value is 1 if the condition is met, and 0 otherwise. (Evidence chain set) This will be returned as additional output along with the answer, making it easier for users to trace the source of the answer.

[0117] This invention discloses a knowledge graph-enhanced question-answering driven event analysis system, comprising multiple functional modules that cooperate to collect, process, and analyze news event information. The system first includes a data acquisition module, which is responsible for obtaining text data related to news events from multiple sources such as news media, social platforms, and online forums, providing rich raw information for subsequent processing stages.

[0118] Based on data collection, the system includes an entity and attribute extraction module. This module identifies core entities in news texts by calling a large language model and further extracts their corresponding attributes. For situations where the same entity has multiple representations, this module incorporates an alias merging mechanism to unify the representation of identical entities, thereby improving the accuracy and consistency of the extraction results.

[0119] The system also includes a relation and event extraction module. This module utilizes the identified entity and attribute information to mine semantic relationships between entities from text and identify the logic of event occurrence, thereby constructing event chains based on triples. This module enables the structured representation of events, providing foundational data for subsequent knowledge graph construction.

[0120] Building upon this foundation, the knowledge graph construction module stores and manages the extracted triples. This module uses schema constraints to ensure the normalization and scalability of the triple representation in the graph database, thereby forming a news event knowledge graph that can be used for querying, reasoning, and updating.

[0121] Finally, the system includes a question-answering and prediction module. This module addresses user interaction needs, and its main functions include question parsing and generation, information requirement decomposition, retrieval and aggregation, and answer generation and evidence traceability. In this module, the user-inputted question is first parsed, and several sub-questions are generated based on semantics. After completing retrieval and aggregation, the system can output results containing the answer and evidence chain, thereby ensuring the transparency and traceability of the results.

[0122] To further improve the accuracy and stability of answer generation, a Retrieval-Augmented Generation (RAG) mechanism is integrated into the question answering and prediction modules. In this mechanism, the system segments the news text library according to sentences or semantic units, and vectorizes and compares the similarity between the sub-questions and the segmented text fragments to filter highly relevant candidate materials. Subsequently, the candidate texts are concatenated with context and combined with related information from the knowledge graph to form complete semantic materials. Redundancy is then eliminated through sorting and deduplication. This processing results in a more concise, accurate, and comprehensive context input to the generation model, enabling the answer generation process to utilize both text libraries and knowledge graphs simultaneously, thus improving the accuracy, stability, and interpretability of the results.

[0123] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the knowledge graph-enhanced question-answering driven event analysis method.

[0124] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the knowledge graph-enhanced question-answering driven event analysis method.

[0125] In summary, this invention can identify and model the relationships of key elements such as people, objects, and equipment in news texts, further analyze the dynamic evolution of their behavior and state, and combine semantic reasoning to predict the future development trend of events. It is applicable to intelligent application scenarios involving fine-grained analysis and prediction of event elements, such as news event analysis, public safety management, emergency command decision-making, and other applications. Its specific advantages are as follows:

[0126] (1) Deep Semantic Association Modeling Capability: Unlike traditional methods that only focus on keyword matching or shallow classification in cross-text and cross-event semantic mining, this invention combines the structured representation of knowledge graphs with the semantic understanding capabilities of large language models to identify and extract elements such as people, objects, and equipment from news texts, and establish attributes, relationships, and event chains on this basis. This enables the accurate construction of logical networks and semantic chains between events, thereby improving the depth and systematicity of news event analysis.

[0127] (2) Controllability and stability of question-answering prediction mechanism: Existing methods that rely solely on large language models to directly generate answers often suffer from uncontrollable output, incoherent logic, and insufficient interpretability of results. This invention proposes a "question-driven" prediction mechanism, which combines templated question generation, information demand decomposition, and evidence retrieval and aggregation to achieve stability and consistency in the question-answering process, effectively reducing logical defects and uncertainty of results.

[0128] (3) Result interpretability and evidence traceability: Unlike existing methods that mostly remain at the level of "black box" output, this invention provides a corresponding chain of evidence while generating the answer. Users can directly trace the source of the answer, thereby improving the transparency of the system, enhancing the interpretability and credibility of the results, and better meeting the reliability requirements of news event analysis and decision support.

[0129] (4) Ability to fuse multi-source heterogeneous data: Traditional solutions have limitations in terms of data sources, usually relying on only a single data source, which makes it difficult to fully reflect the background of news events. The technical solution of this invention can simultaneously collect and process text information from news media, social platforms and other public channels, and achieve the fusion of heterogeneous data through the structured storage and unified schema management of knowledge graphs, thereby ensuring the coverage and systematic nature of event analysis.

[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method of knowledge graph enhanced question-driven event analysis, characterized in that, Includes the following steps: Step S1, data collection and cleaning: Obtain original text data of news events from multi-source channels such as news media, social platforms and network forums, and after noise filtering, format unification and structured sentence processing, construct a cleaned news event text library ; Step S2, Entity and Attribute Extraction: A hybrid strategy combining large language models, rule bases, and dictionary matching is employed to extract entities and attributes from the text database. Extract people, organizations, objects, geography, events, time, and numerical entities from the data; Step S3, Unified Entity Representation: For each entity Perform type determination and extract a set of attributes including role, geographical location, quantity, time, status, and parameters. ; Based on the alias merging mechanism, different representational entities pointing to the same real-world object are mapped to a unified, normalized entity. ; Step S4, Relationship and Event Chain Construction: Based on Normalized Entities Based on the entity's attributes, semantic relationships between entities are extracted through three mechanisms: syntactic dependency analysis, event rule template matching, and large language model completion of latent logic. , generate The event triples, formally represented, constitute the event triple set R, where the relation It belongs to the set of relations P, which includes action, state, cause and effect, time, and space. Step S5: Knowledge graph construction and updating with schema constraints: Predefine a knowledge graph schema S, which specifies the legal entity type-relationship-entity type combinations; perform a validity check on each event triple, store triples that meet the check conditions in the graph database to form a news event knowledge graph KG, and support dynamic updates; Step S6, Question-Answer Driven Event Analysis: Receive the user's natural language question Q. Through three mechanisms—semantic decomposition, event element mapping, and trend prediction extension based on a large language model—question Q is parsed into a set of multiple information demand sub-questions. ; Step S7, Answer Prediction Generation: Employing a retrieval-enhanced generation mechanism, each sub-question q is predicted and generated. j With text library The set of fragments D obtained after sentence segmentation is vectorized and similarity is compared to select the top K most relevant candidate fragments. This is then concatenated with relevant information retrieved from the knowledge graph KG to form a fusion context C; based on this fusion context C and the knowledge graph KG, a conditional probability model is used to calculate and select... The candidate answer A with the highest probability is output; at the same time, the set of text fragments actually used in the process of generating answer A is also output. As a chain of evidence.

2. The knowledge graph-enhanced question-answering driven event analysis method according to claim 1, characterized in that, In step S3, the alias merging mechanism includes character format normalization, semantic similarity calculation based on vector representation, alias dictionary matching, and attribute consistency verification. The specific methods are as follows: Character and format normalization: unify the case and full-width / half-width characters in entity representations, and remove modifiers; Semantic similarity determination: Calculating entity similarity and vector representation and cosine similarity If the similarity is greater than a preset threshold, it is determined to be a candidate for the same entity; Alias ​​dictionary matching steps: Standardize the mapping of entity representations using a pre-built alias dictionary; Attribute consistency check: Compare the attribute sets of candidate entities. and If the similarity exceeds a preset threshold, the entities are ultimately identified as the same entity, and a mapping relationship is established. .

3. The knowledge graph-enhanced question-answering driven event analysis method according to claim 1, characterized in that, In step S4, syntactic dependency analysis is used to identify explicit semantic connections in the subject-verb-object and prepositional phrase structures in the text; event rule template matching uses a predefined relation template library to identify specific event behaviors; Large language model completion is used to identify complex relationships in text that are not explicitly stated but are logically implied.

4. The knowledge graph-enhanced question-answering driven event analysis method according to claim 1, characterized in that, In step S5, a validity check is performed on each event triple, and the judgment formula is as follows:

5. This conditional statement corresponds to the schema validation function. Its definition is: ; in, A value of 1 indicates a ternary combination method, which will be written into the knowledge graph; a value of 0 indicates an invalid method, which will not be written.

6. The knowledge graph-enhanced question-answering driven event analysis method according to claim 1, characterized in that, In step S7, the retrieval enhancement generation mechanism specifically includes: S7-1, Vectorized Representation: Using the same semantic encoder function f_enc() to represent subproblems q j and text fragment d i Encode into vectors respectively and ; S7-2, Similarity Calculation and Ranking: Calculating Subproblem Vectors With each text fragment vector cosine similarity The fragments in the fragment set D are sorted from high to low based on their similarity scores; S7-3, Sorting and Deduplication: Select the K most similar segments to form the initial candidate set C. K ; For C K Execute the redundancy removal function f_dedup(), whereby redundancy removal includes: if any two segments and vector similarity Exceeding the threshold or the overlap of its entity set Exceeding the threshold If a segment is found to be semantically duplicated, one of the segments is removed, resulting in a deduplicated candidate segment set. ; S7-4, Context Concatenation: The candidate fragment set is concatenated using the concatenation function f_Concat(). The information is then fused with the entity, relation, and attribute information related to the sub-problem retrieved from the knowledge graph (KG) to form the final fusion context C.

7. The knowledge graph-enhanced question-answering driven event analysis method according to claim 1 or 5, characterized in that, In step S7, the answer is generated and the evidence chain is output, specifically using the following method: Let C be the fusion context obtained through retrieval enhancement and generation processing, KG be the news event knowledge graph, and Q be the user's original question; The answer generation process is represented as follows: Where A is the generated optimal answer; Chain of evidence The generation process is represented as: ,in For indicator functions, when fragment Used to generate answers The value is 1 when the condition is met, and 0 otherwise; the evidence chain set Output the answer A*.

8. A knowledge graph-enhanced question-answering driven event analysis system, characterized in that, The question-answering driven event analysis method for knowledge graph enhancement according to any one of claims 1-6 includes: The data acquisition and cleaning module is used to perform step S1; The entity, attribute extraction and unified representation module is used to execute steps S2 and S3; The relationship and event chain construction module is used to execute step S4; The knowledge graph construction and update module with schema constraints is used to execute step S5; The question-and-answer driven event analysis and prediction module is used to execute steps S6 and S7.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the knowledge graph enhancement question-answering driven event analysis method as described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the knowledge graph enhancement question-answering driven event analysis method as described in any one of claims 1 to 6.