An event association analysis method and system based on large language model
By combining large language model and event correlation analysis methods with multiple inference methods, the problem of insufficient accuracy and generalization capabilities of the existing technology in the recognition of complex event relationships is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510127770.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The accuracy and generalization ability of existing event association analysis methods are limited in the face of semantic structures and long-distance dependencies of complex events, especially in the case of sparse data and difficulty in labeling.
A large language model-based event correlation analysis method is proposed. Through the combination of semantic analysis, sentence rewriting, sentence segmentation processing, RAG technology search, marginalized reasoning and Softmax classifier, sentence representation of event pairs is generated and multi-step reasoning is carried out to identify causality, succession and superior-understanding relationships.
It improves the accuracy and robustness of event correlation analysis, can capture the complex semantic relationships between events more accurately, and enhances the generalization ability and classification performance of the model.
Smart Images

Figure CN119577139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to an event association analysis method and system based on a large language model. Background Art
[0002] With the development of big data and artificial intelligence technology, event association analysis is increasingly used in the field of natural language processing, especially in scenarios such as news analysis, legal judgments, and policy analysis that require mining of a large number of event relationships. However, traditional event relationship extraction methods often rely on manual rules or simple machine learning models, which have significant deficiencies when faced with the semantic structure and long-distance dependencies of complex events, and it is difficult to accurately identify complex relationships such as causality, sequence, and hierarchy between events. Especially when the event description is vague or the information is incomplete, the accuracy and generalization ability of traditional models are severely restricted.
[0003] Most of the existing event association analysis methods are based on small-scale annotated data sets, while in real applications, event data usually have problems such as data sparsity and difficulty in annotation. Classic methods based on feature engineering or shallow learning have difficulty in dealing with data sparsity, and are prone to losing semantic information in high-dimensional text features. With the rise of pre-trained large language models, researchers have begun to use the powerful feature extraction capabilities of these models to process event texts. However, most existing studies only focus on feature extraction or classification of a single model, ignoring the comprehensive application of reasoning results from different models. In addition, because existing data augmentation strategies rely too much on artificially generated rules, the generated data lacks diversity, which limits the learning ability of the model, especially in low-resource environments. Therefore, developing an event association analysis system that combines a large model and multiple reasoning methods can more accurately capture the complex semantic relationships between events and improve the robustness and generalization ability of the model. Summary of the invention
[0004] In view of the shortcomings in the current event association analysis tasks, the present invention proposes an event association analysis method and system based on a large language model, aiming to more accurately capture the complex relationships between events.
[0005] The present invention also discloses an event association analysis method based on a large language model, which comprises:
[0006] Step 1: Use a large language model to perform semantic analysis on the context of each event pair in the event dataset. The large language model explains and describes the potential relationship between each pair of events based on event triggers, context, and the temporal or logical order between events.
[0007] Step 2: By keeping the original event relationship unchanged, the large language model rewrites the sentence and replaces the event details to generate an event description; the event details include participants, time and place;
[0008] Step 3: After each event is divided into sentences and segments, it is input into the large language model to generate sentence representations of event pairs; the large language model aggregates the global representations of the event at the paragraph and sentence levels;
[0009] Step 4: Based on RAG technology, retrieve semantic paragraphs related to the current event; RAG extracts contextual information related to the event pair from the corpus to provide a basis for reasoning;
[0010] Step 5: Decompose the event reasoning task into multiple steps, each of which corresponds to a specific reasoning task in event association analysis; generate possible reasoning paths through step-by-step analysis, and marginalize the reasoning paths; through marginalized reasoning, evaluate multiple reasoning paths simultaneously and obtain the most reasonable reasoning results; specific reasoning tasks include time sequence deduction and causal relationship judgment;
[0011] Step 6: Using the feature vector generated by the large language model as input, the event pairs are preliminarily classified through the Softmax classifier; the Softmax classifier outputs the probability distribution of the association relationship between the event pairs; the association relationship includes causal relationship, sequential relationship, or superior-female relationship;
[0012] Step 7: Take the most reasonable inference result output from step 5 as another input, and generate the final classification result by weighted fusion of the most reasonable inference result and the classification result of the Softmax classifier;
[0013] Step 8: Determine the specific relationship type between event pairs based on the classification results, and output the determination results of the causal, sequential or hierarchical relationship between events.
[0014] Furthermore, before step 1, the method further includes:
[0015] Assume that the event dataset contains multiple event pairs, each event pair consists of two events, which are used to describe the association relationship between the two; the event dataset is recorded as D={d_1, d_2, …, d_n}, where each event pair d_i represents two related events (e_i1, e_i2); each event text is processed by word segmentation and its related information is annotated, including context information, event trigger words and time sequence.
[0016] Furthermore, the step 2 comprises:
[0017] The large language model generates event pairs that conform to the original logical chain and replaces similar event contexts to increase the richness of the data:
[0018] (1)
[0019] in, Indicates that after rewriting, the event and events The event pair, and Represents two events, Indicated by event and events The event pair, Represents an event rewriting method based on a large language model.
[0020] Furthermore, in step 3, each event is processed into sentences and segments and then input into a large language model to generate a sentence representation of the event pair, including:
[0021] The large language model generates an embedded representation of each token in the context through its self-attention mechanism, and captures the complex semantics and dependencies in the event text by integrating context information. Based on the large language model, the input text is first divided into sentences and segments to generate sentence representations of event pairs:
[0022] (2)
[0023] Among them, X is the sentence representation of the event pair, and x_m is the mth paragraph formed after the input text is segmented and segmented.
[0024] Furthermore, in step 3, the large language model aggregates the global representations of the event at the paragraph and sentence levels, including:
[0025] The global semantic features of paragraphs and sentences are obtained by reading the [CLS] tag output or using the average pooling method; each sentence is encoded into a high-dimensional vector through a large language model, and the large language model generates the final features through symbol embedding, fragment embedding, and position embedding.
[0026] Furthermore, the embedding representation of the large language model is defined as:
[0027] (3)
[0028] in, represents the embedding representation of a large language model, Indicates a sentence or paragraph. Indicates that Model for sentences or paragraphs The word embedding vector of
[0029] The large language model uses a multi-layer self-attention mechanism to generate context-dependent features in each sentence;
[0030] The global representation of each sentence is calculated by [CLS] tagging or average pooling:
[0031] (4)
[0032] in, represents the global representation of each sentence, A word embedding vector representing a sentence or paragraph, Indicates the number of paragraphs formed after the input text is segmented into sentences and paragraphs;
[0033] The feature fusion formula is:
[0034] (5)
[0035] in, represents feature fusion, Represents a sentence or paragraph from the level of RoBERTa The word embedding vector that extracts features from Representation level The fusion weight coefficient of the word embedding vector from which the features are extracted.
[0036] Furthermore, after step 5 and before step 6, the method further includes:
[0037] A reflection mechanism is introduced after each reasoning step; the reflection mechanism verifies and corrects the current reasoning path and adjusts itself based on new evidence to ensure the accuracy of the final reasoning path.
[0038] Further, the step 7 comprises:
[0039] Take a weighted average of the classification results of step 5 and the Softmax classifier to generate the classification result:
[0040] (6)
[0041] in, Indicates an event and events The classification results are obtained by weighted average. Indicates an event and events Through The classification results obtained are Indicates an event and events The classification results obtained by CoT, is the weighting coefficient.
[0042] Furthermore, in step 8, the determination result of the causal, sequential or hierarchical relationship between events is expressed as:
[0043] (7)
[0044] in, It is the result of determining the causal, sequential or hierarchical relationship between events. It means to find the maximum value, Indicates an event and events The classification results are obtained by weighted average.
[0045] The present invention also discloses an event association analysis system based on a large language model, which is applicable to any of the above-mentioned event association analysis methods based on a large language model, and comprises:
[0046] The analysis and description module is used to perform semantic analysis on the context of each event pair in the event dataset through a large language model; the large language model explains and describes the potential relationship between each pair of events based on event triggers, context, and the time or logical sequence between events;
[0047] The generation module is used to generate event descriptions by keeping the original event relationship unchanged and rewriting the sentences and replacing the event details with a large language model; the event details include participants, time and place;
[0048] The sentence representation and aggregation module is used to process each event into sentences and segments and then input them into the large language model to generate sentence representations of event pairs; the large language model aggregates the global representations of the event at the paragraph and sentence levels;
[0049] The extraction module is used to retrieve semantic paragraphs related to the current event based on RAG technology; RAG extracts contextual information related to event pairs from the corpus to provide a basis for reasoning;
[0050] The reasoning module is used to decompose the event reasoning task into multiple steps, each of which corresponds to a specific reasoning task in event association analysis; through step-by-step analysis, possible reasoning paths are generated and marginalized; through marginalized reasoning, multiple reasoning paths are evaluated simultaneously and the most reasonable reasoning result is obtained; specific reasoning tasks include time sequence deduction and causal relationship judgment;
[0051] The probability distribution module is used to use the feature vector generated by the large language model as input to perform preliminary classification of event pairs through the Softmax classifier; the Softmax classifier will output the probability distribution of the association relationship between event pairs; the association relationship includes causal relationship, sequential relationship or superior-female relationship;
[0052] The classification module is used to take the most reasonable inference result as another input, and generate the final classification result by weighted fusion of the most reasonable inference result and the classification result of the Softmax classifier;
[0053] The determination module is used to determine the specific relationship type between event pairs according to the classification results, and output the determination results of the causal, sequential or hierarchical relationship between the events.
[0054] Due to the adoption of the above technical solution, the present invention has the following advantages:
[0055] 1. The present invention can provide higher accuracy and robustness in complex event relationship reasoning and classification tasks, making full use of the generation capabilities of large language models and the reasoning and analysis advantages of traditional models, overcoming the problems of incomplete feature extraction and unclear reasoning chain in existing methods.
[0056] 2. By combining a large language model (such as GLM-130B) with the RoBERTa-BASE model, the present invention proposes a dual feature extraction and data enhancement mechanism, which not only improves the diversity of event data, but also ensures the effective capture of complex semantic relationships, greatly improving the generalization ability of the model.
[0057] 3. The present invention introduces the chain of thought (CoT) and marginal reasoning methods in the reasoning module. By gradually disassembling complex event reasoning tasks and combining marginalized reasoning and reflection mechanisms, a multi-path reasoning chain is generated, thereby more accurately identifying the causal, sequential and hierarchical relationships between events.
[0058] 4. The present invention emphasizes the multi-model fusion of reasoning results. By weighted fusion of RoBERTa feature extraction results and thought chain reasoning results, it ensures that the classifier can maintain stable classification performance in different types of event relations, thus solving the problem of decreased classification accuracy of existing models in complex event association reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0060] Figure 1 The diagram is a schematic diagram of an event association analysis system based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The present invention is further described in conjunction with the accompanying drawings and embodiments, and the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present invention.
[0062] The present invention provides an embodiment of an event association analysis method based on a large language model, which includes:
[0063] Phase 0: Data initialization and preprocessing
[0064] S0: The method of the present invention first relates to the initialization and preprocessing of the original event data. Assume that the event data set contains multiple event pairs, each event pair is composed of two events, which are used to describe the association relationship between the two. The data set is recorded as D={d_1, d_2, …, d_n}, where each event pair d_i represents two associated events (e_i1, e_i2). At this stage, the present invention uses traditional preprocessing methods to process the data, including operations such as removing stop words, noise removal, event pair annotation and sentence segmentation. Each event text is processed by Tokenization segmentation, and its context information, event trigger words, time sequence, etc. are annotated to ensure that the subsequent stages can effectively use this information for event relationship analysis.
[0065] Stage 1: Data Augmentation
[0066] S1: After data preprocessing is completed, the data enhancement phase begins. The present invention expands and enriches the processed data set by combining a large language model (such as GLM-130B) with traditional data enhancement techniques. First, the system performs semantic analysis on the context of each event pair through a large language model. The model interprets and describes the potential relationship between each pair of events based on event triggers, context, and the temporal or logical order between events. Specifically, the system does not infer new relationship types, but generates explanatory analysis through existing semantic information to help enhance the logical coherence between events.
[0067] S2: Next, the system uses event detail reconstruction to enhance data. By keeping the original event relationship unchanged, the model rewrites sentences and replaces event details (such as participants, time, and location) to generate diverse event descriptions. The model generates event pairs that conform to the original logical chain and replaces similar event contexts to further increase the richness of the data.
[0068] (1)
[0069] in, Indicates that after rewriting, the event and events The event pair, and Represents two events, Indicated by event and events The event pair, Represents an event rewriting method based on a large language model.
[0070] Stage 2: Event text structure feature extraction
[0071] S3: This stage extracts the structural features of the event text based on the RoBERTa-BASE model. Before entering the RoBERTa-BASE model, the event text will be processed into sentences and segments to ensure that each paragraph or sentence is extracted as an independent input sequence. RoBERTa-BASE generates an embedded representation of each token in the context through its self-attention mechanism, and captures the complex semantics and dependencies in the event text by integrating contextual information. Based on the RoBERTa-BASE model, the system first sentences and segments the input text to generate sentence representations of event pairs:
[0072] (2)
[0073] Among them, X is the sentence representation of the event pair, and x_m is the mth paragraph formed after the input text is segmented and segmented.
[0074] S4: During the text structure feature extraction process, the RoBERTa-BASE model also aggregates global representations of events at the paragraph and sentence levels. By reading the [CLS] tag output or using the average pooling method, the system can obtain global semantic features of paragraphs and sentences. These high-dimensional vector representations provide accurate input for subsequent classification tasks. Each sentence is encoded into a high-dimensional vector by the RoBERTa-BASE model, and the model generates the final features through the following three embedding representations: Token Embedding, Position Embedding, and Segment Embedding. The embedding representation of the model can be defined as:
[0075] (3)
[0076] in, represents the embedding representation of a large language model, Indicates a sentence or paragraph. Indicates that Model for sentences or paragraphs The word embedding vector of
[0077] RoBERTa-BASE uses a multi-layer self-attention mechanism to generate context-dependent features in each sentence, especially long-distance dependencies. The global representation of each sentence is calculated by [CLS] tagging or average pooling:
[0078] (4)
[0079] in, represents the global representation of each sentence, A word embedding vector representing a sentence or paragraph, Indicates the number of paragraphs formed after the input text is segmented into sentences and paragraphs;
[0080] In order to capture multi-level features more comprehensively, the feature fusion formula is:
[0081] (5)
[0082] in, represents feature fusion, Represents a sentence or paragraph from the level of RoBERTa The word embedding vector that extracts features from Representation level The fusion weight coefficient of the word embedding vector from which the features are extracted.
[0083] Stage 3: Reasoning and decision support based on thought chain
[0084] S5: This paper uses an edge reasoning method based on the Chain of Thought (CoT) prompting technology to decompose the complex reasoning tasks between events. At this stage, the system first retrieves semantic paragraphs related to the current event based on the RAG (Retrieval-Augmented Generation) technology. RAG extracts contextual information related to the event pair from the corpus to provide a basis for reasoning.
[0085] S6: Next, the system breaks down the event reasoning task into multiple Ss, each of which corresponds to a specific reasoning task in event association analysis (such as time sequence derivation, causal relationship judgment, etc.). The system generates possible reasoning paths through step-by-step analysis and marginalizes the reasoning paths. Through marginalized reasoning, the system can evaluate multiple reasoning paths at the same time and obtain the most reasonable reasoning results.
[0086] S7: In order to improve the accuracy of reasoning, the system introduces a reflection mechanism after each reasoning S. The reflection mechanism verifies and corrects the current reasoning path, and makes self-adjustments based on new evidence to ensure the accuracy of the final reasoning path.
[0087] Stage 4: Event relationship classification and result fusion
[0088] S8: After feature extraction and inference, the system will enter the classification stage. First, the system uses the feature vector generated by RoBERTa-BASE as input and performs preliminary classification of event pairs through the Softmax classifier. The Softmax classifier will output the probability distribution of the association relationship between event pairs (such as causal relationship, sequential relationship, or superior-female relationship).
[0089] S9: At the same time, the system will use the edge inference result of S6 as another input, and weightedly fuse the classification results of S6 and Softmax classifier. By weighted averaging the two classification results, the system can generate a more robust classification output.
[0090] (6)
[0091] in, Indicates an event and events The classification results are obtained by weighted average. Indicates an event and events Through The classification results obtained are Indicates an event and events The classification results obtained by CoT, is the weighting coefficient.
[0092] S10: Finally, the system determines the specific relationship type between the event pairs according to the classification results, and outputs the determination results of the causal, sequential or hierarchical relationship between the events. Through the comprehensive classification at this stage, the present invention can accurately identify the complex associations between event pairs and ensure the accuracy and consistency of the classification results.
[0093] (7)
[0094] in, It is the result of determining the causal, sequential or hierarchical relationship between events. It means to find the maximum value, Indicates an event and events The classification results are obtained by weighted average.
[0095] This application is applicable to event relationship extraction and association analysis in natural language processing tasks, especially to the classification and inference of causal, sequential and hierarchical relationships between complex events. Application scenarios include news analysis, legal document processing, and multi-event semantic reasoning in intelligent decision-making systems.
[0096] The present invention also provides an embodiment of an event association analysis system based on a large language model, which aims to more accurately capture the complex relationship between events and is applicable to the event association analysis method based on a large language model described in the above embodiment, and includes:
[0097] The analysis and description module is used to perform semantic analysis on the context of each event pair in the event dataset through a large language model; the large language model explains and describes the potential relationship between each pair of events based on event triggers, context, and the time or logical sequence between events;
[0098] The generation module is used to generate event descriptions by keeping the original event relationship unchanged and rewriting the sentences and replacing the event details with a large language model; the event details include participants, time and place;
[0099] The sentence representation and aggregation module is used to process each event into sentences and segments and then input them into the large language model to generate sentence representations of event pairs; the large language model aggregates the global representations of the event at the paragraph and sentence levels;
[0100] The extraction module is used to retrieve semantic paragraphs related to the current event based on RAG technology; RAG extracts contextual information related to event pairs from the corpus to provide a basis for reasoning;
[0101] The reasoning module is used to decompose the event reasoning task into multiple steps, each of which corresponds to a specific reasoning task in event association analysis; through step-by-step analysis, possible reasoning paths are generated and marginalized; through marginalized reasoning, multiple reasoning paths are evaluated simultaneously and the most reasonable reasoning result is obtained; specific reasoning tasks include time sequence deduction and causal relationship judgment;
[0102] The probability distribution module is used to use the feature vector generated by the large language model as input to perform preliminary classification of event pairs through the Softmax classifier; the Softmax classifier will output the probability distribution of the association relationship between event pairs; the association relationship includes causal relationship, sequential relationship or superior-female relationship;
[0103] The classification module is used to take the most reasonable inference result as another input, and generate the final classification result by weighted fusion of the most reasonable inference result and the classification result of the Softmax classifier;
[0104] The determination module is used to determine the specific relationship type between event pairs according to the classification results, and output the determination results of the causal, sequential or hierarchical relationship between the events.
[0105] For ease of understanding, the present invention also provides a more specific embodiment: the present invention first uses a large language model to perform diversified data enhancement on event data to generate a richer event pair corpus; secondly, a text feature extraction module based on the RoBERTa-BASE model is designed to capture paragraphs, sentence structures and long-distance dependencies in event texts, and generate high-dimensional semantic vectors; then, the system uses the Chain of Thought (CoT) reasoning module to decompose complex event reasoning tasks into multiple steps, and combines edge reasoning methods to perform multi-path reasoning and reflective verification on the logical chains between events to ensure the accuracy of the reasoning results; finally, a comprehensive use of the Softmax classifier and the weighted fusion of reasoning results is used to finally classify the correlation between events and identify causal, sequential, and hierarchical relationships. Figure 1 As shown, this embodiment specifically includes:
[0106] Module 1. Event data processing and enhancement module based on a combination of large language models and traditional technologies: This module combines traditional data preprocessing techniques with the generation and understanding capabilities of large language models (such as GLM-130B) to provide a diverse and high-quality dataset for event relationship extraction tasks. First, the system uses traditional preprocessing techniques to process the raw event data, including stop word removal, denoising, event pair annotation, and sentence segmentation. These steps can clean the data and make it more suitable for subsequent model processing and enhancement. In the data enhancement process, first, the system uses a large language model to analyze the semantic relationship between event pairs, focusing on the context and structural features between event pairs. The model further explains the potential relationship between two events by analyzing the context of the event, event trigger words, and the order in which the events occur. For example, the system will analyze why one event may lead to another event based on contextual information, or why two events are in a temporal or logical order. In this process, the model does not directly infer new relationship types, but provides richer explanations and context through existing semantic clues so that the enhanced data can better reflect the logical chain of events. Through this explanatory analysis, the system is able to generate diverse event descriptions. For example, while maintaining the original event relationship, the model will expand the original data by rewriting sentences, replacing event details, etc. This enhancement method not only includes a rich description of event details, such as changing the participants or scenes of the event, but also includes background explanations of logical relationships such as event causality and sequence. These extended data maintain logical consistency with the original data while providing the model with more diverse training samples. Next, the system generates new sample data through data enhancement technology. Specific methods include rewriting sentences in the original event pair, keeping the event relationship unchanged, but generating new descriptions by replacing sentences, introducing new background or details, etc. This not only enhances the richness of the dataset, but also ensures that the enhanced data is logically and semantically consistent. For example, after analyzing that two events have a causal relationship, the model will generate diversified event pairs consistent with this causal relationship, or replace them with event pairs with similar causal relationships to enrich the dataset. In addition, the model will also use event context to replace or introduce event pairs of the same type, and increase the diversity of the data while keeping the core event relationship unchanged by modifying some content, such as the time, place or participants of the event. Through this series of data enhancement strategies, the model can generate data samples that meet the task requirements and have diverse semantic and structural features. Ultimately, these enhanced data will be integrated with traditional preprocessed data to form a more comprehensive training set, thereby improving the generalization ability and accuracy of the event relationship extraction model.
[0107] Module 2. Event text structural feature extraction module based on RoBERTa-BASE: This module extracts the structural features of event texts, including paragraph and chapter structures, through the RoBERTa-BASE pre-trained model, aiming to provide accurate semantic and structural information for subsequent event relationship extraction. As a powerful pre-trained language model, RoBERTa-BASE can deeply capture the contextual dependencies in event texts and extract complex features in texts through its multi-layer self-attention mechanism. First, the event text will go through a series of preprocessing steps before entering the RoBERTa model. The sentence and segmentation of the text is the basic preprocessing process. The system will split the event text into sentences or paragraphs, and each paragraph is regarded as an independent input unit. Next, the system uses RoBERTa's vocabulary to tokenize the text. This step converts the vocabulary into corresponding tokens through Byte-Pair Encoding (BPE) to ensure that the input text is represented in a form that RoBERTa can process. In addition, paragraph and sentence boundaries in the text are clearly marked with special tags (such as [CLS] and [SEP]) so that the model can accurately distinguish their structures when processing multiple paragraphs or sentences. After preprocessing, the text is fed into the RoBERTa-BASE model. The model encodes the vocabulary, position and sentence structure information of the text by combining three embedding representations - TokenEmbedding, Position Embedding and Segment Embedding - to generate preliminary representations of the event text E1, E2, ..., EN. These embeddings not only provide the model with the position and order information of each word in its context, but also help RoBERTa-BASE capture semantic logical relationships when processing paragraphs and chapters. RoBERTa's feature extraction capability relies on its multi-layer self-attention mechanism Trm. The model generates context-related embedding representations for each token in the input text to capture its semantic and structural features. In this process, the model not only extracts local information of a single sentence, but also crosses sentences and paragraphs to identify long-range dependencies in the text. In particular, through the self-attention mechanism, RoBERTa is able to focus on words at different positions in the text and calculate the dependency weights between these words, thereby accurately capturing the complex relationships between events. When extracting paragraph and chapter structure features, RoBERTa-BASE usually uses the [CLS] tag as the overall representation of the sentence or paragraph. The [CLS] representation of each sentence contains the global semantic information of the sentence. Therefore, by reading the output of [CLS], the system can obtain the semantic features of the paragraph, chapter or sentence.If you are dealing with long text paragraphs, you can also aggregate the representations of all tokens in each paragraph through mean pooling to obtain the overall features of the paragraph. These global representations provide powerful structured feature inputs for subsequent event analysis and classification. In addition, RoBERTa is a deep model whose multiple layers can extract semantic features of different depths. In order to more comprehensively represent the structural features of paragraphs or chapters, this module also performs weighted fusion of outputs from different layers. By extracting features from different layers of RoBERTa (such as the output of the 6th or 12th layer) and performing weighted averaging on these features, the system can combine feature representations of multiple depths to generate a more detailed and comprehensive text structure representation. It is worth emphasizing that RoBERTa's multi-layer self-attention mechanism allows the model to capture long-distance dependencies between events in long texts, especially when events span sentences or paragraphs. The model automatically establishes cross-sentence and cross-paragraph associations through its context-awareness, thereby providing a more complete background and logical chain for event association analysis. This capture of long-distance dependencies is crucial for event relationship extraction, because many associations between events are not limited to a single sentence or paragraph, but span multiple contexts. Ultimately, the features output by RoBERTa-BASE include paragraph-level and sentence-level semantic representations T1, T2, …TN, combining multiple levels of contextual information. Through these high-dimensional feature representations, the system can provide accurate input for subsequent event relationship classification. These features not only reflect the semantic structure of the event text, but also include its logical relationship in the overall context, ensuring that the classification model can make inferences based on a full understanding of the text.
[0108] Module 3. Event association decision support module based on thought chain and marginal reasoning: This module provides decision support for event association analysis by combining the chain-of-thought (CoT) prompt technology and marginal reasoning method. It aims to handle complex event association reasoning tasks and improve the accuracy and reliability of analysis by gradually disassembling and marginalizing reasoning. The key function of this module is to generate a reasonable reasoning chain, verify and optimize each step in the reasoning process, and finally obtain the best reasoning result. First, the module uses RAG (Retrieval-Augmented Generation) technology to retrieve relevant information paragraphs in the process of processing event text. Specifically, when an event pair is input, the system retrieves paragraphs or sentences related to the current event from a large number of document libraries or corpora through RAG as reference information for event reasoning. These retrieved paragraphs provide event background, event context or possible association information, providing a basis for subsequent reasoning. The retrieved paragraphs not only contain direct descriptions of the event, but may also include relevant background information or potential causal clues to help the model establish a preliminary reasoning path. Next, the system decomposes the complex event association reasoning task into multiple actionable steps based on the Chain of Thought (CoT) prompting technology. The core of the Chain of Thought technology lies in the process of step-by-step reasoning. The model does not generate a complete reasoning result at one time, but decomposes the complex reasoning task into multiple stages by gradually disassembling the problem, and each stage handles a subtask. Each subtask corresponds to a specific reasoning step in event association analysis, such as determining the time sequence of events, inferring causal relationships, or analyzing semantic dependencies between events. At each stage of reasoning, the system performs marginalized reasoning, that is, evaluating possible reasoning paths based on existing evidence and contextual information. Marginalized reasoning allows the system to infer on multiple reasoning paths at the same time, and obtain the most likely reasoning result by analyzing the rationality of each path. This method helps to deal with complex event association problems, especially when the relationship between events is unclear or there are multiple possible explanations. Marginalized reasoning can help the model deal with uncertainty more flexibly and avoid premature conclusions. In order to further improve the accuracy and reliability of reasoning, the module introduces a reflection mechanism. After each reasoning step is completed, the system reviews and verifies the current reasoning path to check whether there are potential errors or omissions in the reasoning process. The reflection mechanism allows the model to revisit previous reasoning results at each stage and make adjustments and corrections based on new evidence or background information. This mechanism can effectively reduce the accumulated errors in the reasoning process and ensure that each reasoning step can maintain a high degree of accuracy. The reflection process not only verifies the reasoning results of each subtask, but also provides the system with the ability to self-correct, enabling it to dynamically adjust based on previous reasoning results.After all reasoning paths and reflection processes are completed, the system will synthesize the reasoning results of different paths and select the optimal reasoning path as the final decision basis. In order to obtain the optimal path, the system uses marginalization to reflect on the results of multiple reasoning paths, analyze the confidence and rationality of each path, and finally select the most confident and reasonable reasoning path as the final judgment basis for the association between events. For example, when dealing with the causal relationship between two events, the system first retrieves relevant paragraphs through RAG to identify possible clues related to the event background; then uses the thinking chain prompt to decompose the causal relationship reasoning into multiple steps, such as "whether event A occurs before event B", "whether the occurrence of event A triggers event B", etc. After each reasoning step, the reflection mechanism will review the current conclusion and verify its rationality in combination with new evidence. Finally, through the marginal reasoning and reflection of multiple paths, the system selects the most credible causal chain and derives the causal relationship between events. In summary, this module can provide a flexible and accurate event association decision-making process through the combination of RAG retrieval, thinking chain decomposition, marginalized reasoning and reflection verification. It can not only handle complex reasoning tasks, but also has the ability to self-correct, ensuring the accuracy and reliability of the reasoning results. The output of this module provides strong support for subsequent event classification and association judgment.
[0109] Module 4. Classification module: The main function of this module is to comprehensively classify the extracted event features and identify the specific relationships between events (such as causal relationships, sequential relationships, and hierarchical relationships). This module will combine the structured semantic features extracted from RoBERTa-BASE and the inference results generated by edge reasoning in module three, and use the Softmax classifier to perform the final event relationship classification. First, after RoBERTa-BASE processing, the system has extracted the contextual features, paragraph structure, trigger words and other semantic information of each pair of events. These features represent the semantic content and mutual dependencies of each event in the form of high-dimensional vectors. These vectors provide rich basic information for subsequent classification tasks and can effectively capture the subtle differences in event texts. After the feature extraction is completed, the system will input these semantic features extracted from RoBERTa-BASE into the Softmax classifier for preliminary classification. The Softmax classifier is a multi-classifier that can calculate the probability distribution of the association relationship between events based on the input feature vector. Specifically, the Softmax classifier will classify the event relationship and output the probability of each relationship (such as causal, sequential, hierarchical). The classification result is a probability distribution, in which each relationship type corresponds to a probability value, and the system will select the most likely relationship type as the preliminary classification result based on the maximum probability. At the same time, the edge reasoning result of module three provides another important decision basis for the classification task. Module three obtains the event relationship classification result based on complex reasoning through the process of chain of thought (CoT) reasoning and marginalized reflection. Since edge reasoning has high robustness in analyzing event logic and dependencies across sentences or paragraphs, its output results can supplement the potential deficiencies in RoBERTa-BASE feature extraction to a certain extent. Therefore, in order to improve the accuracy of the overall classification, the system will weight the classification results generated in module three with the results of RoBERTa-BASE and Softmax classifiers. Specifically, the system will first assign weights to the preliminary classification results of RoBERTa-BASE + Softmax and the edge reasoning classification results of module three. The basis for this weighting strategy can be based on prior experience or the model performance obtained during experimental verification (for example, if the reasoning of module three performs better than the RoBERTa model in a certain type of event relationship, the system will give it a higher weight). By weighting the results of the two different models, the system finally generates a comprehensive classification result. This weighted classification result is more robust. It combines the feature extraction advantages of RoBERTa and the deep logical deduction of edge reasoning to ensure the accuracy and reliability of classification. Finally, the system will classify the event pairs and output the relationship type between the event pairs, whether it is a causal relationship, a sequential relationship, or a hierarchical relationship.In this way, this module achieves complementary advantages of models in the event relationship classification task, ensures that the relationship of each event pair can be accurately classified, and provides reliable classification decision support for the entire event association analysis.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An event association analysis method based on a large language model, characterized in that: include: Step 1: Use a large language model to perform semantic analysis on the context of each event pair in the event dataset; The big language model explains and describes the potential relationship between each pair of events based on event triggers, context, and the temporal or logical order between events; Step 2: By keeping the original event relationship unchanged, the large language model rewrites the sentence and replaces the event details to generate an event description; the event details include participants, time and place; Step 3: After each event is divided into sentences and segments, it is input into the large language model to generate sentence representations of event pairs; The large language model aggregates global representations of events at paragraph and sentence levels; Step 4: Based on RAG technology, retrieve semantic paragraphs related to the current event; RAG extracts contextual information related to the event pair from the corpus to provide a basis for reasoning; Step 5: Decompose the event reasoning task into multiple steps, each of which corresponds to a specific reasoning task in event association analysis; generate possible reasoning paths through step-by-step analysis, and marginalize the reasoning paths; through marginalized reasoning, evaluate multiple reasoning paths simultaneously and obtain the most reasonable reasoning results; specific reasoning tasks include time sequence deduction and causal relationship judgment; Step 6: Using the feature vector generated by the large language model as input, the event pairs are preliminarily classified through the Softmax classifier; the Softmax classifier outputs the probability distribution of the association relationship between the event pairs; the association relationship includes causal relationship, sequential relationship, or superior-female relationship; Step 7: Take the most reasonable inference result output from step 5 as another input, and generate the final classification result by weighted fusion of the most reasonable inference result and the classification result of the Softmax classifier; Step 8: Determine the specific relationship type between event pairs based on the classification results, and output the determination results of the causal, sequential or hierarchical relationship between events.
2. The event association analysis method based on a large language model according to claim 1 is characterized in that: Before step 1, the method further includes: Assume that the event dataset contains multiple event pairs, each event pair consists of two events, which are used to describe the association relationship between the two; the event dataset is recorded as D={d_1, d_2, …, d_n}, where each event pair d_i represents two related events (e_i1, e_i2); each event text is processed by word segmentation and its related information is annotated, including context information, event trigger words and time sequence.
3. The event association analysis method based on a large language model according to claim 1 is characterized in that: The step 2 comprises: The large language model generates event pairs that conform to the original logical chain and replaces similar event contexts to increase the richness of the data: (1) in, Indicates that after rewriting, the event and events The event pair, and Represents two events, Indicated by event and events The event pair, Represents an event rewriting method based on a large language model.
4. The event association analysis method based on a large language model according to claim 1 is characterized in that: In step 3, each event is processed into sentences and segments and then input into a large language model to generate a sentence representation of the event pair, including: The large language model generates an embedded representation of each token in the context through its self-attention mechanism, and captures the complex semantics and dependencies in the event text by integrating context information. Based on the large language model, the input text is first divided into sentences and segments to generate sentence representations of event pairs: (2) Among them, X is the sentence representation of the event pair, and x_m is the mth paragraph formed after the input text is segmented and segmented.
5. The event association analysis method based on a large language model according to claim 1 is characterized in that: In step 3, the large language model aggregates the global representation of the event at the paragraph and sentence levels, including: The global semantic features of paragraphs and sentences are obtained by reading the [CLS] tag output or using the average pooling method; each sentence is encoded into a high-dimensional vector through a large language model, and the large language model generates the final features through symbol embedding, fragment embedding, and position embedding.
6. The event association analysis method based on a large language model according to claim 5 is characterized in that: The embedding representation of the large language model is defined as: (3) in, represents the embedding representation of a large language model, Indicates a sentence or paragraph. Indicates that Model for sentences or paragraphs The word embedding vector of The large language model uses a multi-layer self-attention mechanism to generate context-dependent features in each sentence; The global representation of each sentence is calculated by [CLS] tagging or average pooling: (4) in, represents the global representation of each sentence, A word embedding vector representing a sentence or paragraph, Indicates the number of paragraphs formed after the input text is segmented into sentences and paragraphs; The feature fusion formula is: (5) in, represents feature fusion, Represents a sentence or paragraph from the level of RoBERTa The word embedding vector that extracts features from Representation level The fusion weight coefficient of the word embedding vector from which the features are extracted.
7. The event association analysis method based on a large language model according to claim 1 is characterized in that: After step 5 and before step 6, the method further includes: A reflection mechanism is introduced after each reasoning step; the reflection mechanism verifies and corrects the current reasoning path and adjusts itself based on new evidence to ensure the accuracy of the final reasoning path.
8. The event association analysis method based on a large language model according to claim 1, characterized in that: The step 7 comprises: Take a weighted average of the classification results of step 5 and the Softmax classifier to generate the classification result: (6) in, Indicates an event and events The classification results are obtained by weighted average. Indicates an event and events Through The classification results obtained are Indicates an event and events The classification results obtained by CoT are is the weighting coefficient.
9. The event association analysis method based on a large language model according to claim 1, characterized in that: In step 8, the determination result of the causal, sequential or hierarchical relationship between events is expressed as: (7) in, It is the result of determining the causal, sequential or hierarchical relationship between events. It means to find the maximum value, Indicates an event and events The classification results are obtained by weighted average.
10. An event association analysis system based on a large language model, applicable to the event association analysis method based on a large language model according to any one of claims 1 to 9, characterized in that: include: An analysis and description module, which is used to perform semantic analysis on the context of each event pair in the event dataset through a large language model; The big language model explains and describes the potential relationship between each pair of events based on event triggers, context, and the temporal or logical order between events; The generation module is used to generate event descriptions by keeping the original event relationship unchanged and rewriting the sentences and replacing the event details with a large language model; the event details include participants, time and place; The sentence representation and aggregation module is used to process each event into sentences and segments and then input them into the large language model to generate sentence representations of event pairs; The large language model aggregates global representations of events at paragraph and sentence levels; The extraction module is used to retrieve semantic paragraphs related to the current event based on RAG technology; RAG extracts contextual information related to event pairs from the corpus to provide a basis for reasoning; The reasoning module is used to decompose the event reasoning task into multiple steps, each of which corresponds to a specific reasoning task in event association analysis; through step-by-step analysis, possible reasoning paths are generated and marginalized; through marginalized reasoning, multiple reasoning paths are evaluated simultaneously and the most reasonable reasoning result is obtained; specific reasoning tasks include time sequence deduction and causal relationship judgment; The probability distribution module is used to use the feature vector generated by the large language model as input to perform preliminary classification of event pairs through the Softmax classifier; the Softmax classifier will output the probability distribution of the association relationship between event pairs; the association relationship includes causal relationship, sequential relationship or superior-female relationship; The classification module is used to take the most reasonable inference result as another input, and generate the final classification result by weighted fusion of the most reasonable inference result and the classification result of the Softmax classifier; The determination module is used to determine the specific relationship type between event pairs according to the classification results, and output the determination results of the causal, sequential or hierarchical relationship between the events.
Citation Information
Patent Citations
Vertical field financial large model system for realizing function of efficiently processing table data and method of vertical field financial large model system
CN118194988A
Event extraction-oriented big language model data enhancement method and device
CN118551194A