Text event association method based on large language model
Through the text event association method based on the large language model, we deeply explore the implicit association between events in the field of social governance, form a complete event association chain, solve the problem that existing technology is difficult to identify implicit associations and limited early warning capabilities, and achieve accurate and comprehensive correlation analysis and early warning of risk events.
Patent Information
- Application Number
- CN202510267231.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
When the existing technology deals with complex social governance scenarios, it is difficult to identify the hidden relationships between event elements, and it is impossible to effectively explore the causal chain and development context of events. It has limited early warning capabilities, making it difficult to meet the needs of real-time early warning and in-depth analysis of risk events in social governance.
The text event association method based on the large language model is adopted, and the multi-source data collection system is built, and the original data is cleaned, converted and structured by the large language model, core elements such as characters, places, and events are extracted, social relations networks are built, and the implicit associations between events are deeply explored through the reasoning ability of the large language model to form a complete event association chain.
It significantly improves the accuracy of identification of associations between events, fully reflects the deep understanding and accurate identification of associations between different events, effectively improves the depth and accuracy of associations between events, and meets the needs of real-time early warning and in-depth analysis of risk events in social governance.
Smart Images

Figure CN120196741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence applications, and in particular, to a method for associating text events based on a large language model. Background Art
[0002] In recent years, with the continuous increase in the demand for social governance, the field of social governance has faced major challenges in processing a large amount of unstructured data and analyzing the association of complex risk events. These data come from a wide range of sources, including news reports, social media, surveillance videos, sensor data, etc. They are diverse in form and redundant in information, and often lack a unified structured framework, making it extremely difficult to extract and analyze information. To address these problems, existing technical solutions have attempted to solve them through various methods, mainly including traditional methods based on rule matching, feature analysis techniques relying on machine learning, and relationship mining methods based on knowledge graphs. Among them, the method based on rule matching relies on pre-set rules or templates to associate events by directly matching event elements. However, this method cannot adapt to the scenarios of data diversity and dynamic changes, and is prone to limitations in the association effect. The feature analysis technique based on machine learning precisely identifies the connections between events by modeling and classifying data features. However, it relies on a large amount of labeled data and has weak ability to mine implicit associations. In contrast, the relationship mining method based on knowledge graphs provides a more intuitive solution idea. By constructing a network of entities and their relationships, explicit associations between events are discovered. However, these methods often only focus on the surface features of events, and it is difficult to deeply understand the internal connections and complex development laws between events, and they cannot comprehensively reveal the deep logic of events.
[0003] The current technical solution closest to the present invention is a hybrid method that combines a knowledge graph with statistical analysis. This solution discovers explicit associations between risk events by constructing a relationship network containing entities such as people, places, and events, and at the same time uses a statistical model to predict the development trend of events. However, this solution still has many deficiencies. For example, it is difficult to identify implicit associations between event elements, unable to effectively mine the causal chain and development context of events, and insufficient in the ability to analyze and deeply understand complex events. In addition, the early warning ability of this method is relatively limited and difficult to meet the needs of real-time early warning and in-depth analysis of risk events in social governance. Therefore, when dealing with complex social governance scenarios, existing solutions still need to further improve the ability to deeply mine the association and dynamic reasoning of risk events to achieve more accurate and comprehensive governance effects. Summary of the Invention
[0004] The objective of the present invention is to provide a text event correlation method based on large language models, which can be applied to warning and prediction scenarios, etc. First, a multi-source data collection system is constructed, including the unified aggregation of basic data sources such as government affairs data, public opinion data, complaint data, and monitoring data. Then, using the prompting method of large language models, the original data is cleaned, transformed, and structured to accurately extract core elements such as people, places, and events, forming a standardized data foundation. In the element recognition stage, the characteristics and behavior patterns of key figures and key groups are intelligently recognized to construct a social relationship network. In the correlation analysis link, through the reasoning ability of large language models, the implicit correlations between events are deeply mined, including multi-dimensional analysis such as cause correlation, process correlation, and result correlation, forming a complete event correlation chain.
[0005] The technical solution for achieving the objective of the present invention is as follows:
[0006] A text event correlation method based on large language models, including:
[0007] Step 1, using prompting words, use a large language model to extract event elements in the current text event e0 to obtain an event element set;
[0008] Step 2, use a noise word library and regular expressions to filter the event element set;
[0009] Step 3, use the event elements in the filtered event element set to retrieve the associated events of the current text event e0 to obtain an associated event set;
[0010] Step 4, using prompting words, use a large language model to extract the occurrence time of each event in the associated event set, and arrange all the associated events in the associated event set in ascending order of occurrence time to obtain an associated event set in ascending order {e0, e1, e2, e3,...};
[0011] Step 5, using prompting words, use a large language model to extract the relationship between every two events in {e0, e1, e2, e3,...} to obtain an event relationship set; the elements in the event relationship set are "e i , e i and e j 's relationship, e j "; where i, j = 0, 1, 2, 3,... are the numbers of events, and i < j;
[0012] Step 6, calculate the correlation degree of each event in {e0, e1, e2, e3,...} with the current text event e0, and filter out the events with a correlation degree less than or equal to the threshold from {e0, e1, e2, e3,...} to obtain an effective associated event set;
[0013] Step 7: Traverse each element in the event relation set. If both associated events included in the element are in the valid associated event set, retain the element; otherwise, delete the element. After traversal, obtain the valid event relation set.
[0014] Step 8: Integrate all elements in the valid event relation set to obtain the associated event chain of the current text event e0.
[0015] Furthermore, it further includes:
[0016] Step 5.1: Using prompt words, use a large language model to extract the relationships between the characters in all events in {e0, e1, e2, e3,...} to obtain a character relationship set. The elements in the character relationship set are "the relationship between v l , v l and v m ; where l, m = 1, 2, 3,... are the numbers of the characters. m "
[0017] Step 9: Using the character relationship set, use a large language model to optimize the associated event chain of the current text event e0 obtained in Step 8.
[0018] Preferably, the large language model is Tongyi Qianwen Language Model, Zhipu Qingyan, or DeepSeek.
[0019] The present invention adopts a two-stage association strategy of "element - event", accurately extracts key elements such as characters, time, and actions in text events, and through the powerful reasoning and association understanding ability of the large language model, deeply excavates the hidden association relationships between event elements. Compared with the prior art, the present invention can effectively improve the accuracy and depth of event association by combining interference word filtering, regular expression screening, and the sorting of associated events based on time sequence. Especially in the aspect of relationship extraction and association degree calculation between complex events, the present invention significantly improves the recognition accuracy of the association between events, fully reflects the in-depth understanding and accurate recognition of the association relationships between different events, and thus effectively improves the depth and accuracy of the event association relationship. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is the schematic diagram of the principle of the present invention.
[0021] Figure 2 is the schematic flow chart of obtaining the associated event set in the first stage in the embodiment.
[0022] Figure 3 is the character relationship diagram transformed from the character relationship set in the second stage in the embodiment.
[0023] Figure 4 is the event relationship diagram obtained by screening the event relationship set in the second stage in the embodiment.
[0024] Figure 5 It is a schematic diagram of the overall process of the embodiment. Specific implementation manners
[0025] The object of the present invention is to realize in-depth event association by using large language models. This method breaks through the limitations of traditional event association analysis methods, adopts a two-stage association strategy of "element - event", and constructs a systematic event association analysis framework. In the first stage, deeply analyze the risk event data in the field of social governance, use advanced natural language processing technology to accurately extract key elements such as people, places, times, and behaviors in the events, and establish a standardized event structured representation model, laying a solid data foundation for in-depth association analysis. In the second stage, make full use of the powerful knowledge reasoning and association understanding ability of large language models, and through prompt words, guide the model to deeply mine and analyze the hidden associations between event elements, so as to realize the in-depth understanding and accurate identification of complex association relationships between different events.
[0026] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Among them, the content of the embodiment is from publicly disclosed legal documents and does not involve privacy.
[0027] The system principle of the present invention is as Figure 1 shown.
[0028] The first stage, element association.
[0029] Step 1: Element extraction. Design prompt words to enable the large language model to extract event elements such as names of people and places from the description of event text e0, forming an element set N1 = {n1, n2,...}. In this embodiment, the large language model used is Tongyi Qianwen language model launched by Alibaba Cloud (other large language models such as Zhipu Qingyan or Deepseek can also be used). This model can infer and calculate the output results according to the prompt words. Usually, the relationship between the prompt words and the names of people and places is reflected in the semantic association in the context. In this embodiment, the prompt words are abstract language representations of names of people, places or other names. They can be explicit identifiers. For example, the prompt words for names of people include "name", "called", "Mr.", "Ms.", "professor", etc., while the prompt words for places include "located in", "in", "from", "city", "country", "region", etc.
[0030] Step 2: Construct a noise word library c = {c1, c2, c3, ...} that can be dynamically updated according to the text content, and filter the entity set. (The noise word library is generally determined based on domain knowledge. For example, in legal documents, the defendant, he, Chen and other unclear words are used as noise words.) Through text matching, match the noise words from N1 and remove them. Suppose the noise word n7 exists in the matched set N1, remove it from N1, and get the filtered entity set N2 = {n1, n2, n3, n6, n8 ...}.
[0031] Step 3: Use regular expressions (such as "^[\u4e00-\u9fa5]") (regular expressions are a tool for matching and manipulating texts, consisting of a series of characters and special characters that describe the pattern of the text to be matched. For example, runoo+b can match runoob, runooooob, etc., and the + sign means that the preceding character must appear at least once) and fuzzy matching expressions (such as 'someone$') to construct a filtering operator (such as '^[\u4e00-\u9fa5]+someone$'), filter the fuzzy entities in N2, and obtain N3 = {n1, n2, n3, n6, n8...}. Fuzzy entities refer to objects or concepts described in the text that are not clear and distinct, and may be difficult to accurately identify or classify due to lack of specific information, incomplete context or ambiguous expression, such as the fuzzy name "Zhang", etc. After filtering, an accurate entity set N3 is obtained. The conditions in the expression are not unique. This embodiment is based on a legal document dataset, which is usually specified as xxx, China Railway xx Bureau, xx Province xx City.
[0032] Step 4: Use the entity names in entity set N3 to search for related events in the existing database consisting of legal documents, and record the event set corresponding to the event as E1 = {e0, e1, e2, ... e n}.
[0033] like Figure 2As shown in the figure, in this embodiment, the current event e0 = {"On August 1, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang Mou, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management building project."}, and the element set N1 = {"Shenyang No. 14 Construction Engineering Company", "Li Hongjun", "Zhang Mou", "Ma Ji", "Zou Decun", "property"} obtained after step 1 processing. After filtering the set N1 through the interference word library in step 2, "property" was filtered out, and the entity set N2 = {"Shenyang No. 14 Construction Engineering Company", "Li Hongjun", "Zhang Mou", "Ma Ji", "Zou Decun"} was obtained. Subsequently, after filtering through the regular expression in step 3, "Zhang Mou" was filtered out, and the entity set N3 = {"Shenyang No. 14 Construction Engineering Company", "Li Hongjun", "Ma Ji", "Zou Decun"} was obtained. Using the entity set N3 to retrieve related events in the database, E1 = {"On August 1, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang Mou, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management building project.", "On August 7, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang Mou, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management building project. The two parties agreed that the above-ground labor cost would be calculated at 210 yuan per square meter above ground and 260 yuan per square meter underground. However, Zou Decun believed that he couldn't work at the standard of 260 yuan per square meter underground. Therefore, Zhang Mou confirmed Zou Decun's construction workload at the market price standard.", "Afterwards, Zhang Mou returned, and a quarrel broke out between Zhang Mou and Ma Ji.", "Zou Decun went to the company, and Zhang Mou learned that he had failed in the job application."} was obtained.
[0034] The second stage is event association.
[0035] Step 1: Use the prompt words to extract elements such as the time, characters, and the relationship between events in the event set E1 associated in the first stage. In this embodiment, there are four types of relationships between events: "causality", "promotion", "transition", and "none". To extract the above attributes, the prompt words must reflect the output requirements. For example, to enable the large language model to extract the time element of an event, a natural language text statement containing the keyword "event occurrence time" can be used, so that the large language model can extract the required elements according to the prompt words.
[0036] Events and the relationships between characters are inferred through two rules. Specific keywords in the event description can be used. For example, through "contract awarding" and "agreement", the relationship between characters can be determined as a contracting or cooperative relationship. It is also possible to infer the existing relationships through the semantic logic of the text. For example, from the text semantics of "The two parties agreed that the labor cost above ground would be calculated at 210 yuan per square meter and below ground at 260 yuan per square meter, but Zou Decun thought it was impossible to do the work at the standard of 260 yuan per square meter below ground" and "Afterwards, Zhang returned, and a quarrel broke out between Zhang and Ma Ji", it can be seen that due to the disagreement on the price between the two parties, a quarrel occurred. Therefore, there is a causal relationship between these two events.
[0037] The prompt words used in this embodiment are shown in the following table.
[0038] Table 1 Example Table of Prompt Words
[0039]
[0040] Step 2: Input the prompt words designed for each element to be extracted, the current event e0, and the set E1 into the large language model to make it output the relationships between events and characters. Denote the set of event relationships and character relationships as R E and R V respectively, which are represented as R E = {"e0, promotes, e1"; "e1, transitions to, e2"; "e2, causally related to, e3";...}, R V = {"v1, cooperates with, v2"; "v1, contracts with, v3";...}.
[0041] Taking the set E1 output in the first stage as an example:
[0042] R E={"On August 1, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project", promote, "On August 7, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project. The two parties agreed that the labor cost price above ground would be calculated at 210 yuan per square meter and underground at 260 yuan per square meter. However, Zou Decun thought it couldn't be done at the standard of 260 yuan per square meter underground, so Zhang confirmed Zou Decun's construction quantity at the market price standard"; "On August 7, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project. The two parties agreed that the labor cost price above ground would be calculated at 210 yuan per square meter and underground at 260 yuan per square meter. However, Zou Decun thought it couldn't be done at the standard of 260 yuan per square meter underground, so Zhang confirmed Zou Decun's construction quantity at the market price standard", cause and effect, "Afterwards, Zhang returned, and a quarrel broke out between Zhang and Ma Ji"; "Afterwards, Zhang returned, and a quarrel broke out between Zhang and Ma Ji", none, "Zou Decun went to the company, and Zhang learned that he failed in the job application"}. Take R E Regarding each element's event in R as a node and the relationship between events as a directed edge, R E set can be transformed into an event relationship graph.
[0043] R V ={"Shenyang No. 14 Construction Engineering Company", subcontract, "Li Hongjun"; "Li Hongjun", cooperate, "Ma Ji"; "Shenyang No. 14 Construction Engineering Company", subcontract, "Ma Ji"; "Shenyang No. 14 Construction Engineering Company", subcontract, "Zou Decun"; "Li Hongjun", cooperate, "Zou Decun"}. Regarding each element's person in R as a node and the relationship between persons as a directed edge, R V set can be transformed into a person relationship graph, as V shown. Figure 3 shown.
[0044] Step 3: Traverse the correlation degree D between the associated events in E1 and the current event e0 i (i = 1, 2, 3,...). When a given threshold T is set, screen the event set E2 whose correlation degree with e0 satisfies D ≤ T. Integrate the R E obtained in Step 2 of the second stage with the screened event set E2 (traverse R E according to the time priority principle. If R EIf both of the two events contained in the element are in the filtered event set E2, then this element is the fused element), denoted as R' E ={ <Event, Relationship, Event>,...}
[0045] In this embodiment, the correlation degree D is the minimum distance between the associated event and the current event e0 among the nodes of the event relationship graph. In the event relationship graph, node 0 is the current event e0 = "On August 1, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project"; node 1 is the event e1 = "On August 7, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project. The two parties agreed that the above-ground labor cost would be 210 yuan per square meter and the underground labor cost would be 260 yuan per square meter. However, Zou Decun thought that he couldn't do the work at the standard of 260 yuan per square meter underground. Therefore, Zhang confirmed Zou Decun's construction workload at the market price standard"; node 2 is the event e2 = "Afterwards, Zhang returned, and a quarrel broke out between Zhang and Ma Ji"; node 3 is the event e3 = "Zou Decun went to the company, and Zhang learned that he failed in the job application". Therefore, the correlation degree D1 between node 1 and node 0 is 1, the correlation degree D2 between node 2 and node 0 is 2, and the correlation degree D3 between node 3 and node 0 is 3. Given the threshold T = 2, if the correlation degree D <= 2, then the first three nodes including the current event e0 and the relationships between the nodes are filtered out in the event relationship graph. The filtered event relationship graph is as shown in Figure 4 shown
[0046] After being processed by step 3, the obtained R' E ={ "On August 1, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project", Promote, "On August 7, 2019, Shenyang No. 14 Construction Engineering Company subcontracted Buildings 1#, 2#, and 3# of the contracted project to Li Hongjun, Zhang, and Ma Ji. Li Hongjun orally agreed with Zou Decun that Zou Decun would contract the labor cost part of Buildings 1# and 2# and the property management room project. The two parties agreed that the above-ground labor cost would be 210 yuan per square meter and the underground labor cost would be 260 yuan per square meter. However, Zou Decun thought that he couldn't do the work at the standard of 260 yuan per square meter underground. Therefore, Zhang confirmed Zou Decun's construction workload at the market price standard"}
[0047] Step 4: Fuse R' EAll the elements in (i.e., merging the same events in different elements), are sorted according to the time relationship to form the associated event chain of event e0.
[0048] The overall process of this embodiment is as Figure 5 .
[0049] To further improve the accuracy of the associated event chain, the person relationship R obtained in step 2 V and the associated event chain obtained in step 4 can be input into the large language model to further optimize the associated event chain using the person relationship. For example, according to the person relationship, the large language model optimizes the mutual relationship between adjacent events in the associated event chain, making the associated event chain more accurate. The person relationship provides the large language model with an in-depth understanding of the interactions and connections between different people in the event. During the optimization process of the associated event chain, using the relationships between people, the large language model can better identify the associated patterns and logical order in the event, thereby improving the mutual relationship between adjacent events.
Claims
1. A text event association method based on a large language model, characterized in that: include: Step 1, using the prompt words and the large language model to extract event elements in the current text event e0 to obtain an event element set; Step 2, filter the event feature set using the interference word library and regular expressions; Step 3, using the event elements in the filtered event element set, searching for the associated events of the current text event e0 to obtain an associated event set; Step 4: Using the prompt words and the large language model, extract the occurrence time of each event in the associated event set, and arrange all the associated events in the associated event set in positive order according to the occurrence time to obtain the associated event set {e0, e1, e2, e3, ...} in positive order; Step 5, using the prompt words and the large language model to extract the relationship between every two events in {e0, e1, e2, e3, ...}, and obtain an event relationship set; the element in the event relationship set is "e i , e i With e j The relationship between j "; where i,j=0,1,2,3,... is the number of the event, i<j; Step 6, calculate the correlation between each event in {e0, e1, e2, e3, ...} and the current text event e0, filter out events with correlation less than or equal to the threshold from {e0, e1, e2, e3, ...}, and obtain a valid correlation event set; Step 7, traverse each element in the event relationship set, if the two associated events contained in the element are both in the valid associated event set, then keep the element, otherwise delete the element; after traversal, obtain the valid event relationship set; Step 8: merge all elements in the valid event relationship set to obtain the associated event chain of the current text event e0.
2. The text event association method based on a large language model as claimed in claim 1, characterized in that: Also includes: Step 5.1, using the prompt words, using the large language model to extract the relationship between the characters in all events in {e0, e1, e2, e3, ...}, and obtain a character relationship set; the element in the character relationship set is "v l , v l With v m The relationship between m ";in, l,m=1,2,3,... are the character numbers; Step 9, using the character relationship set and the large language model to optimize the associated event chain of the current text event e0 obtained in step 8.
3. The text event association method based on a large language model as described in claim 1 or 2, characterized in that: The large language model is the Tongyi Qianwen language model, the Mass Spectrum Qingyan or Deepseek.
Citation Information
Cited By
Cosmetic adverse reaction user report information automatic acceptance and feature recognition method
CN120376027A