Event context reduction method and system based on vector retrieval and large language model

By combining vector retrieval with a large language model, the problems of poor generalization and inaccurate temporal reasoning in existing event context restoration methods are solved, and efficient and accurate event extraction and timeline verification are achieved, which is suitable for application scenarios such as multi-document summarization and news analysis.

CN120780759AInactive Publication Date: 2025-10-14DATA SPACE RES INST

Patent Information

Application Number
CN202511294812.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing event context restoration methods rely on manual features and rule libraries, have poor generalization capabilities, and are difficult to adapt to complex cross-document, multi-paragraph, and multi-dimensional event data. In addition, the controllability and temporal reasoning accuracy of the content generated by large language models are insufficient, making it difficult to meet application scenarios with high verifiability requirements.

Method used

By combining vector retrieval and large language models, semantic retrieval is used to obtain similar text collections, construct information extraction prompt words, extract event information, and perform causal logic verification on the event timeline to ensure that events are arranged in the actual development order and identify potential contradictions, thereby improving the integrity and logical reliability of the output event chain.

Benefits of technology

It achieves efficient, accurate and structurally unified event extraction, significantly improving the accuracy and structuring level of event information extraction. It is suitable for large-scale document processing and has good general adaptability and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780759A_ABST
    Figure CN120780759A_ABST
Patent Text Reader

Abstract

The invention discloses an event venation reduction method and system based on vector retrieval and a large language model, and relates to the field of natural language processing, and the method comprises the following steps: obtaining a query word of a user and an original structured document corresponding to the query word; performing semantic retrieval in the original structured document by using the query word to obtain a similar text set which is similar to the query word in terms; based on the similar text set, constructing an information extraction cue word; performing event information extraction on the similar text set by utilizing a preset large language model according to the information extraction cue word to obtain structured event information; performing time sequence reconstruction on all structured events in the structured event information to obtain an event timeline; performing causal logic verification on the event timeline to obtain a causal logic verification result; and if the causal logic verification result is passed, outputting the event timeline. According to the invention, the integrity and logic reliability of the event chain in the output event timeline are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, and in particular to an event context restoration method and system based on vector retrieval and large language models. BACKGROUND

[0002] The event context restoration task aims to extract key events from multiple news documents and construct an event clue chain arranged in chronological order. This task is the intersection of time-aware information extraction and text summarization tasks, and is commonly used in multi-document summarization, news analysis and other applications.

[0003] Traditional methods mostly use a combination of information extraction and rule systems to identify event triples, temporal expressions and entity relationships from text, and perform simple sorting through temporal labels. However, such methods highly depend on manual features and rule libraries, have poor generalization ability, and are difficult to adapt to complex cross-document, multi-paragraph, multi-dimensional event data.

[0004] With the development of deep learning and large language models, methods have emerged that use large language models (such as ChatGPT) for event extraction and sorting. By designing natural language prompts (Prompts) to guide the model to generate structured event information, the complexity of feature design and the weakness of temporal reasoning in traditional methods are alleviated to some extent. This method has zero or few-shot adaptation capability, significantly reducing the dependence on labeled data and rule design, and is suitable for application environments with uncertain data distribution. However, this solution still has uncertainties in the controllability of generated content and the accuracy of temporal reasoning, and is prone to problems such as time confusion and factual bias, making it unsuitable for application scenarios with high requirements for output verifiability. SUMMARY

[0005] To solve the technical problems in the background art, the present application proposes an event context restoration method and system based on vector retrieval and large language models.

[0006] In a first aspect, the present application proposes an event context restoration method based on vector retrieval and large language models, comprising: obtaining a user's query word and an original structured document corresponding to the query word; performing semantic retrieval of the query word in the original structured document to obtain a similar text set semantically similar to the query word; based on the similar text set, constructing an information extraction prompt; using a pre-set large language model to perform event information extraction on the similar text set according to the information extraction prompt to obtain structured event information; reconstructing the time sequence of all structured events in the structured event information to obtain an event timeline; The event timeline is subjected to a cause-effect logic check to obtain a cause-effect logic check result; wherein the cause-effect logic check result comprises pass or fail; if the cause-effect logic check result is pass, the event timeline is output.

[0007] Preferably, semantic retrieval is performed on the original structured document using the query word to obtain a similar text set semantically similar to the query word, comprising: Encoding the query word to obtain a query vector; Content field extraction is performed on the original structured document to obtain a plurality of indexing texts; The plurality of indexing texts are encoded to obtain a plurality of high-dimensional semantic feature vectors; According to a predetermined semantic retrieval mechanism, a vector index structure is constructed; According to the query vector and the vector index structure, semantic retrieval is performed on the plurality of high-dimensional semantic feature vectors to obtain a plurality of high-dimensional semantic feature vectors semantically similar to the query vector; and according to the plurality of high-dimensional semantic feature vectors semantically similar to the query vector, a plurality of indexing texts semantically similar to the query word are obtained, and the plurality of indexing texts semantically similar to the query word are combined to form a similar text set.

[0008] Preferably, in the process of performing semantic retrieval on the original structured document using the query word, the same pre-trained semantic encoding model is used to encode the query word and the plurality of indexing texts.

[0009] Preferably, the vector index structure is an L2 distance-based vector index structure constructed using the FAISS semantic retrieval mechanism.

[0010] Preferably, based on the similar text set, an information extraction prompt word is constructed, comprising: The plurality of indexing texts in the similar text set are spliced to form a model input text; According to the model input text and a predetermined task template, an information extraction prompt word is constructed.

[0011] Preferably, all structured events in the structured event information are subjected to time sequence reconstruction to obtain an event timeline, comprising: The event occurrence time of all structured events in the structured event information is extracted, and all event occurrence times are standardized to a fixed time format; All structured events are sorted according to the event occurrence time to obtain an event timeline.

[0012] Preferably, the event timeline is subjected to a cause-effect logic check, specifically comprising: A logic judgment prompt word for judging the event time sequence relationship is constructed; Based on the logical judgment prompt words, the preset large language model is used to determine whether there is a temporal contradiction in the event timeline; if not, the causal logic check result is confirmed to be passed; if so, the causal logic check result is confirmed to be failed.

[0013] Preferably, in the process of reconstructing the time sequence of all structured events in the structured event information, if a structured event lacks the event occurrence time, relative timing inference is first performed based on the contextual reference relationship of the similar text set, entity association analysis or the time directionality of the action verb, and the preset large language model is called to determine the position of the structured event in the event chain of all structured events to complete the missing event occurrence time of the structured event, and then the time sequence of all structured events is reconstructed.

[0014] Preferably, if the result of the causal logic check fails, the event timeline is marked as a logical conflict, and the preset large language model is reused to extract event information from similar text sets based on information extraction prompt words, and the extracted structured event information is reconstructed in time sequence and causal logic checked until the result of the causal logic check passes.

[0015] In a second aspect, the present invention further proposes an event context restoration system based on vector retrieval and a large language model, comprising: An acquisition module is used to obtain the user's query words and the original structured documents corresponding to the query words; The preprocessing module is used to perform semantic retrieval in the original structured documents using the query words to obtain a set of similar texts with similar semantics to the query words; The prompt word construction module is used to construct information extraction prompt words based on similar text sets; The information extraction module is used to extract event information from similar text sets based on information extraction prompt words using a preset large language model to obtain structured event information; The time sequence reconstruction and causal logic verification module is used to reconstruct the time sequence of all structured events in the structured event information to obtain an event timeline; perform causal logic verification on the event timeline to obtain a causal logic verification result; wherein the causal logic verification result includes pass or fail; if the causal logic verification result is pass, the event timeline is output.

[0016] In the present application, the event context restoration method and system based on vector retrieval and large language model are proposed, which combines semantic retrieval with structured information extraction prompt words to guide the large language model to automatically identify, extract and output structured event information in a unified format from the original structured document corresponding to the query word, realizes efficient, accurate and structured unified event extraction, and significantly improves the accuracy and structured level of event information extraction. Moreover, by reconstructing the time sequence of all structured events in the structured event information, an event timeline is obtained. The event timeline is subjected to causal logic verification, which not only ensures that the extracted events are arranged in the actual development order, but also identifies potential time or causal contradictions between events, effectively improving the integrity and logical reliability of the event chain in the output event timeline.

[0017] In addition, the present application is suitable for large-scale document processing tasks and can adjust the extraction target and format requirement according to different application scenarios, and has good general adaptability and scalability. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 The flowchart of the event context restoration method based on vector retrieval and large language model in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0019] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0020] In a first aspect, with reference to Figure 1 The event context restoration method based on vector retrieval and large language model proposed by the present application comprises: obtaining a query word of a user and an original structured document corresponding to the query word; performing semantic retrieval in the original structured document using the query word to obtain a similar text set semantically similar to the query word; constructing information extraction prompt words based on the similar text set; extracting event information from the similar text set according to the information extraction prompt words using a preset large language model to obtain structured event information; reconstructing the time sequence of all structured events in the structured event information to obtain an event timeline; verifying the event timeline for causal logic, and if the causal logic verification is passed, outputting the event timeline that passes the causal logic verification.

[0021] The application combines semantic retrieval with structured information extraction prompt words to guide the large language model to automatically identify, extract and output structured event information in a unified format from the original structured document corresponding to the query word, realize efficient, accurate and structured unified event extraction, and significantly improve the accuracy and structured level of event information extraction. Moreover, by reconstructing the time sequence of all structured events in the structured event information, an event timeline is obtained. The event timeline is checked for causal logic, which not only ensures that the extracted events are arranged in the actual development order, but also identifies potential time or causal contradictions between events, effectively improving the completeness and logical reliability of the event chain in the output event timeline.

[0022] In addition, the application is suitable for large-scale document processing tasks and can adjust the extraction target and format requirement according to different application scenarios, and has good general adaptability and scalability.

[0023] In this embodiment, semantic retrieval is performed on the original structured document using the query word to obtain a similar text set semantically similar to the query word, including: encoding the query word to obtain a query vector; extracting the content field of the original structured document to obtain a plurality of indexable texts; encoding the plurality of indexable texts to obtain a plurality of high-dimensional semantic feature vectors; constructing a vector index structure according to a preset semantic retrieval mechanism; performing semantic retrieval on the plurality of high-dimensional semantic feature vectors according to the query vector and the vector index structure to obtain a plurality of high-dimensional semantic feature vectors semantically similar to the query vector; and obtaining a plurality of indexable texts semantically similar to the query word according to the plurality of high-dimensional semantic feature vectors semantically similar to the query vector, and combining the plurality of indexable texts semantically similar to the query word to form a similar text set.

[0024] In this way, the embodiment can efficiently and accurately obtain a similar text set semantically similar to the query word from the original structured document, thereby facilitating subsequent efficient, accurate and structured unified event extraction.

[0025] In this embodiment, the same pre-trained semantic encoding model is used to encode the query word and the plurality of indexable texts during the semantic retrieval process in the original structured document, to avoid errors caused by using different semantic encoding models.

[0026] In a further embodiment, the query word is encoded to obtain a query vector, specifically including: using a pre-trained semantic encoding model m3e-base to encode the query word to obtain a query vector, so as to retain the semantic features of the query word.

[0027] In a further embodiment, a plurality of texts to be indexed are encoded to obtain a plurality of high-dimensional semantic feature vectors, including: The pre-trained semantic encoding model m3e-base is used to encode several texts to be indexed to obtain several high-dimensional semantic feature vectors to retain the semantic features of the texts to be indexed.

[0028] In one specific embodiment, constructing a vector index structure according to a preset semantic retrieval mechanism includes: constructing a vector index structure based on L2 distance using the FAISS semantic retrieval mechanism.

[0029] With such a configuration, this embodiment can quickly complete nearest neighbor retrieval in large-scale high-dimensional semantic feature vectors, significantly shortening the retrieval time; it can better measure the similarity between semantic vectors, so that the retrieval results are more in line with the semantic relevance requirements, thereby improving the retrieval accuracy; moreover, it can adapt to data sets of different sizes and can be flexibly expanded to a variety of index structures (such as flat index, inverted file index, quantitative index, etc.) to adapt to a variety of application scenarios.

[0030] In a further embodiment, semantic retrieval is performed in multiple high-dimensional semantic feature vectors based on the query vector and the vector index structure to obtain several high-dimensional semantic feature vectors that are semantically similar to the query vector, specifically including: According to the vector index structure of the query vector and the L2 distance, the nearest neighbor search is performed in multiple high-dimensional semantic feature vectors to obtain several high-dimensional semantic feature vectors with similar semantics to the query vector.

[0031] In one specific embodiment, multiple documents in the original structured document are first read in batches and content fields are extracted to form multiple texts to be indexed; then, a locally deployed pre-trained semantic encoding model m3e-base is used to batch vectorize the multiple texts to be indexed, and each text to be indexed is converted into a fixed-length high-dimensional semantic feature vector through an embedding model to retain its semantic features; after completing the construction of the high-dimensional semantic feature vector, FAISS is used to construct an efficient vector index structure (FAISS index) based on L2 distance; the query term entered by the user is also encoded into a semantic vector through the same pre-trained semantic encoding model m3e-base, and a nearest neighbor search is performed in the FAISS index to retrieve several high-dimensional semantic feature vectors that are semantically closest to the query vector, thereby determining several texts to be indexed that are semantically closest to the query term based on the several high-dimensional semantic feature vectors that are semantically closest to the query vector.

[0032] In other embodiments, the distance measurement method based on L2 distance is replaced by calculating inner product or cosine similarity, or the vector search framework is replaced by Elasticsearch + vector search plug-in, which can realize the combination of keyword search and semantic search.

[0033] In this embodiment, based on the similar text set, information extraction prompt words are constructed, including: Splicing several to-be-indexed texts in the similar text set to form model input text; Construct event extraction prompt words based on the model input text and preset task template.

[0034] This configuration of the embodiment can significantly improve the target focusing capability of the generated model and the consistency of the output format.

[0035] The preset task template in this embodiment includes clarifying the task objectives, output requirements and event element structure in a natural language, so that the event collection prompt words can be used to guide the preset large language model to extract the event overview, event occurrence time, and "entity-action-object-time-location" event five-tuple and other core information of each structured event.

[0036] In order to improve domain adaptability, the preset task template also includes feature words or structural requirements for limiting application scenarios to embed specific events, so as to enhance the structural constraints and domain orientation of prompt words, thereby guiding the preset large language model to generate event life cycle descriptions that are more in line with the actual context.

[0037] It should be noted that the preset large language model in this embodiment has the function call capability to perform inference processing on prompt words and similar text sets to obtain structured event information.

[0038] During the inference process, after receiving input, the large language model generates structured event information using a registered function tool (Tool), which defines a unified structure for structured event information. This structured event information includes three fields for each structured event: an event overview, the time of occurrence, and the event quintuple. The quintuple fields sequentially include entity, action, object, time, and location. The large language model automatically parses the task intent based on the prompt word content and completes the corresponding fields. The entire extraction process is completed internally by the large model, without relying on additional regular expressions or post-processing rules.

[0039] In specific implementation, this embodiment supports two generation paths: one is to use the HuggingFace framework to load the locally deployed large language model; the other is to remotely call the large model service through the OpenAI interface-style API to achieve inference generation.

[0040] With such a configuration, this embodiment can automatically identify and extract event subject information from complex texts, and output event structure data in a unified format, significantly improving the accuracy and structuring level of event information extraction.

[0041] In this embodiment, all structured events in the structured event information are reconstructed in time sequence to obtain an event timeline, including: Extract the occurrence time of all structured events in the structured event information and standardize them into a fixed time format to facilitate subsequent rapid sorting; All structured events are sorted according to their occurrence time to obtain an event timeline.

[0042] The fixed time format in this embodiment is YYYY-MM-DD.

[0043] In order to effectively improve the integrity and logical reliability of the event chain in the output event timeline, in this embodiment, a causal logic check is performed on the event timeline, specifically including: Construct logical judgment prompt words for judging the temporal relationship of events; Based on the logical judgment prompt word, a preset large language model is used to determine whether there are any timing contradictions in the event timeline. If not, that is, there are no timing contradictions, the causal logic verification result is confirmed to have passed; if so, that is, there are timing contradictions, the causal logic verification result is confirmed to have failed. It should be understood that the large language model in this embodiment is the same as the large model used in the event information extraction process, but the prompts used for different tasks are different.

[0044] In this embodiment, if the causal logic verification result is failed, the event timeline is marked as a logical conflict, and the preset large language model is reused to extract event information from a similar text set based on the information extraction prompt words, and the extracted structured event information is reconstructed in time sequence and causal logic verification is performed until the causal logic verification result is passed.

[0045] With this arrangement, after identifying the event timeline that fails the causal logic check, this embodiment can also correct potential time or causal contradictions between structured events, effectively improving the integrity and logical reliability of the event chain.

[0046] In this embodiment, in the process of reconstructing the time sequence of all structured events in the structured event information, if a structured event lacks the event occurrence time, relative timing inference is first performed based on the contextual reference relationship of the similar text set, entity association analysis or the time directionality of the action verb, and the preset large language model is called to determine the position of the structured event in the event chain of all structured events to complete the missing event occurrence time of the structured event, and then the time sequence of all structured events is reconstructed.

[0047] This arrangement allows this embodiment to complete structured events that lack temporal information, ensuring that the extracted events are arranged in the order of actual development. This helps improve the temporal rationality and semantic consistency of the subsequent event timeline, further enhancing the integrity and logical reliability of the event chain. It should be understood that the large language model used in this embodiment is the same model as the large model used in the event information extraction process; it's just that the prompts used for different tasks are different.

[0048] In a second aspect, the present invention further proposes an event context restoration system based on vector retrieval and a large language model, comprising: An acquisition module is used to obtain the user's query words and the original structured documents corresponding to the query words; The preprocessing module is used to perform semantic retrieval in the original structured documents using the query words to obtain a set of similar texts with similar semantics to the query words; The prompt word construction module is used to construct information extraction prompt words based on similar text sets; The information extraction module is used to extract event information from similar text sets based on information extraction prompt words using a preset large language model to obtain structured event information; The time sequence reconstruction and causal logic verification module is used to reconstruct the time sequence of all structured events in the structured event information to obtain an event timeline; perform causal logic verification on the event timeline to obtain a causal logic verification result; wherein the causal logic verification result includes pass or fail; if the causal logic verification result is pass, the event timeline is output.

[0049] In the semantic retrieval process of the preprocessing module, the query words are first encoded to obtain a query vector; content fields are extracted from the original structured document to obtain multiple texts to be indexed; the multiple texts to be indexed are encoded to obtain multiple high-dimensional semantic feature vectors; and a vector index structure is constructed according to the preset semantic retrieval mechanism. According to the query vector and the vector index structure, semantic retrieval is performed in several high-dimensional semantic feature vectors to obtain several high-dimensional semantic feature vectors with similar semantics to the query vector; and according to the several high-dimensional semantic feature vectors with similar semantics to the query vector, several to-be-indexed texts with similar semantics to the query word are obtained, and the several to-be-indexed texts with similar semantics to the query word are combined to form a similar text set.

[0050] In the process of semantic retrieval in the original structured document using query words, the same pre-trained semantic encoding model is used to encode the query words and several texts to be indexed respectively.

[0051] Among them, the vector index structure is a vector index structure based on L2 distance constructed using the FAISS semantic retrieval mechanism.

[0052] The process of constructing information extraction prompt words includes: splicing several to-be-indexed texts in a similar text set to form a model input text; and constructing information extraction prompt words based on the model input text and a preset task template.

[0053] In the process of time sequence reconstruction, the event occurrence time of all structured events in the structured event information is first extracted, and all event occurrence times are standardized into a fixed time format; then all structured events are sorted according to the event occurrence time to obtain the event timeline.

[0054] In the process of performing causal logic verification on the event timeline, first construct a logical judgment prompt word for judging the temporal relationship of events; based on the logical judgment prompt word, use the preset large language model to determine whether there is a temporal contradiction in the event timeline; if not, confirm that the causal logic verification result is passed; if so, confirm that the causal logic verification result is failed.

[0055] Among them, in the process of reconstructing the time sequence of all structured events in the structured event information, if a structured event lacks the event occurrence time, relative timing inference is first performed based on the context reference relationship of similar text sets, entity association analysis or the time directionality of action verbs, and the large language model is called to determine the position of the structured event in the event chain of all structured events to complete the missing event occurrence time of the structured event, and then the time sequence of all structured events is reconstructed.

[0056] Among them, if the result of the causal logic verification is failed, the event timeline will be marked as a logical conflict, and the preset large language model will be reused to extract event information from similar text sets based on information extraction prompt words, and the extracted structured event information will be reconstructed in time sequence and causal logic verified until the causal logic verification result is passed.

[0057] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An event context restoration method based on vector retrieval and large language model, characterized in that: include: Obtain the user's query terms and the original structured documents corresponding to the query terms; Use the query words to perform semantic retrieval in the original structured documents to obtain a set of similar texts with similar semantics to the query words; Based on similar text sets, construct information extraction prompt words; Use the preset large language model to extract event information from similar text sets based on information extraction prompt words to obtain structured event information; Reconstruct the time sequence of all structured events in the structured event information to obtain an event timeline; Perform a causal logic check on the event timeline to obtain a causal logic check result; wherein the causal logic check result includes pass or fail; if the causal logic check result is pass, the event timeline is output.

2. The event context restoration method based on vector retrieval and large language model according to claim 1 is characterized in that: Use the query words to perform semantic retrieval in the original structured documents to obtain a set of similar texts with similar semantics to the query words, including: Encode the query word to obtain the query vector; Extract content fields from the original structured document to obtain multiple texts to be indexed; Encode multiple texts to be indexed to obtain multiple high-dimensional semantic feature vectors; Construct a vector index structure based on the preset semantic retrieval mechanism; According to the query vector and the vector index structure, semantic retrieval is performed in multiple high-dimensional semantic feature vectors to obtain several high-dimensional semantic feature vectors with similar semantics to the query vector; and according to the several high-dimensional semantic feature vectors with similar semantics to the query vector, several to-be-indexed texts with similar semantics to the query word are obtained, and the several to-be-indexed texts with similar semantics to the query word are combined to form a similar text set.

3. The event context restoration method based on vector retrieval and large language model according to claim 2 is characterized in that: In the process of semantic retrieval in the original structured document using query words, the same pre-trained semantic encoding model is used to encode the query words and several texts to be indexed respectively.

4. The event context restoration method based on vector retrieval and large language model according to claim 2, characterized in that: The vector index structure is constructed based on L2 distance using the FAISS semantic retrieval mechanism.

5. The event context restoration method based on vector retrieval and large language model according to claim 1 is characterized in that: Based on similar text sets, construct information extraction prompt words, including: Splicing several to-be-indexed texts in the similar text set to form model input text; Construct information extraction prompt words based on the model input text and preset task template.

6. The event context restoration method based on vector retrieval and large language model according to claim 1, characterized in that: Reconstruct the time sequence of all structured events in the structured event information to obtain the event timeline, including: Extract the event occurrence time of all structured events in the structured event information and standardize all event occurrence times into a fixed time format; All structured events are sorted according to their occurrence time to obtain an event timeline.

7. The event context restoration method based on vector retrieval and large language model according to claim 1 is characterized in that: Perform causal logic verification on the event timeline, including: Construct logical judgment prompt words for judging the temporal relationship of events; Based on the logical judgment prompt words, the preset large language model is used to determine whether there is a temporal contradiction in the event timeline; if not, the causal logic check result is confirmed to be passed; if so, the causal logic check result is confirmed to be failed.

8. The event context restoration method based on vector retrieval and large language model according to claim 1 is characterized in that: In the process of reconstructing the time sequence of all structured events in structured event information, if a structured event lacks the event occurrence time, relative timing inference is first performed based on the contextual reference relationship of similar text sets, entity association analysis or the time directionality of action verbs, and the preset large language model is called to determine the position of the structured event in the event chain of all structured events to complete the missing event occurrence time of the structured event, and then the time sequence of all structured events is reconstructed.

9. The event context restoration method based on vector retrieval and large language model according to claim 1, characterized in that: If the causal logic verification result fails, the event timeline will be marked as a logical conflict, and the preset large language model will be reused to extract event information from similar text sets based on information extraction prompt words. The extracted structured event information will be reconstructed in time sequence and causal logic verification will be performed until the causal logic verification result passes.

10. An event context restoration system based on vector retrieval and large language model, characterized in that: include: An acquisition module is used to obtain the user's query words and the original structured documents corresponding to the query words; The preprocessing module is used to perform semantic retrieval in the original structured documents using the query words to obtain a set of similar texts with similar semantics to the query words; The prompt word construction module is used to construct information extraction prompt words based on similar text sets; The information extraction module is used to extract event information from similar text sets based on information extraction prompt words using a preset large language model to obtain structured event information; The time sequence reconstruction and causal logic verification module is used to reconstruct the time sequence of all structured events in the structured event information to obtain an event timeline; perform causal logic verification on the event timeline to obtain a causal logic verification result; wherein the causal logic verification result includes pass or fail; if the causal logic verification result is pass, the event timeline is output.

Citation Information

Patent Citations

  • Intelligent search engine system based on Simbert algorithm

    CN118964524A

  • Event context generation method and device based on large language model and medium

    CN119782520A

  • Event context generation method and device, computer equipment and storage medium

    CN120123602A

  • Text file generation method and device, equipment, medium and product

    CN120561260A

  • Method and system for electronic processing of user queries maintaining factual consistency during processing

    US12204524B1

Cited By

  • Power grid operation event extraction and disposal method and system based on large language model

    CN121210629A

  • Life cycle process list calculation method based on large language model

    CN121279258A