Information extraction method, device, storage medium, and program product

By constructing a directed graph for information extraction and calling a large language model, the problem of "illusion" in the information extraction task is solved, and the accuracy of information extraction results is improved.

CN119918508BActive Publication Date: 2025-07-01JIHUA LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510408346.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-01
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In the information extraction task, the large language models in the prior art may experience 'illusion' phenomenon, generating seemingly reasonable but practically wrong content, resulting in low accuracy of information retrieval.

Method used

By constructing a directed graph of information extraction, the information extraction task is decomposed into the main node and the auxiliary node, the correlation between the main information to be extracted and the auxiliary information is clearly defined, and a large language model is called for information extraction, reducing the freedom of the model and avoiding the generation of irrelevant or fictitious content.

Benefits of technology

It effectively improves the accuracy of information extraction results, reduces the possibility of large language models generating wrong content, and makes the information extraction process more controllable and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918508B_ABST
    Figure CN119918508B_ABST
Patent Text Reader

Abstract

The present application discloses an information extraction method, device, storage medium, and program product, relating to the technical field of information extraction, including: obtaining an information extraction task and a document to be extracted; constructing an information extraction directed graph corresponding to the document to be extracted based on the information extraction task; wherein, the information extraction directed graph includes at least one main node and affiliated nodes associated with the main node, each main node corresponds to a type of main information to be extracted, and each affiliated node corresponds to a type of affiliated information to be extracted; based on the information extraction directed graph, calling a large language model to perform an information extraction operation on the document to be extracted to obtain an information extraction result; wherein, the information extraction result includes a main information extraction result and an affiliated information extraction result. The present application can solve the technical problem of low information extraction accuracy in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of information extraction, and in particular, to an information extraction method, device, storage medium, and program product. Background Art

[0002] In the related art, an extraction task and an original document can be directly used as input information and input into a large language model. The large language model utilizes its powerful parameter base to capture and understand the input information, and then a corresponding information extraction result can be obtained.

[0003] However, in practical applications, for relatively complex extraction tasks, the large language model may exhibit the "hallucination" phenomenon, that is, it may create content that seems reasonable but is actually incorrect, thereby reducing the accuracy of information retrieval. Summary of the Invention

[0004] The main purpose of this application is to provide an information extraction method, device, storage medium, and program product, aiming to solve the technical problem of low information extraction accuracy in the related art.

[0005] To achieve the above object, this application proposes an information extraction method, which includes:

[0006] Obtain an information extraction task and a document to be extracted;

[0007] Based on the information extraction task, construct an information extraction directed graph corresponding to the document to be extracted; wherein, the information extraction directed graph includes at least one main node and affiliated nodes associated with the main node, each main node corresponds to a type of main information to be extracted, and each affiliated node corresponds to a type of affiliated information to be extracted;

[0008] Based on the information extraction directed graph, call the large language model to perform an information extraction operation on the document to be extracted to obtain an information extraction result; wherein, the information extraction result includes a main information extraction result and an affiliated information extraction result.

[0009] In one embodiment, the step of calling the large language model to perform an information extraction operation on the document to be extracted based on the information extraction directed graph to obtain an information extraction result includes:

[0010] Based on the information extraction directed graph, call the large language model to perform a main information extraction operation on the document to be extracted to obtain a main information extraction result; wherein, the main information extraction result includes at least one set of main information sets, and each set of main information sets includes multiple main information;

[0011] Based on the main information extraction result and the information extraction directed graph, call the large language model to extract affiliated information associated with the main information from the document to be extracted to obtain an affiliated information extraction result;

[0012] Based on the backbone information extraction result and the affiliated information extraction result, an information extraction result is obtained.

[0013] In one embodiment, before the step of invoking a large language model to perform backbone information extraction operation on the document to be extracted based on the information extraction directed graph and obtaining the backbone information extraction result, it further includes:

[0014] Based on the information extraction directed graph, determine a combination of prior relationships between multiple backbone information to be extracted;

[0015] The step of invoking a large language model to perform backbone information extraction operation on the document to be extracted based on the information extraction directed graph and obtaining the backbone information extraction result includes:

[0016] For each backbone information to be extracted, invoke a large language model to perform zero-prior information extraction on the document to be extracted to obtain a zero-prior backbone information set; wherein, each zero-prior backbone information set corresponds to a backbone information to be extracted, and each zero-prior backbone information set includes at least one backbone information;

[0017] Based on the combination of prior relationships, determine a conditional backbone information set as prior conditions from all zero-prior backbone information sets;

[0018] Take each backbone information in the conditional backbone information set as prior information respectively, and invoke a large language model to perform prior information extraction on the document to be extracted to obtain a prior backbone information set corresponding to the conditional backbone information set;

[0019] Based on the conditional backbone information set and the prior backbone information set, determine the backbone information extraction result.

[0020] In one embodiment, the step of invoking a large language model to extract affiliated information associated with the backbone information from the document to be extracted based on the backbone information extraction result and the information extraction directed graph and obtaining the affiliated information extraction result includes:

[0021] When the affiliated information to be extracted is preset structure information, for each backbone information in the backbone information extraction result, take the backbone information as a prior condition, and invoke a large language model to perform prior information extraction on the document to be extracted to obtain the first affiliated information;

[0022] When the affiliated information to be extracted is non-preset structure information, determine initial backbone information from the backbone information extraction result;

[0023] Take the initial backbone information as a prior condition, and invoke a large language model to perform prior information extraction on the document to be extracted to obtain the second affiliated information;

[0024] Update the prior condition based on the second subsidiary information, and return to execute the step of calling the large language model to perform prior information extraction on the document to be extracted to obtain the second subsidiary information until the preset stop condition of the information extraction operation is met;

[0025] Obtain the subsidiary information extraction result based on all the first subsidiary information and all the second subsidiary information.

[0026] In one embodiment, the step of calling the large language model to perform information extraction operation on the document to be extracted based on the information extraction directed graph to obtain the information extraction result includes:

[0027] Segment the document to be extracted according to a preset scaling factor to obtain multiple sub-documents;

[0028] Based on the information extraction directed graph, perform information extraction operations on each sub-document respectively to obtain the first information extraction result corresponding to each sub-document;

[0029] For the sub-document to which each first information extraction result belongs, segment the target document according to a preset scaling factor to obtain sub-target documents;

[0030] Update the sub-document based on the sub-target document, and return to execute the step of performing information extraction operations on each sub-document respectively based on the information extraction directed graph to obtain the first information extraction result corresponding to each sub-document until the byte length corresponding to the sub-target document is less than or equal to the preset byte length;

[0031] Obtain the information extraction result based on all the first information extraction results.

[0032] In one embodiment, the step of obtaining the information extraction result based on all the first information extraction results includes:

[0033] Perform average weighting on all the first information extraction results to obtain the weighted scores corresponding to each extracted information in all the first information extraction results;

[0034] Based on the weighted scores, perform a screening and filtering operation on all the extracted information to obtain the information extraction result.

[0035] In one embodiment, the step of obtaining the document to be extracted includes:

[0036] Obtain the initial document; wherein, the initial document is the original text that needs to perform information extraction;

[0037] In the case that the number of bytes of the initial document exceeds the preset number of bytes, segment the initial document to obtain multiple documents to be extracted; wherein, the number of bytes of each document to be extracted does not exceed the preset number of bytes;

[0038] When the number of bytes of the initial document does not exceed a preset number of bytes, the initial document is determined as the document to be extracted.

[0039] In addition, to achieve the above object, the present application also provides an information extraction device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the information extraction method as described above.

[0040] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the information extraction method as described above.

[0041] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the information extraction method as described above.

[0042] One or more technical solutions proposed by the present application have at least the following technical effects:

[0043] In the information extraction method proposed by the present application, an information extraction task and a document to be extracted can be obtained; and based on the information extraction task, an information extraction directed graph corresponding to the document to be extracted is constructed; wherein, the information extraction directed graph includes at least one main node and affiliated nodes associated with the main node, each main node corresponds to a piece of main information to be extracted, and each affiliated node corresponds to a piece of affiliated information to be extracted; then, according to the information extraction directed graph, a large language model is called for information extraction to obtain a final information extraction result including the main information extraction result and the affiliated information extraction result. The present application decomposes a complex information extraction task into an information extraction directed graph including main nodes and affiliated nodes, whereby the association relationship between each piece of main information to be extracted and each piece of affiliated information to be extracted can be clarified according to the directed graph, so that when calling the large model for information extraction, the large model can consider the logical relationship between information as a whole, play a role in restricting and guiding the extraction process of the large model, reduce the freedom degree of the large language model, avoid generating irrelevant or fictional inaccurate content, and thus can effectively improve the accuracy of the information extraction result. Description of the Drawings

[0044] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 It is a schematic flowchart provided for the first embodiment of the information extraction method of the present application;

[0047] Figure 2 It is a directed graph of information extraction for an example;

[0048] Figure 3 It is a schematic flowchart for the refinement of step S300;

[0049] Figure 4 It is a schematic flowchart for the refinement of step S310;

[0050] Figure 5 It is a schematic flowchart for the refinement of step S320;

[0051] Figure 6 It is a schematic flowchart provided for the second embodiment of the information extraction method of the present application;

[0052] Figure 7 It is a schematic diagram of the device structure of the hardware operating environment involved in the information extraction method in the embodiments of the present application.

[0053] The realization of the purpose of the present application, functional features and advantages will be further described in combination with the embodiments with reference to the drawings. Specific Embodiments

[0054] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0055] To better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.

[0056] The main solution of the embodiment of the present application is: obtaining an information extraction task and a document to be extracted; based on the information extraction task, constructing a directed graph of information extraction corresponding to the document to be extracted; wherein, the directed graph of information extraction includes at least one main node and affiliated nodes associated with the main node, each main node corresponds to a type of backbone information to be extracted, and each affiliated node corresponds to a type of affiliated information to be extracted; based on the directed graph of information extraction, calling a large language model to perform an information extraction operation on the document to be extracted to obtain an information extraction result; wherein, the information extraction result includes a backbone information extraction result and an affiliated information extraction result.

[0057] In the field of information extraction, especially information extraction based on large models, it is in a stage of rapid technological development. The goal of Information Extraction (IE) is to extract structured knowledge such as entities, relationships, events, etc. from natural language text. This field faces challenges in specific task patterns and complex text expressions. With the emergence of Large Language Models (LLMs), we see a new possibility of using the natural language understanding capabilities of these models to handle information extraction tasks. Through their powerful parameter base, large language models can capture and understand complex patterns and relationships in text, thus showing great potential in information extraction tasks.

[0058] In related technologies, information extraction methods mainly extract from the already identified target paragraphs, and the information to be extracted needs to be concentrated in that paragraph. This not only increases the screening work but also limits the application scenarios of this method, lacking sufficient flexibility and generalization ability. Using large language models can solve the problems of flexibility and generalization ability. The extraction task and the original document are directly used as input information and input into the large language model. Through its powerful parameter base, the large language model can capture and understand the input information and obtain the corresponding information extraction result.

[0059] In practical applications, although the flexibility and generalization ability of large language models are relatively better, for relatively complex extraction tasks, large language models still exhibit the "hallucination" phenomenon, that is, they may create content that seems reasonable but is actually incorrect, thus reducing the accuracy of information retrieval.

[0060] This application provides a solution. A complex information extraction task can be decomposed into an information extraction directed graph containing a main node and subsidiary nodes. From this, the association relationship between each main information to be extracted and each subsidiary information to be extracted can be clarified. When calling a large model for information extraction, the large model can consider the logical relationship between information as a whole, play a role in constraining and guiding the extraction process of the large model, reduce the degree of freedom of the large language model, and avoid generating irrelevant or fictional inaccurate content, thereby effectively improving the accuracy of the information extraction result.

[0061] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. The following takes an information extraction device as an example to illustrate this embodiment and the following embodiments.

[0062] Based on this, an information extraction method is provided in an embodiment of this application, referring to Figure 1 ,Figure 1 This is a schematic flowchart of the first embodiment of the information extraction method of this application.

[0063] In this embodiment, the information extraction method includes steps S100 to S300:

[0064] Step S100, obtain an information extraction task and a document to be extracted.

[0065] Step S200, based on the information extraction task, construct an information extraction directed graph corresponding to the document to be extracted.

[0066] Among them, the information extraction directed graph includes at least one main node and affiliated nodes associated with the main node. Each main node corresponds to a type of main information to be extracted, and each affiliated node corresponds to a type of affiliated information to be extracted.

[0067] Step S300, based on the information extraction directed graph, call a large language model to perform an information extraction operation on the document to be extracted, and obtain an information extraction result.

[0068] Among them, the information extraction result includes a main information extraction result and an affiliated information extraction result.

[0069] Specifically, when there is an information extraction requirement, the user can input an information extraction task and a document to be extracted into the information extraction device; among them, the document to be extracted can be a text file such as an enterprise report, an academic paper, or social media; the information extraction task indicates the specific content that needs to be obtained from the document to be extracted this time. For example, the document to be extracted can be an academic paper on the surface reaction of organic molecules, and the information extraction task can be to extract the relevant content of the surface reaction of the experiment in this academic paper.

[0070] According to the information extraction task, it is possible to determine the main information to be extracted and the affiliated information to be extracted that the information extraction task may correspond to, and construct an information extraction directed graph including at least one main node and affiliated nodes according to the relationship between each main information to be extracted and the affiliated information to be extracted; in the information extraction directed graph, one main node corresponds to a type of main information to be extracted, and one affiliated node corresponds to a type of affiliated information to be extracted. Generally, the affiliated information to be extracted depends on the main information to be extracted; the main node and the affiliated node can be organized into a directed graph according to the logical relationship of the information they correspond to, and the edges of the directed graph can represent the dependency or association relationship between the nodes. For example, when the information extraction task is surface reaction, it can be determined that such as Figure 2The directed graph for information extraction shown. In this directed graph for information extraction, reactants, products, and substrates are three types of main information to be extracted, while the reactive functional groups of reactants, reaction sites of products, reaction intermediate products (intermediate product 1, intermediate product 2, etc.), and reaction conditions (reaction condition 1, reaction condition 2, etc.) are auxiliary information to be extracted. It should be noted that in practical applications, the information extraction device can analyze and process the information extraction task using existing large models to obtain a possible directed graph for information extraction; or the user can, based on manual experience, analyze the information extraction task in advance to obtain the directed graph for information extraction, and then input this directed graph for information extraction into the information extraction device during information extraction, enabling the information extraction device to perform relevant information extraction operations based on this directed graph and the document to be extracted.

[0071] After determining the directed graph for information extraction, a large language model with text processing capabilities and / or image capabilities and capable of information extraction can be directly called, such as GPT, Wenxin Yiyan, etc. These models have powerful language understanding and generation capabilities and can handle complex natural language tasks. The document to be extracted, the directed graph for information extraction, and the corresponding task instructions are passed to the large language model. Thus, the large language model can perform information extraction based on the logical relationships between the nodes in the directed graph for information extraction, obtaining an information extraction result that includes the main information extraction result and the auxiliary information extraction result. The directed graph can clarify the goals and logical associations during extraction, play a role in restricting and guiding the information extraction process of the large language model, reduce the degree of freedom of the large model, and avoid generating irrelevant or fictional inaccurate content.

[0072] In a feasible implementation manner, step S300 may specifically include steps S310 to S330, as Figure 3 shown, Figure 3 is a schematic diagram of the refined process of step S300.

[0073] Step S310, based on the directed graph for information extraction, call the large language model to perform main information extraction operations on the document to be extracted, obtaining the main information extraction result.

[0074] Among them, the main information extraction result includes at least one set of main information sets, and each set of main information sets includes multiple main information.

[0075] Step S320, based on the main information extraction result and the directed graph for information extraction, call the large language model to extract auxiliary information associated with the main information from the document to be extracted, obtaining the auxiliary information extraction result.

[0076] Step S330, based on the main information extraction result and the auxiliary information extraction result, obtain the information extraction result.

[0077] Specifically, the backbone information to be extracted is clarified in the information extraction directed graph. Based on the backbone information to be extracted, a large language model can be called to process the document to be extracted, and the possible backbone information can be extracted from the document to be extracted. For example, the backbone information to be extracted includes three types: reactants, substrates, and products. The large language model can extract all specific reactants, substrates, and products from the document to be extracted input by the user. Information of the same type can be regarded as the same set of backbone information. For example, all the extracted reactants constitute a set of backbone information, and the multiple backbone information included in this set of backbone information are multiple specific reactants. In a feasible implementation, the large language model can be called in the way of zero-prior extraction and prior information extraction to obtain the backbone information extraction result. Therefore, before step S310, the prior relationship combinations between multiple backbone information to be extracted can also be determined based on the information extraction directed graph. In this implementation, step S310 may specifically include steps S311 to S314, as Figure 4 shown Figure 4 is a schematic diagram of the refined process of step S310.

[0078] Step S311, for each backbone information to be extracted, call the large language model to perform zero-prior information extraction on the document to be extracted, and obtain a zero-prior backbone information set.

[0079] Among them, each zero-prior backbone information set corresponds to a backbone information to be extracted, and each zero-prior backbone information set includes at least one backbone information.

[0080] Step S312, based on the prior relationship combination, determine the conditional backbone information set as the prior condition from all zero-prior backbone information sets.

[0081] Step S313, take each backbone information in the conditional backbone information set as the prior information, call the large language model to perform prior information extraction on the document to be extracted, and obtain the prior backbone information set corresponding to the conditional backbone information set.

[0082] Step S314, based on the conditional backbone information set and the prior backbone information set, determine the backbone information extraction result.

[0083] It should be noted that the information extraction directed graph can reflect the logical relationship between multiple backbone information to be extracted. According to this logical relationship, multiple prior relationship combinations can be determined for subsequent prior information extraction. For example, in a surface reaction, both the reactant and the substrate have a clear causal relationship with the product, which can be expressed as: reactant → product ← substrate. At this time, the following three prior relationship combinations can be determined: 1) Using the reactant and the substrate as the prior to infer the product; 2) Using the reactant and the product as the prior to infer the substrate; 3) Using the product and the substrate as the prior to infer the reactant.

[0084] For each piece of backbone information to be extracted in the information extraction directed graph, a large language model can be called to perform zero-prior information extraction on the document to be extracted, extracting all possible backbone information in the document to be extracted to obtain a zero-prior backbone information set; zero-prior information extraction means that without prior knowledge, a large model is directly called to extract the required information to be extracted from the document to be extracted; the backbone information of the same type extracted can be grouped into the same zero-prior backbone information set. Further, according to the prior relationship combination between the backbone information to be extracted, a conditional backbone information set as a prior condition is determined from all zero-prior backbone information sets. Each backbone information in the conditional backbone information set is used as a prior condition respectively, and the backbone information related to these prior conditions is further extracted. The result of the prior information extraction is called the prior backbone information set. Integrating the conditional backbone information set and its corresponding prior backbone information set can form the final backbone information extraction result.

[0085] To facilitate the understanding of the above content, here T is used to represent the document to be extracted, and T contains the target information d corresponding to multiple groups of information extraction tasks (for example, if the information extraction task is to extract the surface reaction in a paper, the paper contains multiple groups of surface reactions d), that is, T = {d 1 ,…,d k}; where k represents the number of target information d; in subsequent representations, superscripts are used to distinguish different groups of information within a document to be extracted, and subscripts are used to distinguish different information within a group (for example, reactant, substrate, and product are three different pieces of information, and the subscripts can be represented by 1, 2, and 3 respectively). Sometimes, for simplicity of description, the superscript will be ignored. When the superscript is ignored, it is default to describe the information within the same group. The set of backbone information included in the i-th group of target information can be represented as M i ; M i = {m1 i ,..,m n i} (m is the backbone information, and n is the number of backbone information in this group), and the corresponding set of accessory information can be represented as S i ; S i = {s1 i ,..,s l i} (s is the accessory information, and l is the number of accessory information); d 1 = [M 1 ,S 1 , d 2 = [M 2 ,S 2 , d 3 = [M 3 ,S 3 … Note that M i = {m1,..,m nThe structure of} is fixed, while S i may vary. For example, in a chemical reaction, there may be multiple intermediate reaction steps. Therefore, we can cluster a specific type of backbone information M1 = {m1 1 , m1 2 , m1 3 , m1 4 …} to facilitate further information extraction from the dimension of a single backbone information accordingly.

[0086] When performing the operation of extracting backbone information, a function for extracting target information based on a large language model can be defined as f. Then, for the set of backbone information M1, it can be expressed as f(M1|T). This method requires no prior knowledge, that is, zero prior knowledge. For prior information extraction, prior knowledge needs to be provided as a prior condition for information extraction. For example, given m2 i , if the corresponding possible M1 needs to be extracted, it can be expressed as (M1|T, m2 i ). Here, three types of backbone information to be extracted, M1, M2, and M3, are used as examples for illustration. Among them, M1 and M2 are the causes, and M3 is the result, that is, M1→M3←M2. For each type of backbone information to be extracted, zero prior information extraction is performed on T, and M 1,zero = f(M1|T); M 2,zero = f(M2|T); M 3,zero = f(M3|T); where M 1,zero , M 2,zero , and M 3,zero respectively represent three zero-prior backbone information sets obtained based on M1, M2, and M3.

[0087] On this basis, prior information extraction is carried out, that is, using the results of zero-prior extraction as prior conditions to constrain the output space of the large language model. For the sake of understanding, taking M1→M2 as an example, given M1 to extract M2 (that is, using M1 as the conditional backbone information set for prior extraction to obtain the corresponding M2), and then given M2 to extract M1 (that is, using M2 as the conditional backbone information set for prior extraction to obtain the corresponding M1). Thus, for each backbone information m1 1,zero in M i = f(M1|T), M 2,prior = f(M2|T, m1 i ); for each m2 2,zero in M i = f(M2|T), M 1,prior = f(M1|T, m2 i ); M 1,prior and M 2,prior are two obtained prior backbone information sets; that is, taking M1,zero Each main information m1 in i is used as a prior condition to extract the prior information of M2 respectively, and each main information m1 is obtained i The corresponding prior main information. The set of these prior main information is M 2,prior ; The acquisition of M 1,prior is the same reason, which will not be elaborated here. Combining the set of conditional main information and the corresponding set of prior main information, the result of main information extraction can be obtained. In the above example, the result of main information extraction is two groups {m1 i , m2 i |M 1,zero} (representing the set of zero prior main information of M1 and the prior main information M2 obtained based on M1) and {m1 i , m2 i |M 2,zero} (representing the set of zero prior main information of M2 and the prior main information M1 obtained based on M2).

[0088] Based on this, for the aforementioned situation of M1→M3←M2, there are three combinations of prior relationships: 1) M1 and M2 are used as priors to infer M3; 2) M1 and M3 are used as priors to infer M2; 3) M2 and M3 are used as priors to infer M1. The combination of using M1 and M2 as priors to infer M3 is illustrated (using other combinations or all combinations are feasible). Specifically, taking a certain m1 1,zero in M i as a prior condition to extract M2 and M3: For each m1 1,zero in M i = f(M1|T), M 2,prior = f(M2|T, m1 i ); Then use the main information m2 2,prior in M i to combine with the corresponding m1 i to extract M3, that is, for each m1 1,zero in M i = f(M1|T) and each m2 2,prior in M i = f(M2|T), M 3,prior = f(M3|T, m1 i , m2 i ); From this, a set of extraction results {m1 i , m2 i , m3 i |M 1,zero , M 2,prior} can be obtained; Similarly, taking M 2,zero as a prior condition and performing the above operations, another set {m1i , m2 i , m3 i |M 2,zero , M1,prior}; Finally, two sets of results can be obtained. To obtain as accurate a result of the backbone information extraction as possible, the multiple sets of results obtained can be processed by weighting, taking the intersection, taking the union, or other methods. By choosing different merging methods, the coverage rate (i.e., obtaining as many results as possible) and the accuracy rate (i.e., obtaining as accurate results as possible) can be balanced. Taking {m1 i , m2 i , m3 i |M 1,zero , M 2,prior} and {m1 i , m2 i , m3 i |M 2,zero , M 1,prior} as an example for the two sets of results, the intersection of the two sets can be selected or other confidence metrics can be used for the merging operation. This operation balances the final results of the extracted information in terms of recall and accuracy, and can increase the accuracy of the extracted information results while ensuring the coverage rate of the information extraction. For example, the intersection of the two sets can be selected to maximize the accuracy, and the final result of the backbone information extraction is {m1 i , m2 i , m3 i |M 1,zero , M 2,prior} ∩ {m1 i , m2 i , m3 i |M 2,zero , M 1,prior}. In addition, for cases where there are multiple steps of backbone information to be extracted, such as M1 → M2 → M3 → M4 → M5, the aforementioned steps can be executed step by step in segments, such as first M1 → M2 and then M2 → M3, until the entire extraction operation of the backbone information to be extracted is completed to obtain the final result of the backbone information extraction.

[0089] Based on the zero-prior backbone information set, calling the large language model for prior backbone information extraction can constrain the output space of the large language model, reduce the difficulty of task implementation, thereby reducing the possibility of "hallucinations" of the large language model (i.e., the results output by the large language model are not based on any factual data and there is information that does not conform to the facts), and increasing the accuracy of information extraction.

[0090] After obtaining the backbone information extraction result, it is possible to further use the information extraction directed graph and the above-mentioned backbone information extraction result as a basis to call a large language model to extract subsidiary information associated with the backbone information from the document to be extracted, obtaining a subsidiary information extraction result; the subsidiary information extraction result is a supplement and detailed description of the backbone information, which can help enhance the integrity of the entire information extraction. In a feasible implementation manner, step S320 may specifically include steps S321 to S325, as Figure 5 shown, Figure 5 is a schematic diagram of the refined process of step S320.

[0091] Step S321, when the subsidiary information to be extracted is preset structure information, for each backbone information in the backbone information extraction result, using the backbone information as a prior condition, call a large language model to perform prior information extraction on the document to be extracted, and obtain the first subsidiary information.

[0092] Step S322, when the subsidiary information to be extracted is non-preset structure information, determine the initial backbone information from the backbone information extraction result.

[0093] Step S323, using the initial backbone information as a prior condition, call a large language model to perform prior information extraction on the document to be extracted, and obtain the second subsidiary information.

[0094] Step S324, update the prior condition based on the second subsidiary information, and return to execute the step of calling a large language model to perform prior information extraction on the document to be extracted to obtain the second subsidiary information until the preset stop condition for the information extraction operation is met.

[0095] Step S325, based on all the first subsidiary information and all the second subsidiary information, obtain the subsidiary information extraction result.

[0096] Specifically, after extracting the backbone information, subsidiary information extraction can be performed according to the type of the subsidiary information to be extracted. The subsidiary information to be extracted can be preset structure information, that is, subsidiary information with a fixed structure similar to molecular properties, experimental conditions, etc.; when the subsidiary information to be extracted is preset structure information, the aforementioned prior information extraction method can be adopted, taking each backbone information in the backbone information extraction result as a prior condition respectively, and combining the preset structure information to call a large language model for information extraction to obtain the first subsidiary information. For example, when extracting the preset structure subsidiary information s1 to be extracted, for each backbone information m1 in the backbone information extraction result, s1 = f(s1|T, m1).

[0097] The additional information to be extracted may also be non - preset structural information, that is, the structure of the additional information to be extracted is not fixed, and the latter additional information to be extracted may depend on the result of the previous additional information to be extracted. For example, in surface reactions, the reaction process, reaction conditions, etc. are usually chain structures of variable length; when the additional information to be extracted is non - preset structural information, the initial backbone information can be determined from the result of the backbone information extraction and used as a prior condition to call the large - language model for prior information extraction to obtain the second additional information; since there may be a front - back dependency relationship between the additional information to be extracted, the obtained second additional information can be used as a new prior condition, and then the step of calling the large - language model to perform prior information extraction on the document to be extracted to obtain the second additional information is executed until the preset stop condition is met to stop the information extraction operation. For example, in addition to constructing the information extraction function f, a preset stop condition function g can also be constructed. After each information extraction, g is used to judge whether the stop condition is reached. If not, the current extracted information (i.e., the extracted second additional information) is used as a new prior condition to perform information extraction again, and the above process is continuously looped until g reaches the preset condition. The following takes surface reaction as an example to illustrate: The initial state (i.e., the initial backbone information) can be determined as t initial ; t initial = (reactant m1 i , substrate m2 i ); The final state is t final ; t final = (product m3 i , substrate m2 i ); The function f extracts the next intermediate state t i according to the current state t i+1 = f(T, t i , t initial , t final ); The function g judges whether the final state is reached according to the current state t i if_end = g(T, t i , t initial , t final ) (that is, the preset stop condition is the current state = the final state). It can be expressed in pseudocode as:

[0098] "The additional information of the reaction process r = [[]

[0099] t i <- t initial

[0100] While not g(if_end|T, t i , t initial , t final ):

[0101] t i+1 = f(t i+1 |T, t i , t initial , t final )

[0102] r.append(t i )

[0103] t i <- t i+1 ”;

[0104] Using the above steps, the extraction of all non-fixed-structure accessory information (i.e., the second accessory information) can be completed; then, by integrating all the first accessory information and all the second accessory information, the extraction result of the accessory information of all accessory information can be obtained. Finally, by comprehensively combining the main information extraction result and the accessory information extraction result, a complete and structured information extraction result can be obtained.

[0105] In addition, since this application needs to call a large language model, and the current large language models usually support the processing of input texts with lengths from 16k to 128k, corresponding to 50k to 300k characters; for most academic articles (except reviews), these large language models can directly process them. However, for articles that exceed the one-time processing length of the large language model, they need to be pre-processed. Therefore, in a feasible implementation manner, the steps for obtaining the document to be extracted can be specifically: obtaining an initial document; where the initial document is the original text that needs to be information-extracted; in the case where the number of bytes of the initial document exceeds the preset number of bytes, the initial document is segmented to obtain multiple documents to be extracted; where the number of bytes of each document to be extracted does not exceed the preset number of bytes; in the case where the number of bytes of the initial document does not exceed the preset number of bytes, the initial document is determined as the document to be extracted. That is, when the number of bytes of the initial document meets the processing requirements of the large language model, the large language model is directly called to perform the aforementioned information extraction operation on the initial document; when the number of bytes of the initial document exceeds the processable range of the large language model, it is necessary to further divide the initial document into multiple documents to be extracted whose number of bytes does not exceed the processable range of the large language model, so as to call the large language model to perform the aforementioned information extraction operation. For example, the map-reduce approach can be adopted to divide the original article into multiple parts, and after independently processing each part, the results are merged. For example, for the extraction of the main information M1, after separately extracting M1 for each part, the union of the results is taken to obtain the complete result.

[0106] It is not difficult to understand that the information extraction method provided in this application can directly input the original data (such as the original text of an article, without excerpting the target paragraph) and output the final extraction result, realizing end-to-end information extraction without manual intervention. And by constructing a directed graph for the information extraction task, it is divided into the main information to be extracted and the subsidiary information to be extracted, and the extraction of complex information is transformed into sub-problems that are easier to solve. From this, the association relationship between each main information to be extracted and each subsidiary information to be extracted can be clarified according to the directed graph. When calling the large model for information extraction, the large model can consider the logical relationship between information as a whole. It can not only clarify the association relationship between each main information to be extracted and each subsidiary information to be extracted according to the directed graph, so that when calling the large model for information extraction, the large model can consider the logical relationship between information as a whole, which plays a role of constraint and guidance in the extraction process of the large model, reduces the degree of freedom of the large language model, avoids generating irrelevant or fictional inaccurate content, but also avoids information omission and incompleteness caused by local information extraction, thereby effectively improving the accuracy of the information extraction result.

[0107] Since in the extraction of some more complex tasks, due to the limitation of the large model's capabilities, there may still be a certain degree of "hallucination" problem in the output result of the model, and the extracted information is still not accurate enough. Therefore, based on Embodiment 1 of this application, Embodiment 2 of this application is proposed. In Embodiment 2 of this application, the same or similar content as that in the above Embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , step S300 may include steps A100~A500 to improve the accuracy of the information extraction process:

[0108] Step A100, divide the document to be extracted according to a preset scaling factor to obtain multiple sub-documents.

[0109] Step A200, based on the information extraction directed graph, perform information extraction operations on each sub-document respectively to obtain the first information extraction result corresponding to each sub-document.

[0110] Step A300, for each sub-document to which the first information extraction result belongs, divide the target document according to a preset scaling factor to obtain sub-target documents.

[0111] Step A400, update the sub-document based on the sub-target document, and return to execute the step of performing information extraction operations on each sub-document respectively based on the information extraction directed graph to obtain the first information extraction result corresponding to each sub-document until the byte length corresponding to the sub-target document is less than or equal to the preset byte length.

[0112] Step A500, obtain the information extraction result based on all the first information extraction results.

[0113] Specifically, when performing information extraction, the document to be extracted can be segmented according to a preset scaling ratio. For example, if the document to be extracted is 50k bytes and the preset scaling ratio is 0.2, then 5 sub-documents of 10k bytes can be obtained by segmenting according to the writing order in the document. Then, according to the information extraction directed graph, information extraction operations are respectively performed on each sub-document to obtain the first information extraction results of each sub-document (specifically, the information extraction process in the first embodiment described above can be referred to). Further, the information extraction range is shrunk, that is, for the sub-documents to which the first information extraction results belong, the target document is segmented again according to the preset scaling factor to obtain sub-target documents. For example, for a 10k-byte sub-document, the sub-document is further segmented into 5 sub-target documents of 2k bytes according to the preset scaling ratio of 0.2. The sub-target documents are used as new sub-documents, and then based on the information extraction directed graph, the above-mentioned information extraction operations are respectively performed on each sub-document to obtain the first information extraction results corresponding to the current sub-documents. The process of segmenting the document according to the preset scaling ratio is continuously repeated until the byte length corresponding to the sub-target document is less than or equal to the preset byte length. By integrating all the first information extraction results, the information extraction result can be obtained. The above division of the document to be extracted according to the preset scaling ratio can realize the gradual shrinkage of the information extraction range, that is, it can realize the segmentation from coarse to fine. By restricting the range of fine-grained segmentation, it is ensured to the greatest extent that the information corresponding to the information extraction task will not be wrongly removed.

[0114] In practical applications, the above operation of dividing the document to be extracted can be implemented based on the adaptive retrieval enhanced method (RAG: Retrieval Augmented Generation). Its basic principle is described through the following three steps:

[0115] 1) Define the task: For an information extraction task with a prior condition p, its conventional extraction method is generally o = f(o|t,p), where t is the input document, p is the prior condition, and o is the extraction result corresponding to the information extraction task.

[0116] 2) Construct a retrieval question based on p: Based on the prior condition p, construct a retrieval question for RAG, query(p), and then based on the retrieval question, use RAG to obtain relevant content from the original text. For the retrieval method, considering that the prior condition may usually include specific terms, a hybrid retrieval method is adopted, that is, the results of sparse and dense retrieval are combined; among them, sparse retrieval is usually based on information such as keywords and term frequencies (such as the BM25 algorithm), and dense retrieval is based on the similarity of feature vectors.

[0117] 3) Adaptive segmented construction of data to be retrieved: In RAG, the original text is usually segmented with a fixed length or retrieved through multi-level segmentation. However, considering the knowledge-intensive nature of the task, this application adopts an adaptive approach of gradually shrinking the retrieval scope. For example, a coarse-to-fine segmentation is used. In the first step, the segmentation uses 1200 words, in the second step 600 words, and in the third step 300 words until the preset fine-grained range is reached, to ensure that the target information is not erroneously removed to the greatest extent. The following gives a sample pseudocode for illustration: The minimum segmentation length (i.e., the preset byte length) is l min , and the scaling factor of each segmentation length (i.e., the preset scaling ratio) is α. The corresponding pseudocode can be expressed as:

[0118] "The extraction result o of the target extraction task of RAG rag =

[0119] The content t retrieved by RAG rag =

[0120] l i <- len(T) * α # Segmentation length

[0121] t i <- T # Document to be extracted # Content input to the large language model

[0122] o = f(o|t i ,p) # Result extracted based on the document to be extracted

[0123] o rag .append(o)

[0124] while l i >l min :

[0125] t i+1 =retrieval(t i ,l i ,query(p))

[0126] l i <- len(t i+1 )*α

[0127] o = f(o|t i+1 ,p)

[0128] o rag .append(o)

[0129] t rag .append(t i+1 )

[0130] ti <- t i+1 ”;

[0131] Finally, the information extraction results obtained according to the information extraction task based on different article paragraphs can be obtained.o rag and the corresponding article paragraphs t rag . Assume that there are 4 stages of division during extraction (such as 1200 bytes, 600 bytes, 300 bytes, and 150 bytes in sequence). At this time,o rag there are results of 4 stages in it,o rag = [[a1,a2,a3],[a1,a3],[a2,a3],[a1,a3]]; where a1, a2, and a3 are the extracted information in the extraction results; [a1,a2,a3] is the first information extraction result corresponding to the 1200 - byte division, [a1,a3] is the first information extraction result corresponding to the 600 - byte division, [a2,a3] is the first information extraction result corresponding to the 300 - byte division, and [a1,a3] is the first information extraction result corresponding to the 150 - byte division. It should be noted that the above - mentioned adaptive method of gradually shrinking the information extraction range can be used in the main information extraction process and also in the subsidiary information extraction process. Integrating all the first information extraction results can obtain the final information extraction result.

[0132] In order to further improve the accuracy of the information extraction results, obtain more effective information extraction results, and avoid redundant extracted information in the extraction results, in a feasible implementation manner, step A500 may specifically include steps A510 - A520:

[0133] Step A510, perform average weighting on all the first information extraction results to obtain the weighted scores corresponding to each extracted information in all the first information extraction results.

[0134] Step A520, based on the weighted scores, perform a screening and filtering operation on all the extracted information to obtain the information extraction result.

[0135] Specifically, the importance or credibility of each extracted information in the above-obtained first information extraction result can be evaluated by the method of average weighting, providing a basis for subsequent screening; for example, corresponding weights can be assigned to each extraction result according to the information type and the correlation between information, and then the corresponding weighted scores are calculated for each extracted information; information screening can be performed accordingly. For example, a screening threshold is preset, and for all extracted information, screening is performed according to its weighted score, filtering out information with lower scores. For example, assuming the preset screening threshold is 0.75, only the extracted information with a weighted score higher than 0.75 is retained, and the finally retained extracted information is integrated to obtain the information extraction result; or, the weighted score can reflect the confidence of the corresponding extracted information, and the corresponding weighted score can also be calculated according to the extracted information, the confidence of each information extraction result is calculated, and information screening processing can be performed according to this confidence to obtain the final information extraction result.

[0136] Taking the foregoing content as an example, currently, the extraction result of o rag = [[a1,a2,a3],[a1,a3],[a2,a3],[a1,a3]] is obtained. By using the weighted average method, that is, the weighted scores of each step are the same, the weighted result can be obtained as follows: the weighted score of a1 is 0.75, the weighted score of a2 is 0.5, and the weighted score of a3 is 1.0. Or for some extraction tasks, the first information extraction result obtained by setting the finest-grained division steps according to actual needs can also correspond to a higher weighted score. According to the weighted scores of each extracted information, the confidence of [a1,a2,a3] can be calculated as (0.75 + 0.5 + 1) / 3 = 0.75; the confidence of [a1,a3] is (0.75 + 1) / 2 = 0.875; the confidence of [a2,a3] is (0.5 + 1) / 2 = 0.75; and then for o rag = [[a1,a2,a3],[a1,a3],[a2,a3],[a1,a3]], the sum of the confidences of all extraction results is averaged to obtain the confidence mean of o rag as (0.75 + 0.875 + 0.75 + 0.875) / 4 = 0.815. This confidence reflects the reliability of this extraction result and can provide a reference for subsequent processing.

[0137] It can be understood that the information extraction method provided in this embodiment can gradually shrink the extraction range of the document to be extracted. For various information to be extracted corresponding to the information extraction task, weighted statistics can be performed through the corresponding retrieval enhancement results (i.e., the first information extraction result) to obtain the confidence of the corresponding extraction result, thereby further improving the result accuracy.

[0138] The present application provides an information extraction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the information extraction method in the first embodiment above.

[0139] Reference is made below Figure 7 , which shows a schematic structural diagram of an information extraction device suitable for implementing the embodiments of the present application. The information extraction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), etc., and fixed terminals such as desktop computers. Figure 7 The information extraction device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0140] As Figure 7 shown, the information extraction device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the information extraction device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the information extraction device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an information extraction device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0141] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0142] The information extraction device provided in the present application adopts the information extraction method in the above embodiments, and can solve the technical problem of low accuracy of information extraction in the related art. Compared with the related art, the beneficial effects of the information extraction device provided in the present application are the same as those of the information extraction method provided in the above embodiments, and other technical features in the information extraction device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0143] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0144] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0145] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the information extraction method in the above embodiments.

[0146] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0147] The above computer-readable storage medium can be included in the information extraction device; or it can exist separately without being assembled into the information extraction device.

[0148] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the information extraction device, the information extraction device is caused to: obtain an information extraction task and a document to be extracted; based on the information extraction task, construct an information extraction directed graph corresponding to the document to be extracted; wherein the information extraction directed graph includes at least one main node and affiliated nodes associated with the main node, each main node corresponding to a type of main information to be extracted, and each affiliated node corresponding to a type of affiliated information to be extracted; based on the information extraction directed graph, call a large language model to perform an information extraction operation on the document to be extracted to obtain an information extraction result; wherein the information extraction result includes a main information extraction result and an affiliated information extraction result.

[0149] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0151] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0152] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned information extraction method, and can solve the technical problem of low accuracy of information extraction in the related art. Compared with the related art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the information extraction method provided in the above embodiments, and will not be elaborated here.

[0153] The present application also provides a computer program product, including a computer program, which implements the steps of the information extraction method as described above when executed by a processor.

[0154] The computer program product provided by the present application can solve the technical problem of low accuracy of information extraction in the related art. Compared with the related art, the beneficial effects of the computer program product provided by the present application are the same as those of the information extraction method provided by the above embodiments, and will not be elaborated here.

[0155] The above are only partial embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. An information extraction method, characterized in that: The information extraction method comprises: Obtain information extraction tasks and documents to be extracted; Based on the information extraction task, construct an information extraction directed graph corresponding to the document to be extracted; wherein the information extraction directed graph includes at least one main node and subsidiary nodes associated with the main node, each of the main nodes corresponds to a type of main information to be extracted, and each of the subsidiary nodes corresponds to a type of subsidiary information to be extracted; Based on the information extraction directed graph, calling the large language model to perform information extraction operation on the document to be extracted, and obtaining information extraction results; wherein the information extraction results include main information extraction results and subsidiary information extraction results; The step of calling a large language model to perform an information extraction operation on the document to be extracted based on the information extraction directed graph to obtain an information extraction result comprises: Based on the information extraction directed graph, calling the large language model to perform a trunk information extraction operation on the document to be extracted, and obtaining the trunk information extraction result; wherein the trunk information extraction result includes at least one group of trunk information sets, and each group of the trunk information sets includes multiple trunk information; Based on the trunk information extraction result and the information extraction directed graph, calling a large language model to extract the subsidiary information associated with the trunk information from the document to be extracted, to obtain the subsidiary information extraction result; Based on the main information extraction result and the subsidiary information extraction result, the information extraction result is obtained; wherein the performing of the main information extraction operation and / or extracting the subsidiary information associated with the main information from the document to be extracted is obtained by performing prior information extraction based on a priori conditions; Before the step of calling a large language model to perform a main information extraction operation on the document to be extracted based on the information extraction directed graph and obtaining the main information extraction result, the step further includes: Determine a priori relationship combinations between a plurality of trunk information to be extracted based on the information extraction directed graph; The step of calling a large language model to perform a main information extraction operation on the document to be extracted based on the information extraction directed graph to obtain a main information extraction result comprises: For each type of backbone information to be extracted, a large language model is called to perform zero-prior information extraction on the document to be extracted to obtain a zero-prior backbone information set; wherein each zero-prior backbone information set corresponds to one type of backbone information to be extracted, and each zero-prior backbone information set includes at least one backbone information; Based on the prior relationship combination, determining a conditional backbone information set as a priori condition from all the zero prior backbone information sets; Each of the backbone information in the conditional backbone information set is used as prior information, and a large language model is called to extract prior information from the document to be extracted, so as to obtain a prior backbone information set corresponding to the conditional backbone information set; Based on the conditional backbone information set and the priori backbone information set, the backbone information extraction result is determined.

2. The information extraction method according to claim 1, characterized in that: The step of extracting the ancillary information associated with the main information from the document to be extracted by calling a large language model based on the main information extraction result and the information extraction directed graph to obtain the ancillary information extraction result comprises: In the case where the to-be-extracted subsidiary information is preset structural information, for each of the trunk information in the trunk information extraction result, the trunk information is used as a priori condition, and a large language model is called to perform priori information extraction on the to-be-extracted document to obtain first subsidiary information; In the case where the auxiliary information to be extracted is non-preset structural information, determining initial trunk information from the trunk information extraction result; Taking the initial trunk information as a priori condition, calling the large language model to extract priori information from the document to be extracted, and obtaining the second subsidiary information; The prior condition is updated based on the second subsidiary information, and the step of calling the large language model to extract the prior information of the document to be extracted to obtain the second subsidiary information is returned to be executed until a preset stop condition of the information extraction operation is met; The auxiliary information extraction result is obtained based on all the first auxiliary information and all the second auxiliary information.

3. The information extraction method according to claim 1, characterized in that: The step of calling a large language model to perform an information extraction operation on the document to be extracted based on the information extraction directed graph to obtain an information extraction result comprises: Segmenting the document to be extracted according to a preset zoom factor to obtain a plurality of sub-documents; Based on the information extraction directed graph, performing information extraction operations on each of the sub-documents respectively to obtain a first information extraction result corresponding to each of the sub-documents; For each sub-document to which the first information extraction result belongs, segmenting the target document according to the preset scaling factor to obtain sub-target documents; The sub-document is updated based on the sub-target document, and the step of returning to the step of executing the information extraction directed graph based on the information extraction, performing information extraction operations on each of the sub-documents respectively, and obtaining a first information extraction result corresponding to each of the sub-documents, until the byte length corresponding to the sub-target document is less than or equal to a preset byte length; Based on all the first information extraction results, the information extraction result is obtained.

4. The information extraction method according to claim 3, characterized in that: The step of obtaining the information extraction result based on all the first information extraction results comprises: Performing average weighting on all the first information extraction results to obtain weighted scores corresponding to each extracted information in all the first information extraction results; Based on the weighted scores, all the extracted information is screened and filtered to obtain the information extraction result.

5. The information extraction method according to any one of claims 1 to 4, characterized in that: The step of obtaining the document to be extracted comprises: Obtaining an initial document; wherein the initial document is an original text from which information extraction is required; When the number of bytes of the initial document exceeds the preset number of bytes, the initial document is segmented to obtain a plurality of the documents to be extracted; wherein the number of bytes of each of the documents to be extracted does not exceed the preset number of bytes; When the number of bytes of the initial document does not exceed a preset number of bytes, the initial document is determined as the document to be extracted.

6. An information extraction device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the information extraction method according to any one of claims 1 to 5.

7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the information extraction method according to any one of claims 1 to 5 are implemented.

8. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the information extraction method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Document level relationship extraction fusing enhanced entity and multi-level representation

    CN119599019A