Information extraction method and device, electronic equipment and storage medium
By acquiring target documents and rules, using named entity recognition and large models for information extraction, and combining semantic similarity and rule matching, the problem of incomplete performance task extraction in existing technologies is solved, and more accurate and closed-loop performance task management is achieved.
Patent Information
- Application Number
- CN202510772756.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-10
AI Technical Summary
When extracting information from documents, especially when extracting performance tasks from contract texts, existing technologies are prone to omitting information in the target rules, resulting in inaccurate and incomplete extraction.
By obtaining the target documents and related target rules, using named entity recognition and large models to extract information, combining semantic similarity and rule matching, supplementing and updating the information extracted from the documents, and building a task dependency graph to store and manage fulfillment tasks.
It improves the accuracy and completeness of information extraction, avoids missing information in the target rules, ensures that the extracted fulfillment tasks comply with the target role permissions, reduces fulfillment disputes, and realizes a closed loop of information management.
Smart Images

Figure CN120764657A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to the field of artificial intelligence such as natural language processing, deep learning, large models, and specifically relates to an information extraction method and device, an electronic device, and a storage medium. BACKGROUND
[0002] Extracting information from a document is one of the important tasks in the field of natural language. For example, performance tasks can be extracted from contract text. Contract performance tasks usually refer to related tasks that both parties need to complete within the contract period. In order to facilitate the management of contract performance tasks, performance tasks in electronic contracts can be extracted. For example, in the after-sales service of software service privatization delivery, the performance tasks in the contract need to be extracted. SUMMARY
[0003] The present application provides an information extraction method, device, electronic device, and storage medium. The specific solutions are as follows:
[0004] According to an aspect of the present application, an information extraction method is provided, comprising:
[0005] obtaining a target document and a target rule related to the target document;
[0006] extracting first association information associated with a target role from the target document;
[0007] obtaining second association information associated with the target role from the target rule;
[0008] updating the first association information according to the second association information to obtain target association information of the target role.
[0009] According to another aspect of the present application, an information extraction device is provided, comprising:
[0010] a first obtaining module configured to obtain a target document and a target rule related to the target document;
[0011] an extraction module configured to extract first association information associated with a target role from the target document;
[0012] a second obtaining module configured to obtain second association information associated with the target role from the target rule;
[0013] an updating module configured to update the first association information according to the second association information to obtain target association information of the target role.
[0014] According to another aspect of the present application, an electronic device is provided, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.
[0018] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.
[0019] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.
[0022] Figure 1 A flowchart of an information extraction method provided in one embodiment of the present application;
[0023] Figure 2 A flowchart of an information extraction method provided in another embodiment of the present application;
[0024] Figure 3 A flowchart of an information extraction method provided in another embodiment of the present application;
[0025] Figure 4 A flowchart of an information extraction method provided in another embodiment of the present application;
[0026] Figure 5 A schematic diagram of the structure of an information extraction device provided in one embodiment of the present application;
[0027] Figure 6 It is a block diagram of an electronic device used to implement the information extraction method of the embodiment of the present application. DETAILED DESCRIPTION
[0028] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.
[0030] The following describes the information extraction method, device, electronic device, and storage medium of the embodiments of the present application with reference to the accompanying drawings.
[0031] Figure 1 A flowchart of an information extraction method provided in one embodiment of the present application.
[0032] The information extraction method of the embodiment of the present application can be executed by the information extraction device of the embodiment of the present application, and the device can be configured in an electronic device.
[0033] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.
[0034] like Figure 1 As shown, the information extraction method includes:
[0035] Step 101: Obtain a target document and target rules related to the target document.
[0036] The target rule related to the target document may refer to a rule requirement that matches the role association information targeted by the target document.
[0037] For example, the target document may be a contract text, an agreement, a service agreement, etc., and the relevant target rules may include but are not limited to one or more of industry specifications, business specifications, company regulations, legal regulations, etc.
[0038] The target document may be a document in plain text format, a document in PDF format, or a document in other formats, which is not limited.
[0039] To facilitate information extraction, in this application, OCR (Optical Character Recognition) can be used to convert text-containing images or PDF files into a target document in text format, preserving the hierarchical structure such as paragraphs and clause numbers. This eliminates formatting interference and provides standardized input for subsequent semantic analysis. For example, OCR can be used to convert a contract document in image or PDF format into contract text.
[0040] Step 102: extract first association information associated with the target role from the target document.
[0041] The target role can be a designated role or any of the roles involved in the target document. For example, if the target document is a contract, the target role could be Party B. Since Party B may have multiple responsible persons to fulfill the contract, the target role could also be a specific responsible person.
[0042] For example, the target roles in the target document can be identified based on the named entity recognition model. For example, the named entity recognition model can be used to identify roles such as Party A and supplier in the after-sales service contract for the private delivery of software services.
[0043] In this application, the first associated information may be information related to the target role in the target document, such as the role name, tasks, rights, restrictions, time requirements, triggering events, etc. related to the target role.
[0044] Taking the target document as a contract text as an example, the first related information can be the task that the target role in the target document needs to perform, that is, the performance task. The performance task may include but is not limited to the target role (such as the responsible person), performance task content, process steps and other information.
[0045] Taking the target document as an agreed agreement as an example, the first associated information may be the rights that the target role in the agreed agreement can execute, or the first associated information may be the obligations that the target role in the agreed agreement needs to fulfill, etc.
[0046] For example, in the scenario of extracting performance tasks from a contract text, performance tasks can include explicit performance tasks and implicit performance tasks. Explicit performance tasks can be tasks explicitly described in the target document, such as "Party B must deliver the goods before June 1, 2024." Implicit performance tasks can be tasks inferred through context analysis, such as "pay the balance after delivery" which implies that "acceptance completion" is a precondition, meaning that "acceptance completion" is an implicit performance task.
[0047] As a possible implementation method, the first associated information associated with the target role may be extracted from the target document by keyword matching.
[0048] Taking the target document as a contract text as an example, the performance tasks of the after-sales technical support personnel of the second party can be extracted from the after-sales service contract of the software service privatization delivery, such as through the keywords "bug fixing", "security vulnerability", "upgrade", and the like.
[0049] As another possible implementation manner, the first association information associated with the target role can be extracted from the target document through a pre-trained extraction model.
[0050] It should be noted that the first association information can be one or more, which is not limited in the present application.
[0051] In step 103, the second association information associated with the target role is obtained from the target rule.
[0052] Since the association information associated with the role is not mentioned in the target document, there can be a requirement in the target rule, and based on this, in the present application, the second association information associated with the target role can be obtained from the target rule to supplement the first association information extracted from the target document with the second association information.
[0053] Taking the target document as a contract text as an example, since the performance tasks are not agreed in the contract text, there can be a requirement in the target rule, and in the present application, the performance tasks related to the target role can be obtained from the target rule to supplement the performance tasks extracted from the target document with the performance tasks extracted from the target rule.
[0054] For example, the target rule includes business specifications, the task list in the business specifications is stored in a JSON format, and contains a trigger condition (such as, contract type = procurement AND role = supplier → add "inspection application" task), the target role can be matched with the roles in each performance task in the task list, and then the performance tasks of the roles in the task list that match the target role are the performance tasks of the target role.
[0055] For another example, the target rule includes legal provisions, and then the performance tasks of the target role can be extracted from the legal provisions through a pre-trained extraction model.
[0056] It should be noted that the second association information of the target role can be one or more, which is not limited in the present application.
[0057] In step 104, the first association information is updated according to the second association information to obtain the target association information of the target role.
[0058] In the present application, the first association information can be supplemented and updated according to the second association information, and the updated first association information can be used as the target association information of the target role.
[0059] Taking the target document as a contract text as an example, the performance task of the target role in the target document can be supplemented and updated according to the performance task of the target role obtained from the target rule, and the updated performance task can be taken as the target performance task of the target role. Thus, the performance task in the target rule can be avoided to be missed, and the accuracy of the performance task extraction can be improved.
[0060] In the embodiments of the present application, the associated information of the target role extracted from the target document is updated to obtain the final associated information of the target role by using the associated information of the target role obtained from the target rule, so that the associated information in the target rule can be avoided to be missed, and the accuracy of the information extraction can be improved.
[0061] Figure 2 The flowchart of the information extraction method provided by another embodiment of the present application is shown.
[0062] As shown in Figure 2 , the information extraction method comprises:
[0063] Step 201, obtaining a target document and a target rule related to the target document.
[0064] In the present application, step 201 can adopt any one of the implementation manners of the embodiments of the present application, and therefore will not be repeated here.
[0065] Step 202, extracting first associated information associated with the target role from the target document.
[0066] In the present application, step 202 can adopt any one of the implementation manners of the embodiments of the present application, and therefore will not be repeated here.
[0067] In some embodiments, a large model can be used to extract the first associated information associated with the target role from the target document.
[0068] For example, the prompt template can be filled according to the target document to generate prompt information, and the large model can be used to perform associated information extraction processing on the prompt information to obtain the first associated information of the target role.
[0069] The large model can be obtained by fine-tuning an initial large model using an annotated target document. Here, the annotated target document can be annotated with associated information associated with the role. Taking the extraction of the performance task as an example, the annotated contract text is annotated with the performance task, including the explicit performance task and the implicit performance task.
[0070] The prompt template can include the associated information extraction requirements corresponding to the target role.
[0071] For example, taking the after-sales service contract for private delivery of software services as an example, the target role is Party B's after-sales technical support personnel. The corresponding performance task extraction requirements for this role may include: 1. Only output the complete original text directly related to "Party B's obligation to provide security vulnerability notification and repair"; 2. Keyword limitation: the clause contains at least one of "vulnerability notification", "vulnerability", "vulnerability repair", "security vulnerability", and "security management"; 3. Retain the complete sentences of the original text (such as clause number, scope of responsibility, etc.), etc.
[0072] Optionally, the prompt template may further include model output requirements, which may include output format requirements, output content requirements, etc.
[0073] Taking the extraction of fulfillment tasks as an example, the model output requirement includes outputting a structured task list in JSON format. The structured task list includes task content, responsible persons, deadlines, dependencies, etc. The responsible persons here can be understood as target roles.
[0074] For example, if the target role is a specified role, the prompt template used can be a prompt template selected from the prompt template library that matches the target role. If the target role is each role in the target document, then the prompt template can be a prompt template selected from the prompt template library that matches the target document and is used to extract the associated information of each role.
[0075] Therefore, by extracting relevant information from target documents based on the relevant information requirements corresponding to the target role and leveraging the semantic understanding capabilities of the large model, the accuracy of the extracted relevant information can be improved. For example, the large model can understand complex expressions in contract texts (such as mapping "quarterly performance tuning" to "performance optimization"), resolve missed detection issues such as ignoring synonyms, and improve task extraction accuracy.
[0076] Step 203: Acquire second association information associated with the target role from the target rule.
[0077] In the present application, step 203 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.
[0078] Step 204: Determine the semantic relevance between the second association information and the target document.
[0079] In the present application, for any second associated information, the second associated information and the target document can be vectorized respectively to obtain the vector of the second associated information and the vector of the target document, and the semantic association between the second associated information and the target document can be determined based on the vector of the second associated information and the vector of the target document.
[0080] In some embodiments, the first semantic similarity between the second association information and the target document can be determined according to a vector of the second association information and a vector of the target document, the rule matching degree between the second association information and the target role can be determined, and the semantic association degree between the second association information and the target document can be determined according to the first semantic similarity and the rule matching degree.
[0081] Therefore, when determining the semantic association degree between the second association information and the target document, not only the semantic similarity between the second association information and the target document is considered, but also the rule matching degree between the second association information and the target role is considered, so that the accuracy of the semantic association degree can be improved.
[0082] For example, the rule matching degree between the second association information and the target role can be determined as follows: whether the target role has an execution right on the second association information can be determined first, if the target role has an execution right on the second association information, the second semantic similarity between the second association information and the role permission of the target role can be determined, and the rule matching degree can be determined according to the second semantic similarity, if the target role does not have an execution right on the second association information, the rule matching degree can be determined as a preset value.
[0083] For example, the target document is a contract text, and the association information is a performance task, whether the target role has an execution right on the performance task obtained from the target rule can be determined first, if the target role has an execution right on the performance task, the second semantic similarity between the performance task and the role permission of the target role can be determined, and the rule matching degree can be determined according to the second semantic similarity.
[0084] For example, the rule matching degree can be a Boolean value of 0 or 1, if the target role has no execution right on the second association information at all, the rule matching degree can be determined as 0, if the target role has an execution right on the second association information, the second semantic similarity can be calculated.
[0085] For example, the role “financial personnel” cannot execute the task “construction acceptance”, so the rule matching degree between the role “financial personnel” and the task “construction acceptance” is 0. For example, the role “supplier” has the right to execute the task “delivery of goods”, so the semantic similarity between the role “supplier” and the task “delivery of goods” can be calculated.
[0086] For example, if the second semantic similarity is greater than a first threshold value, the rule matching degree can be determined as 1, if the second semantic similarity is less than or equal to the first threshold value, the rule matching degree can be determined as 0.
[0087] It should be noted that the first threshold value can be determined according to actual needs, which is not limited.
[0088] Exemplarily, the first semantic similarity and the rule matching degree may be weighted to obtain the semantic relevance.
[0089] As an example, the following formula (1) can be used to determine any second association information t i The semantic relevance S(t i ):
[0090] S(t i )=α·Sim LLM (t i ,C)+β·Match Rule (t i ,R) (1)
[0091] Among them, Sim LLM (t i , C) represents the second associated information t i Semantic similarity with the target document C; Match Rule (t i , R) represents the second associated information t i The rule matching degree with the target role R, for example, the rule matching degree can be a Boolean value of 0 or 1; α and β represent weights.
[0092] Therefore, when the target role has execution authority for the second associated information, the rule matching degree is determined based on the semantic similarity between the second associated information and the role authority of the target role, thereby improving the accuracy of the rule matching and further improving the accuracy of the semantic association degree.
[0093] Step 205: Determine third association information from the second association information according to the semantic association degree.
[0094] In the present application, the association information with a semantic association greater than a second threshold in the second association information may be used as the third association information.
[0095] For example, the second threshold is 0.7, and the second association information with a semantic association greater than 0.7 can be used as the third association information.
[0096] It should be noted that the second threshold can be determined according to actual needs and is not limited thereto.
[0097] Step 206: Update the first association information according to the third association information to obtain target association information.
[0098] In this application, the third association information can be used as the association information of the target role to supplement the first association information and obtain the target association information of the target role, that is, the association information not mentioned in the contract but included in the target rules can be supplemented.
[0099] In an embodiment of the present application, third related information is filtered out from the second related information based on the semantic relevance between the second related information and the target document, and the first related information is updated using the filtered third related information. Thus, the first related information is updated using the second related information that has a high semantic relevance to the target document, thereby improving the accuracy of the target related information.
[0100] Figure 3 A flowchart of an information extraction method provided in another embodiment of the present application.
[0101] like Figure 3 As shown, the information extraction method includes:
[0102] Step 301: Obtain a target document and target rules related to the target document.
[0103] Step 302: extract first association information associated with the target role from the target document.
[0104] Step 303: Acquire second association information associated with the target role from the target rule.
[0105] In the present application, steps 301 to 303 can be implemented in any of the embodiments of the present application, so they will not be described in detail here.
[0106] Step 304: Update the first association information according to the second association information to obtain updated first association information.
[0107] In the present application, the method described in the above embodiment can be used to filter out the third related information from the second related information, and use the third related information to update the first related information to obtain the updated first related information.
[0108] Taking the target document as a contract text as an example, one or more performance tasks can be screened out from the performance tasks obtained from the target rules, and the performance tasks extracted from the target document can be updated using the screened performance tasks to obtain updated performance tasks.
[0109] Step 305: Determine the role authority of the target role according to the target rule.
[0110] In the present application, the target rule may include a mapping relationship between roles and permissions, and the role permissions of the target role may be determined by querying the mapping relationship based on the target role.
[0111] Step 306: Filter the updated first association information according to the role authority to obtain target association information.
[0112] In this application, the role permissions of the target role can be used to filter the associated information in the updated first associated information that the role permissions do not have execution permissions to obtain the target associated information. This can solve the problems of information redundancy and permission leakage.
[0113] For example, hide the "Party A Internal Approval" task from the role "Supplier".
[0114] In an embodiment of the present application, by determining the role permissions of the target role according to the target rules and using the role permissions to filter the updated first associated information, the accuracy of information extraction can be improved. For example, in the scenario of extracting fulfillment tasks, using role permissions to filter the updated fulfillment tasks can not only ensure that the filtered fulfillment tasks are tasks that the target role has the authority to perform, but also generate differentiated task lists based on role permissions, thereby improving the accuracy of task extraction.
[0115] In the scenario of extracting performance tasks from the contract text, the target association information can be the target performance tasks, and there may be dependencies between performance tasks. In order to improve the efficiency of information storage, a task dependency graph can be constructed to store the performance tasks in the form of a task dependency graph. Figure 4 To explain, Figure 4 A flowchart of an information extraction method provided in another embodiment of the present application.
[0116] like Figure 4 As shown, the information extraction method may further include:
[0117] Step 401: Obtain attribute information of the target fulfillment task.
[0118] Among them, the attribute information may include the identification of the target fulfillment task, task type, estimated time, person in charge, task priority, task status and other attributes.
[0119] Taking contracts in the software delivery field as an example, the identifier of the performance task can be the task ID, the task type can include development, testing, documentation, etc., the estimated time can be in hours or days, the task priority can be high, medium, or low, and the task status can include not started, in progress, completed, etc.
[0120] Step 402: Obtain the node feature matrix of the target fulfillment task based on the attribute information.
[0121] In this application, a fulfillment task can be represented by a node, and the directed edges between nodes represent the dependency relationship between the fulfillment tasks. The attribute information of the edges between nodes may include dependency type, delay time, whether it is a strong dependency, etc.
[0122] Exemplarily, dependency types may include finish to start, start to start, etc., such as task A completed → task B started; delay time may refer to the delay after the dependency is triggered, such as task B delayed for 2 days after task A is completed; whether the dependency is forced can be a Boolean value, such as 0 for no forced dependency and 1 for forced dependency.
[0123] In this application, the attribute information of the target fulfillment task can be feature encoded to uniformly convert different attributes into numerical form to obtain the node feature matrix of the target fulfillment task.
[0124] Taking contracts in the software delivery field as an example, task types and task statuses are classified and encoded, the names of responsible persons are converted into vectors through word embedding models, and task priorities are mapped to numerical values, such as high, medium, and low are mapped to 1, 0.5, and 0.2 respectively.
[0125] For example, if there are multiple target fulfillment tasks, the node feature matrix is obtained by performing feature encoding on the attribute information of the multiple target fulfillment tasks respectively.
[0126] For example, the node feature matrix X∈R N×D , where N represents the number of nodes, i.e., the number of target fulfillment tasks, and D represents the feature dimension. A node feature corresponds to an attribute of a fulfillment task.
[0127] Step 403: Based on the node feature matrix of the target fulfillment tasks, a graph neural network is used to perform prediction to obtain the probability of existence of the first edge and the first dependency type between the target fulfillment tasks.
[0128] Graph neural networks can be used to predict the edge existence probability and dependency type between fulfillment tasks. The edge existence probability represents the possibility of a dependency relationship between fulfillment tasks. The greater the edge existence probability, the higher the possibility of a dependency relationship.
[0129] In this application, a graph neural network can be used to process the node feature matrix to obtain the node representation corresponding to each target fulfillment task, and the node representation of each target fulfillment task can be decoded to obtain the probability of the first edge existence and the first dependency type between the target fulfillment tasks.
[0130] The existence probability of the first edge may indicate the possibility of a dependency relationship between target fulfillment tasks, and the first dependency type may indicate the dependency type between target fulfillment tasks.
[0131] For example, if the probability of the first edge between the target fulfillment tasks is greater than the third threshold, it can be considered that there is a dependency relationship between the target fulfillment tasks. Based on the node representation of the target fulfillment tasks, the dependency type between the target fulfillment tasks is further predicted. If the probability of the first edge between the target fulfillment tasks is less than or equal to the third threshold, it can be considered that there is no dependency relationship between the target fulfillment tasks.
[0132] In some embodiments, the graph neural network can be trained in the following manner: the node feature matrix and edge feature matrix corresponding to the sample fulfillment task can be obtained, and the initial graph neural network can be used to process the node feature matrix and edge feature matrix of the sample fulfillment task to obtain the node feature representation of the sample fulfillment task, and the node feature representation of the sample fulfillment task is decoded to obtain the probability of the second edge existence and the second dependency type between the sample fulfillment tasks, and then the initial graph neural network is trained based on the second edge existence probability and the second dependency type to obtain the graph neural network.
[0133] For example, historical task data can be obtained from a task management system and processed to obtain sample fulfillment tasks. The dependencies between the sample fulfillment tasks can then be inferred based on the completion times of the fulfillment tasks in the historical task data. Feature encoding can be performed on the attribute information of the sample fulfillment tasks to obtain a node feature matrix for the sample fulfillment tasks. Feature encoding can also be performed on the attribute information of the dependencies between the sample fulfillment tasks to obtain an edge feature matrix. Here, one edge feature corresponds to one attribute of the edge.
[0134] For example, the edge feature matrix E∈R M×F , where M represents the number of edges and F represents the edge feature dimension.
[0135] Taking the initial neural network as the graph attention network as an example, the node feature representation of the sample fulfillment task output by the previous graph convolutional layer can be linearly transformed to obtain the transformed features. According to the transformed features corresponding to any sample fulfillment task and its adjacent sample fulfillment tasks, the attention weight of any sample fulfillment task and its adjacent sample fulfillment tasks can be obtained. According to the attention weight and the node feature representation output by the previous layer corresponding to any sample fulfillment task, the node feature representation of any sample fulfillment task output by the current graph convolutional layer is determined. In this way, the node feature representation of the sample fulfillment task output by the last graph convolutional layer can be obtained. By decoding the node feature representation of any two sample fulfillment tasks, the probability of the existence of the second edge and the second dependency type between any two sample fulfillment tasks can be obtained.
[0136] Exemplarily, if the probability of the existence of the second edge is greater than the third threshold, it is determined that the edge between the sample fulfillment tasks exists. For the existing edge, it is decoded according to the node feature representation of the sample fulfillment task output by the last graph convolution layer to obtain the second dependency type.
[0137] Exemplarily, the edge loss can be determined based on the probability of existence of the second edge corresponding to the sample fulfillment task and the actual edge existence probability, and the dependency loss can be determined based on the second dependency type and the actual dependency type. The model loss can be determined based on the edge loss and the dependency loss, and then the parameters of the initial graph neural network can be adjusted based on the model loss. The graph neural network with adjusted parameters can be continued to be trained until the training end conditions are met to obtain the graph neural network.
[0138] For example, if there is a dependency relationship between two sample fulfillment tasks, then the probability of the actual edge existing is 1; if there is no dependency relationship between the two sample fulfillment tasks, then the probability of the actual edge existing is 0.
[0139] Therefore, by using the node feature matrix and edge feature matrix corresponding to the sample fulfillment tasks, the initial graph neural network is trained so that the graph neural network can learn how to capture whether there is a dependency relationship between fulfillment tasks and the dependency type of the dependency relationship from the node feature matrix, thereby obtaining a graph neural network that can predict the probability of edge existence and dependency type between fulfillment tasks.
[0140] Step 404: construct a task dependency graph corresponding to the target role according to the first edge existence probability and the first dependency type.
[0141] In this application, valid edges can be screened out based on the probability of existence of the first edge, and then a task dependency graph corresponding to the target role can be constructed based on the valid edges and the first dependency type corresponding to the valid edges, so that the target role's fulfillment tasks can be stored in the form of a task dependency graph.
[0142] Exemplarily, if the existence probability of the first edge between the target fulfillment tasks is greater than a third threshold, it can be determined that the edge between the target fulfillment tasks is a valid edge.
[0143] In an embodiment of the present application, a graph neural network can also be used to construct a task dependency graph corresponding to the target role based on the attribute information of the target fulfillment task. By storing the role's fulfillment tasks through the task dependency graph, storage space can be compressed and storage efficiency can be improved.
[0144] As service demands change, the role's performance tasks may increase. In one embodiment of the present application, based on the above embodiment, if the target role has newly added performance tasks, a graph neural network can be used to obtain the probability of the existence of the third edge and the third dependency type between the newly added performance tasks of the target role and the performance tasks in the target role's task dependency graph, and update the task dependency graph based on the probability of the existence of the third edge and the third dependency type to obtain an updated task dependency graph.
[0145] The method for obtaining the third edge existence probability and the third dependency type can refer to the method for obtaining the first edge existence probability and the first dependency type in the above embodiment, so it will not be repeated here.
[0146] For example, valid edges between newly added fulfillment tasks and fulfillment tasks in the task dependency graph can be screened based on the probability of the third edge existing. The task dependency graph can then be updated based on the valid edges and the third dependency types corresponding to the valid edges to obtain an updated task dependency graph. The method for screening valid edges can be found in the description of the above embodiment and will not be repeated here.
[0147] For example, if the edge between a newly added fulfillment task and a fulfillment task in the task dependency graph is a valid edge, then a node for the newly added fulfillment task can be added to the task dependency graph, and an edge can be added between the node of the newly added fulfillment task and the node of the fulfillment task, and the direction of the edge can be determined based on the dependency type between the newly added fulfillment task and the fulfillment task.
[0148] Therefore, when the target role has new fulfillment tasks, a graph neural network can be used to incrementally update the task dependency graph. Compared with reconstructing the task dependency image based on the new fulfillment tasks and the existing fulfillment tasks in the task dependency graph, it can reduce the amount of calculation and improve the updating efficiency of the task dependency graph.
[0149] Since the association information of the same role in the target document and the target rule may not be inconsistent, for example, the contract is signed after negotiation between the signatories, then the performance tasks of a role in the contract text may be inconsistent with the performance tasks of the role in the target rule. Based on this, in one embodiment of the present application, a conflict detection is performed on the first association information and the target rule to determine the difference information between the first association information and the target rule.
[0150] The difference information may refer to the difference between the first association information and the target rule.
[0151] Exemplarily, the first association information may be compared with the target rule one by one to determine difference information.
[0152] Taking the target document as a contract text as an example, the types of difference information may include time contradictions, process reversals, role overstepping, content violations, etc.
[0153] For example, there are time conflicts, for example, the contract requires "payment within 3 days", but the company stipulates "payment within at least 5 days"; the process is reversed, for example, the contract requires "payment first and then inspection of goods", but the industry standard requires "inspection first and then payment"; role overstepping, for example, the contract requires "salesperson signs the acceptance form", but the company stipulates that "quality inspection department must sign"; content violations, for example, the contract states "no safety inspection is required", but the law stipulates that "inspection is required".
[0154] Therefore, by performing conflict detection on the first association information and the target rule and determining the difference information between the two, contract performance disputes can be reduced.
[0155] In one embodiment of the present application, after determining the difference information between the first association information and the target rule, a conflict detection report can be generated based on the difference information, or a modification suggestion for the first association information can be generated based on the difference information, or the first association information in the target document can be highlighted based on the difference information to obtain the processed target document.
[0156] The conflict detection report may include but is not limited to difference information, conflict detection time, location of the first associated information in the target document, location of clauses in the target rule that differ from the first associated information, etc.
[0157] For example, based on the performance tasks of a role in the target rule in the difference information, a modification suggestion for the performance tasks of that role in the contract text can be generated. For example, if the contract requires "payment within 3 days", but the company stipulates "payment within at least 5 days", the modification suggestion can be "payment time should be ≥ 5 days".
[0158] Exemplarily, the highlighting process may include but is not limited to one or more of the following: highlighting, marking the first fulfillment task with a mark box, underlining the first fulfillment task, etc.
[0159] It should be noted that one or more of generating a conflict detection report, generating modification suggestions, and highlighting processing may be performed, and this is not limited.
[0160] Therefore, based on the difference information, a conflict detection report is generated, which can facilitate obtaining the conflict between the contract terms and the target rules based on the conflict detection report; generating modification suggestions for related information can facilitate users to execute related information based on the modification suggestions to avoid performance disputes; the first related information in the target document is highlighted to facilitate users to pay attention to contract terms that are likely to cause performance disputes.
[0161] In one embodiment of the present application, after obtaining the target-related information of the target role, the target-related information can also be pushed to the information management system, which manages the target role's related information, supports information status writeback, realizes a closed loop of document management and information execution, and avoids information islands.
[0162] Taking the extracted related information as the performance task as an example, the target performance task is pushed to the task management system, and the task management system manages the performance task of the target role, supports status writeback, realizes the closed loop of contract management and task execution, and realizes two-way linkage between the task management system and the contract performance progress (for example, the task "completed" automatically updates the contract performance dashboard), avoids information silos, and reduces communication costs.
[0163] For example, a role can view the fulfillment tasks that need to be performed and update the status of the fulfillment tasks through the task management system. For example, when a role performs a certain associated information, the status of the associated information can be updated on the user interface, so that the fulfillment management system can simultaneously update the status of the stored associated information.
[0164] Since some rights of a role can only be executed after other rights are executed, or can be executed at the same time, it can be understood that in the scenario of extracting the executable rights of a role from a document, the above-mentioned method of constructing a task dependency graph and the method of updating a task dependency graph are also applicable to the construction of a rights dependency graph and the updating of a task dependency graph, etc., and will not be repeated here.
[0165] In order to implement the above embodiment, the embodiment of the present application also proposes an information extraction device. Figure 5 A schematic diagram of the structure of an information extraction device provided in one embodiment of the present application.
[0166] like Figure 5 As shown, the information extraction device 500 includes:
[0167] A first acquisition module 510 is configured to acquire a target document and target rules related to the target document;
[0168] An extraction module 520 is configured to extract first association information associated with a target role from the target document;
[0169] A second acquisition module 530 is configured to acquire second association information associated with the target role from the target rule;
[0170] The first updating module 540 is configured to update the first association information according to the second association information to obtain target association information of the target role.
[0171] Optionally, the first updating module 540 is configured to:
[0172] Determining a semantic relevance between the second association information and the target document;
[0173] determining third association information from the second association information according to the semantic association degree;
[0174] The first association information is updated according to the third association information to obtain the target association information.
[0175] Optionally, the first updating module 540 is configured to:
[0176] Determining a first semantic similarity between the second association information and the target document;
[0177] Determining a rule matching degree between the second association information and the target role;
[0178] The semantic relevance is determined according to the first semantic similarity and the rule matching degree.
[0179] Optionally, the first updating module 540 is configured to:
[0180] In response to the target role having execution authority for the second association information, determining a second semantic similarity between the second association information and the role authority of the target role;
[0181] The rule matching degree is determined according to the second semantic similarity.
[0182] Optionally, the first updating module 540 is configured to:
[0183] updating the first association information according to the second association information to obtain updated first association information;
[0184] Determine the role permissions of the target role according to the target rule;
[0185] The updated first association information is filtered according to the role authority to obtain the target association information.
[0186] Optionally, the extraction module 520 is configured to:
[0187] Filling a prompt template according to the target document to generate prompt information; wherein the prompt template includes the associated information extraction requirements corresponding to the target role;
[0188] A large model is used to perform associated information extraction processing on the prompt information to obtain the first associated information.
[0189] Optionally, the target association information is a target fulfillment task, and the device may further include:
[0190] A second acquisition module is used to obtain attribute information of the target fulfillment task;
[0191] A third acquisition module is used to obtain a node feature matrix of the target fulfillment task based on the attribute information;
[0192] A prediction module, configured to perform prediction using a graph neural network based on a node feature matrix of the target fulfillment tasks to obtain a first edge existence probability and a first dependency type between the target fulfillment tasks;
[0193] A construction module is used to construct a task dependency graph corresponding to the target role according to the existence probability of the first edge and the first dependency type.
[0194] Optionally, the graph neural network is trained using the following method:
[0195] Obtain the node feature matrix and edge feature matrix corresponding to the sample fulfillment task;
[0196] Using an initial graph neural network, the node feature matrix and the edge feature matrix of the sample fulfillment task are processed to obtain a node feature representation of the sample fulfillment task;
[0197] Decoding the node feature representations of the sample fulfillment tasks to obtain the second edge existence probability and the second dependency type between the sample fulfillment tasks;
[0198] The initial graph neural network is trained according to the second edge existence probability and the second dependency type to obtain the graph neural network.
[0199] Optionally, the device may further include:
[0200] A fourth acquisition module is used to acquire the newly added fulfillment tasks of the target role;
[0201] The prediction module is further configured to use the graph neural network to obtain a probability of existence of a third edge and a third dependency type between the newly added fulfillment task and the fulfillment task in the task dependency graph;
[0202] The second updating module is configured to update the task dependency graph according to the third edge existence probability and the third dependency type to obtain an updated task dependency graph.
[0203] Optionally, the device may further include:
[0204] The conflict detection module is configured to perform conflict detection on the first association information and the target rule to determine difference information between the first association information and the target rule.
[0205] Optionally, the device may further include at least one of the following modules:
[0206] A first generating module, configured to generate a conflict detection report based on the difference information;
[0207] a second generating module, configured to generate a modification suggestion for the first associated information based on the difference information;
[0208] A highlighting processing module is used to perform highlighting processing on the first associated information in the target document according to the difference information to obtain a processed target document.
[0209] It should be noted that the explanation of the aforementioned information extraction method embodiment is also applicable to the information extraction device of this embodiment, and therefore will not be repeated here.
[0210] In an embodiment of the present application, by utilizing the association information of the target role obtained from the target rule, the association information of the target role extracted from the target document is updated to obtain the final association information of the target role, thereby avoiding missing the association information in the target rule and improving the accuracy of association information extraction.
[0211] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.
[0212] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0213] like Figure 6As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.
[0214] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0215] The computing unit 601 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the information extraction method. For example, in some embodiments, the information extraction method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the information extraction method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the information extraction method in any other appropriate manner (eg, by means of firmware).
[0216] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0217] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0218] In the context of this application, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electronic storage, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0219] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0220] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front-end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.
[0221] The computer system can include clients and servers. This relationship can be between two computers, or between computers and servers located throughout the network, depending on the context in which the term is used. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services (Virtual Private Server, Virtual Private Server). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0222] According to the embodiments of the present application, the present application also provides a computer program product, when the instruction processor in the computer program product executes, executes the information extraction method proposed in the above embodiments of the present application.
[0223] It should be understood that the steps shown above can be reordered, added or deleted using various forms of flow. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, which is not limited herein.
[0224] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent replacement and improvement within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An information extraction method comprising: Obtaining a target document and target rules related to the target document; extracting first association information associated with the target role from the target document; Acquire second association information associated with the target role from the target rule; The first association information is updated according to the second association information to obtain target association information of the target role.
2. The method according to claim 1, wherein The updating of the first association information according to the second association information to obtain target association information of the target role includes: Determining a semantic relevance between the second association information and the target document; determining third association information from the second association information according to the semantic association degree; The first association information is updated according to the third association information to obtain the target association information.
3. The method according to claim 2, wherein: The determining of the semantic association between the second association information and the target document includes: Determining a first semantic similarity between the second association information and the target document; Determining a rule matching degree between the second association information and the target role; The semantic relevance is determined according to the first semantic similarity and the rule matching degree.
4. The method according to claim 3, wherein: The determining of a rule matching degree between the second association information and the target role includes: In response to the target role having execution authority for the second association information, determining a second semantic similarity between the second association information and the role authority of the target role; The rule matching degree is determined according to the second semantic similarity.
5. The method according to claim 1, wherein The updating of the first association information according to the second association information to obtain target association information of the target role includes: updating the first association information according to the second association information to obtain updated first association information; Determine the role permissions of the target role according to the target rule; The updated first association information is filtered according to the role authority to obtain the target association information.
6. The method of claim 1, wherein: The step of extracting the first associated information of the target role from the target document includes: Filling a prompt template according to the target document to generate prompt information; wherein the prompt template includes the associated information extraction requirements corresponding to the target role; A large model is used to perform associated information extraction processing on the prompt information to obtain the first associated information.
7. The method according to any one of claims 1 to 6, wherein The target-related information is a target fulfillment task, and the method further includes: Obtaining attribute information of the target fulfillment task; Obtaining a node feature matrix of the target fulfillment task based on the attribute information; According to the node feature matrix of the target fulfillment tasks, a graph neural network is used to perform prediction to obtain the first edge existence probability and the first dependency type between the target fulfillment tasks; A task dependency graph corresponding to the target role is constructed based on the first edge existence probability and the first dependency type.
8. The method of claim 7, wherein: The graph neural network is trained using the following method: Obtain the node feature matrix and edge feature matrix corresponding to the sample fulfillment task; Using an initial graph neural network, the node feature matrix and the edge feature matrix of the sample fulfillment task are processed to obtain a node feature representation of the sample fulfillment task; Decoding the node feature representations of the sample fulfillment tasks to obtain a second edge existence probability and a second dependency type between the sample fulfillment tasks; The initial graph neural network is trained according to the second edge existence probability and the second dependency type to obtain the graph neural network.
9. The method of claim 7, further comprising: Obtaining newly added fulfillment tasks for the target role; Using the graph neural network, obtaining a probability of existence of a third edge and a third dependency type between the newly added fulfillment task and the fulfillment tasks in the task dependency graph; The task dependency graph is updated according to the third edge existence probability and the third dependency type to obtain an updated task dependency graph.
10. The method according to any one of claims 1 to 6, further comprising: A conflict detection is performed on the first association information and the target rule to determine difference information between the first association information and the target rule.
11. The method of claim 10, further comprising at least one of the following: generating a conflict detection report based on the difference information; generating a modification suggestion for the first associated information based on the difference information; According to the difference information, the first associated information in the target document is highlighted to obtain a processed target document.
12. An information extraction device comprising: A first acquisition module is used to acquire a target document and target rules related to the target document; An extraction module, configured to extract first association information associated with a target role from the target document; A second acquisition module, configured to acquire second association information associated with the target role from the target rule; An updating module is configured to update the first association information according to the second association information to obtain target association information of the target role.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.