Contract information extraction method and device, electronic equipment and storage medium

By enabling user-driven annotation and feedback integration, the method optimizes contract extraction models to reduce training costs and improve adaptability and accuracy, addressing the limitations of existing models.

CN120317221APending Publication Date: 2025-07-15QUNJE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510493480.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing contract information extraction models face high training costs and are difficult to adapt to individual user needs, limiting their performance improvement due to the need for frequent retraining with large datasets and inability to incorporate user feedback.

Method used

A method that allows users to annotate contract documents directly, generating user-specific data for model optimization, which is then used to refine the contract extraction model, incorporating user feedback and performance evaluation to ensure continuous improvement.

Benefits of technology

This approach reduces training data dependency, shortens model update cycles, and enhances model adaptability and accuracy by aligning with user-specific needs, ensuring high-quality contract information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317221A_ABST
    Figure CN120317221A_ABST
Patent Text Reader

Abstract

The invention provides a contract information extraction method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the steps that a contract document is labeled, the contract document added with user labeling data is obtained, and the user labeling data comprises elements selected by a user, labels of the elements selected by the user and position information of the elements selected by the user in the contract document; and according to the contract document added with the user annotation data, optimizing the contract extraction model to obtain an optimized contract extraction model for performing information extraction on the contract document. According to the method, by introducing a user labeling and feedback mechanism, the user is allowed to directly label the contract document, and the contract extraction model is optimized by using the contract document added with the user labeling data, so that the contract extraction model can be adjusted according to the personalized demand of the user, and the user experience is improved. The adaptability and accuracy of the optimized contract extraction model are improved, and the pertinence of the extraction result is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, electronic device, and storage medium for extracting contract information. Background Art

[0002] With the acceleration of digital transformation, the demand for automation in the contract management process is increasing day by day, which has promoted the development of contract information extraction technology based on deep learning. Currently, contract information extraction mainly relies on pre-trained deep learning models, such as pre-trained models (Bidirectional Encoder Representations from Transformers, BERT), pre-trained models (Robustly Optimized BERT Approach, RoBERTa), etc. These models can efficiently identify and extract key information such as contract parties, contract amounts, and contract terms from contract texts through training on a large amount of training data.

[0003] However, in order to maintain the accuracy and effectiveness of the model, it is usually necessary to update the model regularly, which often requires re-collecting a large amount of training data and consuming resources for model retraining, thereby increasing the model training cost. In addition, the above models are mostly general-purpose models, which are difficult to meet the personalized needs of different users, and cannot optimize and improve the model in a timely manner according to user feedback, thus limiting the continuous improvement of model performance. Summary of the Invention

[0004] The purpose of the present application is to provide a method, apparatus, electronic device, and storage medium for extracting contract information to solve the problems in the prior art that the model training cost is relatively high, it is difficult to meet the personalized needs of different users, and the model cannot be optimized and improved in a timely manner according to user feedback, thereby limiting the continuous improvement of model performance.

[0005] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for extracting contract information, the method including:

[0007] Obtain a contract document, where the contract document includes a plurality of clause paragraphs;

[0008] Perform annotation processing on the contract document to obtain a contract document with user annotation data added, where the user annotation data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document;

[0009] Optimize the contract extraction model according to the contract document with added user annotation data to obtain an optimized contract extraction model, which is used to extract information from the contract document to obtain the target extraction result.

[0010] As a possible implementation, the process of performing annotation processing on the contract document to obtain a contract document with added user annotation data includes:

[0011] Respond to the selection operation on the elements in the contract document, and display an optional label menu in the preset area where the user-selected element is located;

[0012] Respond to the touch operation on the optional labels in the optional label menu, and assign a label to the user-selected element, where the label is used to indicate the element category of the user-selected element;

[0013] Obtain the position information of the user-selected element in the contract document, and use the position information, the user-selected element, and the label of the user-selected element as the user annotation data to obtain a contract document with added user annotation data, where the position information includes the paragraph index of the clause paragraph to which the user-selected element belongs, and the start position and end position of the user-selected element in the clause paragraph.

[0014] As a possible implementation, before responding to the selection operation on the elements in the contract document and displaying an optional label menu in the preset area where the user-selected element is located, the method further includes:

[0015] Use an identification algorithm to identify the annotatable elements in the contract document, and generate each annotatable element into an optional control.

[0016] As a possible implementation, the process of optimizing the contract extraction model according to the contract document with added user annotation data to obtain an optimized contract extraction model includes:

[0017] Perform model training on the contract extraction model based on the contract document with added user annotation data to obtain a trained contract extraction model;

[0018] Evaluate the performance of the trained contract extraction model, determine the performance index value of the trained contract extraction model, and compare the performance index value of the trained contract extraction model with the performance index value of the contract extraction model;

[0019] If the performance metric value of the trained contract extraction model is greater than the performance metric value of the contract extraction model, then use the trained contract extraction model as the optimized contract extraction model and perform model update. Otherwise, re-annotate the contract document to continuously optimize the trained contract extraction model.

[0020] As a possible implementation, the re-annotating the contract document to continuously optimize the trained contract extraction model includes:

[0021] In response to correcting the user annotation data and incrementally annotating the contract document, new user annotation data is obtained;

[0022] Optimize the trained contract extraction model according to the new user annotation data until the performance metric value of the optimized contract extraction model is greater than the performance metric value of the trained contract extraction model.

[0023] As a possible implementation, after optimizing the contract extraction model according to the contract document with added user annotation data to obtain the optimized contract extraction model, it further includes:

[0024] Perform information extraction on the contract document based on the optimized contract extraction model to obtain an initial extraction result, where the initial extraction result includes multiple key entities;

[0025] Arrange the initial extraction result according to the position information of each key entity in the contract document to obtain the target extraction result.

[0026] As a possible implementation, the arranging the initial extraction result according to the position information of each key entity in the contract document to obtain the target extraction result includes:

[0027] Determine the character offset of each key entity according to the start position and end position of each key entity in the contract document;

[0028] Determine the context distance of each key entity according to the character offset of each key entity and the position index of the preset key anchor point;

[0029] Arrange and combine each key entity according to the context distance of each key entity to obtain the target extraction result.

[0030] In a second aspect, an embodiment of the present application provides a contract information extraction device, and the device includes:

[0031] An acquisition module, configured to acquire a contract document, where the contract document includes multiple clause paragraphs;

[0032] A marking module for marking the contract document to obtain a contract document with user marking data added, where the user marking data includes the elements selected by the user, the labels of the elements selected by the user, and the position information of the elements selected by the user in the contract document;

[0033] A model optimization module for optimizing the contract extraction model according to the contract document with user marking data added to obtain an optimized contract extraction model, where the optimized contract extraction model is used to extract information from the contract document to obtain a target extraction result.

[0034] As a possible implementation, the marking module is specifically used for:

[0035] Responding to the selection operation on the elements in the contract document, and displaying an optional label menu in a preset area where the user-selected elements are located;

[0036] Responding to the touch operation on the optional labels in the optional label menu, and assigning labels to the user-selected elements, where the labels are used to indicate the element categories of the user-selected elements;

[0037] Obtaining the position information of the user-selected elements in the contract document, and using the position information, the user-selected elements, and the labels of the user-selected elements as the user marking data to obtain a contract document with user marking data added, where the position information includes the paragraph index of the clause paragraph to which the user-selected elements belong and the start position and end position of the user-selected elements in the clause paragraph to which they belong.

[0038] As a possible implementation, the marking module is further used for:

[0039] Using an identification algorithm to identify the markable elements in the contract document, and generating each of the markable elements into an optional control.

[0040] As a possible implementation, the model optimization module is specifically used for:

[0041] Training the contract extraction model based on the contract document with user marking data added to obtain a trained contract extraction model;

[0042] Evaluating the performance of the trained contract extraction model, determining the performance index value of the trained contract extraction model, and comparing the performance index value of the trained contract extraction model with the performance index value of the contract extraction model;

[0043] If the performance metric value of the trained contract extraction model is greater than the performance metric value of the contract extraction model, then use the trained contract extraction model as the optimized contract extraction model and perform model update. Otherwise, re-annotate the contract document to continuously optimize the trained contract extraction model.

[0044] As a possible implementation, the model optimization module is specifically configured to:

[0045] In response to correcting the user annotation data and incrementally annotating the contract document, obtain new user annotation data;

[0046] Optimize the trained contract extraction model according to the new user annotation data until the performance metric value of the optimized contract extraction model is greater than the performance metric value of the trained contract extraction model.

[0047] As a possible implementation, the model optimization module is further configured to:

[0048] Perform information extraction on the contract document based on the optimized contract extraction model to obtain an initial extraction result, where the initial extraction result includes multiple key entities;

[0049] Arrange the initial extraction result according to the position information of each key entity in the contract document to obtain the target extraction result.

[0050] As a possible implementation, the model optimization module is further configured to:

[0051] Determine the character offset of each key entity according to the start position and end position of each key entity in the contract document;

[0052] Determine the context distance of each key entity according to the character offset of each key entity and the position index of a preset key anchor point;

[0053] Arrange and combine each key entity according to the context distance of each key entity to obtain the target extraction result.

[0054] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the contract information extraction method according to any one of the first aspects described above.

[0055] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the contract information extraction method according to any one of the above first aspects.

[0056] In the contract information extraction method provided by the present application, by obtaining a contract document including multiple clause paragraphs, performing annotation processing on the contract document to obtain a contract document with user annotation data added, where the user annotation data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document. According to the contract document with user annotation data added, optimizing the contract extraction model to obtain an optimized contract extraction model, and this optimized contract extraction model is used to extract information from the contract document to obtain a target extraction result. Specifically, it realizes that the user is allowed to directly perform annotation processing on the contract document to obtain a contract document with user annotation data added, and the user annotation data can be directly used for model retraining, reducing the dependence on model training data annotation, and effectively shortening the model update cycle. And because the user annotation data reflects the personalized needs of users in the actual business scenario, using the contract document with user annotation data added to optimize the contract extraction model enables the contract extraction model to be adjusted according to the personalized needs of users, while improving the adaptability and accuracy of the optimized contract extraction model, and also enhancing the pertinence of the extraction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 It shows a schematic flow chart of a contract information extraction method provided by an embodiment of the present application;

[0059] Figure 2 It shows a schematic flow chart of an annotation processing method provided by an embodiment of the present application;

[0060] Figure 3 It shows a schematic interface diagram of an annotation processing provided by an embodiment of the present application;

[0061] Figure 4 It shows a schematic flow chart of a contract extraction model optimization method provided by an embodiment of the present application;

[0062] Figure 5The flowchart shows another method for extracting contract information provided by an embodiment of the present application;

[0063] Figure 6 The flowchart shows a method for iteratively optimizing a contract extraction model based on user feedback provided by an embodiment of the present application;

[0064] Figure 7 The structural diagram shows a contract information extraction device provided by an embodiment of the present application;

[0065] Figure 8 The structural diagram shows an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application show operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.

[0067] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0068] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the subsequently stated features, but does not exclude the addition of other features.

[0069] In view of the problems existing in the prior art, the present application introduces a user annotation and feedback mechanism, and realizes the continuous optimization and improvement of the contract extraction model by combining model retraining and performance evaluation. Moreover, by continuously iteratively optimizing the contract extraction model based on user feedback, while reducing the model update cost, the adaptability and accuracy of the contract extraction model are improved. Furthermore, the optimized and improved contract extraction model is used to extract information from the contract document, ensuring the accuracy of the finally output extraction result.

[0070] Figure 1 FIG. shows a schematic flow chart of a contract information extraction method provided by an embodiment of the present application. Referring to Figure 1 as shown, the method specifically includes the following steps:

[0071] S101. Obtain a contract document, where the contract document includes a plurality of clause paragraphs.

[0072] Optionally, the user uploads one or more contract documents as samples through the front-end page. The contract document is composed of a plurality of clause paragraphs, and each clause paragraph contains different contract details, such as important information like contract parties, amounts, dates, etc. Among them, each clause paragraph is a logical paragraph. After the user uploads the contract document, the Natural Language Processing Engine (NLP engine) can disassemble the contract document into a plurality of logical paragraphs. This logical paragraph is different from the ordinary paragraph divided according to formatting elements such as punctuation marks, indents, and line spacing. The logical paragraph focuses on the logical structure and semantic association of the content, rather than simply format features. By disassembling the contract document into a plurality of logical paragraphs, the subsequent contract extraction model can more accurately extract key information by identifying the logical paragraphs, ensuring that the contract extraction model can correctly understand the true intention of the contract document.

[0073] S102. Perform annotation processing on the contract document to obtain a contract document with user annotation data added.

[0074] Optionally, the user annotation data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document. Among them, the user-selected elements refer to the key information fragments selected by the user from the contract document, such as specific amount, date, or participant name and other element information. The label of the user-selected element refers to a predefined label assigned to the user-selected element to identify the type of the user-selected element. The position information of the user-selected element in the contract document includes the paragraph index of the clause paragraph to which the user-selected element belongs, as well as the start position and end position of the user-selected element in the clause paragraph to which it belongs.

[0075] Optionally, the process of annotating a contract document is to convert unstructured text into structured information. The user interacts with the contract document through the front-end page. The user can select some elements in the contract document by dragging the mouse or other means, and record the position information of the selected elements in the contract document at this time, so as to determine an area for displaying an optional label menu or a pop-up window for display annotation options based on this position information, so as to dynamically display the optional label menu or the pop-up window for display annotation options in this area. Further, based on the predefined labels listed in the optional label menu or the pop-up window for display annotation options, the user can assign appropriate labels to the selected text according to the actual semantics, that is, mark the selected elements with predefined labels, and record the user annotation data to obtain a contract document with user annotation data added.

[0076] S103. Optimize the contract extraction model according to the contract document with user annotation data added to obtain an optimized contract extraction model, and the optimized contract extraction model is used to extract information from the contract document to obtain a target extraction result.

[0077] Optionally, use the contract document with user annotation data added to optimize the contract extraction model, that is, use the user annotation data to train the contract extraction model to adjust the model parameters of the contract extraction model, improve the accuracy and recall rate of the contract extraction model, so that the optimized contract extraction model can more accurately extract the required contract information from the contract document.

[0078] Optionally, convert the contract document with user annotation data added into an input format acceptable to the contract extraction model (such as JSON, CSV, etc.), and divide the user annotation data into a training set, a validation set and a test set, which are used for model training and evaluation respectively. Then select the model architecture of the contract extraction model (such as pre-trained models like BERT, RoBERTa, etc.), use the training set to iteratively train the contract extraction model, and continuously adjust the model parameters to improve the model prediction accuracy to obtain an optimized contract extraction model. Further, use the validation set to evaluate the performance of the optimized contract extraction model to prevent overfitting, and use the test set to evaluate the final performance of the optimized contract extraction model, such as indicators like accuracy and recall rate.

[0079] Optionally, the contract extraction model is any version of the contract extraction model that has been trained and put into use. In this application, the contract document with user annotation data added is used to iteratively optimize any version of the contract extraction model, and the accuracy and recall rate of the optimized contract extraction model are evaluated after the optimization is completed. Only when the extraction effect of the optimized contract extraction model is better than any version of the contract extraction model before optimization, the model is put online for replacement.

[0080] Based on this, the contract information extraction method according to the embodiments of the present application allows users to directly perform annotation processing on contract documents, obtaining contract documents with user annotation data added, and the user annotation data can be directly used for model retraining, reducing the dependence on model training data annotation and effectively shortening the model update cycle. Moreover, since the user annotation data reflects the personalized needs of users in the actual business scenario, using the contract documents with user annotation data added to optimize the contract extraction model enables the contract extraction model to be adjusted according to the personalized needs of users, improving the adaptability and accuracy of the optimized contract extraction model while also enhancing the pertinence of the extraction results.

[0081] Figure 2 shows a schematic flowchart of a marking processing method provided by an embodiment of the present application. Refer to Figure 2 As shown, the above step S102 performs marking processing on the contract document to obtain a contract document with user marking data added, specifically including the following steps:

[0082] S201. In response to a selection operation on an element in the contract document, display an optional label menu in a preset area where the user selects the element.

[0083] Optionally, in order to achieve the dynamic display of the optional label menu and ensure its position is always correct, the Selection API interface of JavaScript can be used to obtain the text range selected by the user, then find the last node (DOM element) in the selection range, and call the getBoundingClientRect() method of this node to obtain its position information relative to the viewport. And since the position information returned by the getBoundingClientRect() method is relative to the viewport, even if the page scrolls, this position is relative to the current visible area. However, when the page scrolls, the position needs to be recalculated to ensure the accuracy of the display position of the optional label menu. Further, according to the position information of the selection range, set the style of the optional label menu and position it to ensure that the optional label menu is positioned relative to the viewport and is not affected by scrolling. It should be noted that when the page scrolls, the position of the selection area changes relative to the viewport, so the position of the optional label menu needs to be updated in the scroll event.

[0084] Exemplarily, refer to Figure 3As shown, on the front-end page, the user can select specific text in the contract document via a mouse or touch device, such as "50% of the contract price, i.e., RMB Thirty-four thousand one hundred yuan in full (¥: 34,100 yuan)". After the user completes this selection, in response to the user's selection operation on the elements in the contract document, an optional label menu pops up in the preset area where the user selects the element. This optional label menu may include the following label options: Party A, Party B, Amount, Deposit Amount, Down Payment Amount, Total Amount. These labels are used to indicate the category of the element selected by the user.

[0085] Optionally, before displaying the optional label menu in the preset area where the user selects the element, it further includes: using an identification algorithm to identify the labelable elements in the contract document and generating each labelable element into an optional control.

[0086] Exemplarily, to help the user perform the labeling work more efficiently and reduce the workload of the user to select important information from a large amount of text, the present application provides a direct interaction method to quickly complete the labeling task. For example, before displaying the optional label menu in the preset area where the user selects the element, an identification algorithm is used to pre-identify the labelable elements in the contract document and generate these elements into optional controls.

[0087] Specifically, first, text preprocessing is performed on the contract document, such as word segmentation, stop word removal, etc., to better analyze and understand the text content of the contract document. Then, natural language processing technology NLP, especially named entity recognition technology (NER), is used to identify the key entities or elements in the contract document, including but not limited to contract parties (such as Party A, Party B), amounts (such as deposit amount, total amount), dates (such as signing date, payment date), contact information (such as phone number, email address), etc. Further, relationship extraction technology can also be used to identify the relationships between different entities. For example, what is the "contact information" of "a certain party", or whether a certain amount is used as "deposit" or "balance", etc.

[0088] Exemplarily, once the labelable elements in the contract document are identified, the labelable elements are converted into interface elements that are easy for the user to operate, that is, optional controls. For example, the identified elements are highlighted in different colors or styles in the contract document, or an underline or border is added to each identified element to make it more prominent, and a click area is created for each element. When the user hovers or clicks on this area, a corresponding label menu will pop up for the user to select the appropriate label.

[0089] Exemplarily, a tag library can be defined and constructed first, which is used to identify the categories of key information that may be involved in the contract document, such as Amount - Deposit Amount, Amount - Receipt and Payment Amount, Contract Party - Party A, Contract Party - Party B, Date - Signing Date, etc. These tags are predefined to provide a set of standardized classification systems for users. On this basis, when the user hovers the mouse over these highlighted elements, a small menu will automatically pop up, listing the tag options applicable to the selected element (such as "Contract Party", "Amount", etc.). The user only needs to click on the corresponding tag to complete the annotation of the element, without manually selecting the text range. And since the key elements that are most likely to need annotation have been pre - identified, not only the annotation efficiency can be improved during annotation, but also the consistency and accuracy of the annotation are ensured. Especially for some complex or long - length contract documents, this method significantly reduces the burden on users and makes the entire annotation process more intuitive and efficient.

[0090] S202. In response to a touch operation on an optional tag in the optional tag menu, assign a tag to the user - selected element.

[0091] Exemplarily, the tag is used to indicate the element category of the user - selected element. Continuing to refer to Figure 3 As shown, if the user selects the "Deposit Amount" tag from the popped - up optional tag menu, the "Deposit Amount" tag is associated with the user - selected element "50% of the contract price, i.e., RMB Thirty - four thousand one hundred yuan (¥: 34,100 yuan)", indicating that the text "50% of the contract price, i.e., RMB Thirty - four thousand one hundred yuan (¥: 34,100 yuan)" belongs to the "Deposit Amount" category.

[0092] S203. Obtain the position information of the user - selected element in the contract document, and use the position information, the user - selected element, and the tag of the user - selected element as user annotation data to obtain a contract document with user annotation data added.

[0093] Exemplarily, continuing to refer to Figure 3As shown, the element selected by the user is "50% of the contract price, i.e., RMB Thirty-four thousand one hundred yuan (¥: 34,100 yuan)". The label of the element selected by the user is "Deposit amount". The position information of the element selected by the user in the contract document includes the paragraph index of the clause paragraph to which the element selected by the user belongs, as well as the start position and end position of the element selected by the user in the clause paragraph. Among them, the paragraph index indicates which paragraph of the contract document the element selected by the user is located in. For example, if the element selected by the user "50% of the contract price, i.e., RMB Thirty-four thousand one hundred yuan (¥: 34,100 yuan)" is in the first article of the part "II. Payment Method" in the contract document, the paragraph index can be expressed as 2.1. The start position and end position refer to the character start and end positions of the element selected by the user in the clause paragraph. For example, if "50% of the contract price" is the 10th character in the paragraph, and "yuan (¥: 34,100 yuan)" is the 30th character, then the start position is 10 and the end position is 30. Finally, the above information is saved as user annotation information to form a contract document with structured annotations, that is, a contract document with user annotation data added is obtained.

[0094] Based on this, the user annotates the contract document on the front-end page and records the user annotation operation to generate a contract document with structured information of user annotation data, which not only facilitates subsequent information extraction and analysis but also improves the automation level of contract management.

[0095] Figure 4 The flowchart of an optimization method for a contract extraction model provided by an embodiment of the present application is shown. Refer to Figure 4 As shown, the above step S103 optimizes the contract extraction model according to the contract document with user annotation data added to obtain an optimized contract extraction model, which specifically includes the following steps:

[0096] S401. Perform model training on the contract extraction model based on the contract document with user annotation data added to obtain a trained contract extraction model.

[0097] Exemplarily, before using the contract documents with added user-annotated data for model training, the user-annotated data is first cleaned and format-converted, and then the user-annotated data is divided into a training set, a validation set, and a test set, with the division ratio being, for example, 70% for the training set, 15% for the validation set, and 15% for the test set. In this application, for the clause classification task, a model architecture of biLSTM+Softmax can be selected, and for the element recognition task, a model architecture of biLSTM+CRF can be selected. On this basis, the text is converted into a vector representation using word embedding or subword embedding, and the labels annotated by the user are converted into a format acceptable to the model using the annotation method (Begin, Inside, Outside, BIO). Then, an appropriate loss function, such as the cross-entropy loss function, and an optimization algorithm (such as Adam, SGD) are selected to minimize the loss function, and during the training process, the model parameters are iteratively adjusted to improve the model's ability to extract key information from the contract documents.

[0098] S402. Evaluate the performance of the trained contract extraction model, determine the performance metric values of the trained contract extraction model, and compare the performance metric values of the trained contract extraction model with those of the contract extraction model.

[0099] Exemplarily, during the training process, the performance of the model is regularly evaluated using the validation set to prevent overfitting, and after the training is completed, the final performance of the trained contract extraction model is evaluated using an independent test set. For example, calculate the performance metrics of the trained contract extraction model, such as the comprehensive evaluation metric (F1 value), accuracy, and recall rate, etc. Among them, the comprehensive evaluation metric (F1 value) is used to measure the performance of the model in identifying key contract information, which takes into account both the accuracy of the model's prediction (accuracy) and whether the model misses some key information (recall rate), and the higher the F1 value, the better the overall performance of the model. In addition, in this application, a threshold value of the F1 value can be set, for example, 0.80. Only when the F1 value of the optimized contract extraction model is greater than this threshold will the optimized contract extraction model be put online to replace the old version of the contract extraction model, so as to ensure that the performance of the contract extraction model put online always remains at a high level.

[0100] Furthermore, compare the performance metric values of the newly trained contract extraction model with those of any previous version of the contract extraction model to determine whether the newly trained contract extraction model has a significant improvement in performance.

[0101] S403. If the performance metric value of the trained contract extraction model is greater than that of the contract extraction model, then use the trained contract extraction model as the optimized contract extraction model and perform model update. Otherwise, re-annotate the contract documents to continuously optimize the trained contract extraction model.

[0102] Exemplarily, if the performance metric value (e.g., F1 value) of the new trained contract extraction model is higher than that of the old model, it is considered that the new model is better, which also means that the new trained contract extraction model can extract the required key information from the contract documents more accurately and should be used as the optimized contract extraction model online. Specifically, save the optimized contract extraction model as a file and deploy it to the production environment to replace the old version of the contract extraction model.

[0103] Optionally, re-annotating the contract documents to continuously optimize the trained contract extraction model includes: in response to correcting the user annotation data and incrementally annotating the contract documents to obtain new user annotation data; optimizing the trained contract extraction model according to the new user annotation data until the performance metric value of the optimized contract extraction model is greater than that of the trained contract extraction model.

[0104] Exemplarily, if the new trained contract extraction model is not significantly better than the old model, further optimization and improvement are needed, such as re-annotating the data, checking and correcting the errors in the existing user annotation data, adding more high-quality annotation data, or adjusting the model parameters or architecture, such as trying different model configurations, hyperparameter settings, and adopting other advanced model architectures, or performing data augmentation to generate more training samples through data augmentation techniques to improve the generalization ability of the model.

[0105] Specifically, carefully review the existing user annotation data to identify possible errors or inconsistencies, such as some entities may be mislabeled as different categories, or the use of labels is not unified enough. And once a problem is found, it should be corrected immediately, such as manually adjusting the annotation to ensure that each entity is correctly classified and all labels meet the predefined standards. In addition, consistency checks can be performed on the existing user annotation data to ensure that the same type of entities in the entire user annotation data always use the same label. For example, "Party A" should always be labeled as "contract subject" instead of sometimes being labeled as "company name".

[0106] In addition, to increase the diversity and quantity of training data, some new contract documents can be selected for annotation. The new contract documents can come from different fields, industries, or have different text formats to help the model generalize better. For the selected new contract documents, annotation is carried out according to the same criteria as before to increase the amount of user-annotated data, and the newly added user-annotated data is merged into the existing user-annotated data, updated and formed into a larger and more comprehensive dataset, and then the contract extraction model is retrained using the updated high-quality annotated data. Since the quality of the training data has been improved, the contract extraction model has the opportunity to learn more accurate feature representations, thus being able to improve the performance of the contract extraction model. After each retraining, the performance metrics of the new model need to be evaluated using an independent test set and compared with the contract extraction model after the previous training. If the performance of the new model still does not meet the expectations, the above steps are repeated, and the annotated data is continuously corrected and incrementally annotated until a new model with significantly better performance than the previous version model is found. Once a new model with significantly improved performance is found, it can be deployed online as the final optimized contract extraction model to replace the old model.

[0107] Based on this, the present application provides a cyclic optimization mechanism that uses user feedback, i.e., user-annotated data, to continuously improve the performance of the contract extraction model. This can not only ensure that the contract extraction model is always in the best state but also gradually improve the accuracy and reliability of the contract extraction model as the data grows and the quality improves. Especially in contract processing scenarios that require high precision and reliability, it can effectively support business requirements.

[0108] Figure 5 The flowchart of another contract information extraction method provided by an embodiment of the present application is shown. Refer to Figure 5 As shown, after obtaining the optimized contract extraction model, the method further includes:

[0109] S501. Perform information extraction on the contract document based on the optimized contract extraction model to obtain an initial extraction result, where the initial extraction result includes multiple key entities.

[0110] Exemplarily, use the optimized contract extraction model to perform automated information extraction on the existing contract document. The optimized contract extraction model will output a set containing multiple key entities, and each key entity has attributes such as entity text, entity type, and location information in the contract document. Among them, the entity text is, for example, "RMB 10,000 in full", the entity type is, for example, "amount", and the location information in the contract document includes the start position and the end position. For example, the start position of "RMB 10,000 in full" may be the 23rd character, and the end position is the 30th character.

[0111] S502. Arrange the initial extraction results according to the position information of each key entity in the contract document to obtain the target extraction results.

[0112] Optionally, since the initial extraction results may be disordered in sequence. For example, the "Name of Party A" is extracted, but it is not clear whether this Name of Party A belongs to the beginning or the end of the contract, and the "Contact Information" extracted is for Party A or Party B. To make the initial extraction results more in line with the logical structure and correctly reflect the relationships between various key entities, it is necessary to further arrange the key entities in the initial extraction results.

[0113] Optionally, the above step S502 specifically includes determining the character offset of each key entity according to the start position and end position of each key entity in the contract document; determining the context distance of each key entity according to the character offset of each key entity and the position index of the preset key anchor; arranging and combining each key entity according to the context distance of each key entity to obtain the target extraction results.

[0114] Exemplarily, the start position and end position of a key entity in the contract document clarify the exact position of the key entity in the contract document. The character offset of the key entity can be determined based on this exact position, which can be specifically achieved by recording the start and end character positions of the key entity. For example, the start position of "Contract Party - Party A" is the 45th character, and the end position is the 47th character, so the character offset is (45, 47).

[0115] Exemplarily, the preset key anchor refers to certain specific positions or paragraphs in the contract document, which can be used as reference points to measure the distance of other key entities relative to the preset key anchor. For example, in a contract document, the "Signing Date" can be an important key anchor. Further, the context distance refers to the relative distance between each key entity and the preset key anchor. Based on this relative distance, the relative positions of each key entity in the entire contract document and the logical relationships between key entities can be understood. For example, if a certain amount entity appears immediately after the "Signing Date", there may be some association between these two entities.

[0116] Exemplarily, based on the context distances of each key entity, the order of each key entity is reorganized so that the final extraction result not only reflects the content of each key entity itself but also embodies the logical relationships among the key entities in the contract document. For example, all entities related to "payment terms" (such as amount, payment method) are grouped together, while entities related to "liability for breach of contract" are placed in another group. In this way, the target extraction result obtained after arrangement is easier to understand and analyze. It is not just a simple list of entities but an information set organized according to a certain logical structure, which also helps in subsequent application scenarios such as automated auditing, clause comparison, etc.

[0117] Based on this, the optimized contract extraction model obtained through iterative optimization is used for information extraction, and the extracted information is arranged to ensure the logical coherence and accuracy of the extraction result.

[0118] Figure 6 The flowchart of a method for iterative optimization of a contract extraction model based on user feedback provided by an embodiment of the present application is shown. As a possible implementation, referring to Figure 6 As shown, in the present application, a natural language processing engine (NLP engine) can disassemble a contract document into multiple logical paragraphs. For a contract document containing multiple logical paragraphs, the user interacts with the contract document through a front-end page and performs annotation processing on the contract document, including but not limited to highlighting the selected elements by the user and tagging (adding labels), to obtain a contract document with user annotation data added, and storing the user annotation data in a back-end database. Additionally, before performing a training task using the contract document with user annotation data added, the content of the user annotation data is first reviewed to standardize the format of the user annotation data, ensuring that high-quality user annotation data enters the training set. Then, the reviewed user annotation data is used to train the existing old version of the contract extraction model, and it is determined whether the accuracy of the optimized contract extraction model obtained through training has increased, that is, it is determined whether there is a significant improvement in the model performance of the new optimized contract extraction model compared to the previous version of the old model. If so, the model is updated online to replace the old version model. Otherwise, the model update is postponed, and the user annotation data and incremental annotation data are re-reviewed to iteratively perform model training to achieve continuous optimization of the contract extraction model.

[0119] Exemplarily, after reviewing the user annotation data, the user submits the reviewed user annotation data to the back-end. After the back-end standardizes the reviewed user annotation data, the data obtained after the standardization process is transmitted to the NLP engine in the form of a file. The transmission and reception interface can be as shown in Table 1 below:

[0120] Table 1 Annotation Data Receiving Interface

[0121]

[0122] Exemplarily, after receiving all user-annotated data, the NLP engine pre-processes the user-annotated data, including but not limited to noise removal, format conversion, etc., so that the user-annotated data meets the input requirements of model training, and then stores the cleaned annotated data in the background database to provide a stable data source for subsequent model training. After storage is completed, if the preconditions are met, the training process is started, and the training log is recorded in real time for subsequent monitoring and analysis. After the training is completed, the training results, such as training success or failure, are fed back to the backend to ensure that the business party can obtain the training status in a timely manner. Among them, the preconditions include whether the user-annotated data has been stored and whether the amount of user-annotated data meets the minimum training requirements.

[0123] Specifically, the interfaces involved in data transmission between the NLP engine and the backend are shown in Table 2 below, and the request parameters involved are shown in Table 3 below:

[0124] Table 2 Interface parameters

[0125]

[0126] Table 3 Request parameter diagram

[0127]

[0128] Based on this, the present application realizes the adaptive optimization of the contract extraction model by introducing user feedback annotation, review mechanism, incremental training and model evaluation online. Compared with the prior art, the present application introduces user annotation and feedback mechanism, allowing users to directly annotate the contract text, and use the user annotation data for model retraining, thereby reducing the dependence on data annotation. The user annotation data can be directly used for model training without the need for additional large-scale annotation data, and after the user annotation data is reviewed, the model retraining can be quickly started, shortening the model update cycle. Moreover, the user annotation data reflects the personalized needs of users, so that the contract extraction model can be adjusted in a targeted manner according to the personalized needs of users, thereby improving the adaptability and accuracy of the contract extraction model. In addition, after each model retraining, the model performance is evaluated to determine whether the optimized model needs to be put on the shelf, and some users with unsatisfactory prediction results can be annotated again and the optimization process can be repeated, forming a closed-loop model dynamic continuous optimization mechanism, avoiding the problem of model performance degradation over time, and the user feedback mechanism forms a closed-loop optimization, so that the contract extraction model can continuously learn and improve.

[0129] Based on the same inventive concept, an embodiment of the present application further provides a contract information extraction device corresponding to the contract information extraction method. Since the principle of solving problems by the contract information extraction device in the embodiment of the present application is similar to the above contract information extraction method in the embodiment of the present application, the implementation of the contract information extraction device can refer to the implementation of the contract information extraction method, and the repeated parts will not be elaborated.

[0130] Referring Figure 7 As shown, it is a schematic structural diagram of a contract information extraction device provided by an embodiment of the present application. The contract information extraction device 700 includes: an acquisition module 701, a marking module 702, and a model optimization module 703, where:

[0131] The acquisition module 701 is configured to acquire a contract document, and the contract document includes a plurality of clause paragraphs;

[0132] The marking module 702 is configured to perform marking processing on the contract document to obtain a contract document with user marking data added. The user marking data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document;

[0133] The model optimization module 703 is configured to optimize the contract extraction model according to the contract document with user marking data added to obtain an optimized contract extraction model. The optimized contract extraction model is used to perform information extraction on the contract document to obtain a target extraction result.

[0134] Based on this, according to the contract information extraction device of the embodiment of the present application, it allows users to directly perform marking processing on the contract document to obtain a contract document with user marking data added, and the user marking data can be directly used for model retraining, reducing the dependence on model training data marking and effectively shortening the model update cycle. And since the user marking data reflects the personalized needs of users in the actual business scenario, using the contract document with user marking data added to optimize the contract extraction model enables the contract extraction model to be adjusted according to the personalized needs of users, improving the adaptability and accuracy of the optimized contract extraction model while also enhancing the pertinence of the extraction result.

[0135] In a possible implementation manner, the above marking module 702 is specifically configured to:

[0136] Respond to the selection operation on the elements in the contract document, and display an optional label menu in the preset area where the user-selected elements are located;

[0137] Respond to the touch operation on the optional labels in the optional label menu, and assign labels to the user-selected elements. The labels are used to indicate the element categories of the user-selected elements;

[0138] Obtain the position information of the user-selected elements in the contract document, and use the position information, the user-selected elements, and the labels of the user-selected elements as user annotation data to obtain the contract document with user annotation data added. Among them, the position information includes the paragraph index of the clause paragraph to which the user-selected element belongs, and the start position and end position of the user-selected element in the clause paragraph to which it belongs.

[0139] In a possible implementation manner, the above-mentioned annotation module 702 is further configured to:

[0140] Use an identification algorithm to identify the annotatable elements in the contract document, and generate each annotatable element into an optional control.

[0141] In a possible implementation manner, the above-mentioned model optimization module 703 is specifically configured to:

[0142] Perform model training on the contract extraction model based on the contract document with user annotation data added to obtain a trained contract extraction model;

[0143] Evaluate the performance of the trained contract extraction model, determine the performance index value of the trained contract extraction model, and compare the performance index value of the trained contract extraction model with the performance index value of the contract extraction model;

[0144] If the performance index value of the trained contract extraction model is greater than the performance index value of the contract extraction model, use the trained contract extraction model as the optimized contract extraction model and perform model update. Otherwise, re-annotate the contract document to continuously optimize the trained contract extraction model.

[0145] In a possible implementation manner, the above-mentioned model optimization module 703 is specifically configured to:

[0146] In response to the correction of the user annotation data and the incremental annotation of the contract document, obtain new user annotation data;

[0147] Optimize the trained contract extraction model according to the new user annotation data until the performance index value of the optimized contract extraction model is greater than the performance index value of the trained contract extraction model.

[0148] In a possible implementation manner, the above-mentioned model optimization module 703 is further configured to:

[0149] Perform information extraction on the contract document based on the optimized contract extraction model to obtain an initial extraction result, and the initial extraction result includes multiple key entities;

[0150] Arrange the initial extraction result according to the position information of each key entity in the contract document to obtain a target extraction result.

[0151] In a possible implementation manner, the above-mentioned model optimization module 703 is further configured to:

[0152] Determine the character offsets of each key entity according to the start position and end position of each key entity in the contract document;

[0153] Determine the context distances of each key entity according to the character offsets of each key entity and the position indexes of preset key anchors;

[0154] Arrange and combine each key entity according to the context distance of each key entity to obtain the target extraction result.

[0155] Descriptions of the processing flows of the modules in the device and the interaction flows between the modules may refer to the relevant descriptions in the above method embodiments, and will not be elaborated here.

[0156] An embodiment of the present application further provides an electronic device 800, as Figure 8 shown, which is a schematic structural diagram of the electronic device 800 provided by the embodiment of the present application, including: a processor 801, a memory 802, and optionally, a bus 803 may also be included. The memory 802 stores machine-readable instructions executable by the processor 801. When the electronic device 800 runs, the processor 801 communicates with the memory 802 through the bus 803. When the machine-readable instructions are executed by the processor 801, the steps in the contract information extraction method described in any one of the above are executed.

[0157] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps in the contract information extraction method described in any one of the above are executed.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the method embodiments, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical or other forms.

[0159] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0160] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.

Claims

1. A method for extracting contract information, characterized in that, Including: Obtain a contract document, where the contract document includes multiple clause paragraphs; Perform annotation processing on the contract document to obtain a contract document with user annotation data added, where the user annotation data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document; Optimize a contract extraction model according to the contract document with user annotation data added to obtain an optimized contract extraction model, where the optimized contract extraction model is used to perform information extraction on the contract document to obtain a target extraction result.

2. The method according to claim 1, wherein The performing annotation processing on the contract document to obtain a contract document with user annotation data added includes: Respond to a selection operation on an element in the contract document, and display an optional label menu in a preset area where the user-selected element is located; Respond to a touch operation on an optional label in the optional label menu, and assign a label to the user-selected element, where the label is used to indicate the element category of the user-selected element; Obtain the position information of the user-selected element in the contract document, and use the position information, the user-selected element, and the label of the user-selected element as the user annotation data to obtain a contract document with user annotation data added, where the position information includes the paragraph index of the clause paragraph to which the user-selected element belongs and the start position and end position of the user-selected element in the clause paragraph to which it belongs.

3. The method according to claim 2, characterized in that, Before the responding to a selection operation on an element in the contract document and displaying an optional label menu in a preset area where the user-selected element is located, the method further includes: Use an identification algorithm to identify the annotatable elements in the contract document, and generate each of the annotatable elements into an optional control.

4. The method according to claim 1, wherein The optimizing a contract extraction model according to the contract document with user annotation data added to obtain an optimized contract extraction model includes: Perform model training on the contract extraction model based on the contract document with user annotation data added to obtain a trained contract extraction model; Perform performance evaluation on the trained contract extraction model, determine the performance index value of the trained contract extraction model, and compare the performance index value of the trained contract extraction model with the performance index value of the contract extraction model; If the performance index value of the trained contract extraction model is greater than the performance index value of the contract extraction model, use the trained contract extraction model as the optimized contract extraction model and perform model update. Otherwise, re-annotate the contract document to continuously optimize the trained contract extraction model.

5. The method according to claim 4, characterized in that, The re-annotating the contract document to continuously optimize the trained contract extraction model includes: Respond to the correction of user annotation data and the incremental annotation of the contract document to obtain new user annotation data; Optimize the trained contract extraction model according to the new user annotation data until the performance index value of the optimized contract extraction model is greater than the performance index value of the trained contract extraction model.

6. The method according to claim 1, characterized in that, After optimizing the contract extraction model according to the contract document with the added user annotation data to obtain the optimized contract extraction model, it further includes: Performing information extraction on the contract document based on the optimized contract extraction model to obtain an initial extraction result, where the initial extraction result includes multiple key entities; Arranging the initial extraction result according to the position information of each key entity in the contract document to obtain the target extraction result.

7. The method according to claim 6, wherein The arranging the initial extraction result according to the position information of each key entity in the contract document to obtain the target extraction result includes: Determining the character offset of each key entity according to the start position and end position of each key entity in the contract document; Determining the context distance of each key entity according to the character offset of each key entity and the position index of a preset key anchor point; Arranging and combining each key entity according to the context distance of each key entity to obtain the target extraction result.

8. A contract information extraction device, characterized in that, Includes: An acquisition module for acquiring a contract document, where the contract document includes multiple clause paragraphs; A annotation module for performing annotation processing on the contract document to obtain a contract document with added user annotation data, where the user annotation data includes user-selected elements, labels of the user-selected elements, and position information of the user-selected elements in the contract document; A model optimization module for optimizing the contract extraction model according to the contract document with the added user annotation data to obtain an optimized contract extraction model, where the optimized contract extraction model is used to perform information extraction on the contract document to obtain a target extraction result.

9. An electronic device, characterized in that, Includes: A processor and a memory, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor executes the machine-readable instructions to perform the steps of the contract information extraction method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it performs the steps of the contract information extraction method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Contract element extraction method and related device

    CN121542444A