Long contract information extraction method and device, equipment, medium and product
By performing text recognition, segmentation, and large-scale model processing on long contracts, the problem of low efficiency in extracting information from long contract elements through manual intervention in existing technologies has been solved, achieving efficient and accurate automated information extraction.
Patent Information
- Application Number
- CN202511101796.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies require human intervention to extract key contract elements when identifying complex and diverse long contracts, resulting in inefficiency.
A large model is used to perform text recognition and segmentation on long contract documents. By combining a semantically aware text segmenter and an embedding model, contract information vectors are extracted, and feature information is obtained through customized extraction prompts, reducing data annotation and model training.
It enables the efficient and accurate extraction of key information from long contracts without manual annotation or model training, improving information extraction efficiency and ensuring the accuracy of element information.
Smart Images

Figure CN120997859A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, particularly the application of large models in intelligent services in the fintech field, and especially to a method, apparatus, device, medium, and product for extracting information from long contracts. Background Technology
[0002] In actual transactions, trade contracts are primarily designed from the perspective of individual companies and enterprises. Contract styles vary greatly across industries and among different companies, with no standardized document format. According to relevant requirements, it is sometimes necessary to systematically input key contract elements such as contract number, transaction date, transaction amount, and performance method online.
[0003] Existing technical solutions mainly use traditional optical character recognition (OCR) schemes, annotate data as needed, and use the annotated data to train the OCR model in a targeted manner, so as to recognize the text of long contract documents through the targeted trained OCR model.
[0004] However, while OCR models are mainly used for recognizing contract texts, for the recognition results, especially for long contracts with diverse and complex formats, manual retrieval and analysis of the recognition results are still required to obtain key information. Summary of the Invention
[0005] This invention provides a method, apparatus, equipment, medium, and product for extracting information from long contracts, so as to achieve efficient extraction of element information in long contracts.
[0006] According to one aspect of the present invention, a method for extracting information from long contracts is provided, comprising:
[0007] The contract information in the long contract document is obtained by performing text recognition on the long contract document;
[0008] The contract information is segmented to obtain at least two contract information vectors;
[0009] The large model is invoked to extract information from the contract information vector to obtain the element information in the contract information.
[0010] According to another aspect of the present invention, an information extraction device for long contracts is provided, comprising:
[0011] The recognition module is used to perform text recognition on long contract documents to obtain contract information in the long contract documents;
[0012] The segmentation module is used to segment the contract information to obtain at least two contract information vectors;
[0013] The extraction module is used to call the large model to extract information from the contract information vector and obtain the element information in the contract information.
[0014] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the information extraction method for long contracts according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the long contract information extraction method according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the information extraction method for long contracts as described in any embodiment of the present invention.
[0017] This invention utilizes a large model to extract element information from long contracts without requiring data annotation and model training, significantly reducing the workload of manual annotation. Furthermore, based on the large amount of information in long contracts, the contract information is segmented as necessary before extraction to ensure that the large model can successfully extract element information. This improves the efficiency of information extraction while ensuring that the extracted element information is sufficiently accurate.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of an information extraction method for long contracts according to an embodiment of the present invention;
[0021] Figure 2 This is a flowchart of a method for extracting information from long contracts according to another embodiment of the present invention;
[0022] Figure 3 This is a schematic diagram of the structure of an information extraction device for long contracts according to another embodiment of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] Figure 1 This is a flowchart illustrating a method for extracting information from long contracts according to an embodiment of the present invention. This embodiment is applicable to situations requiring the extraction of element information from long contract documents. The method can be executed by a long contract information extraction device, which can be implemented in hardware and / or software. This device can be configured in an electronic device with corresponding data processing capabilities, such as a contract management system. Figure 1 As shown, the method includes:
[0027] S110. Perform text recognition on the long contract document to obtain the contract information in the long contract document.
[0028] S120. The contract information is segmented to obtain at least two contract information vectors.
[0029] S130. Call the large model to extract information from the contract information vector to obtain the element information in the contract information.
[0030] Among them, long-term contract documents refer to electronic scans or images of long-term contracts. Essential information refers to the basic components necessary to constitute a complete and valid contract. These elements ensure the legal validity of the contract and the clarity of the rights and obligations of both parties, such as the effective date and the amount of liquidated damages.
[0031] Specifically, users upload long contract documents to the system. The system preprocesses these documents, performing operations such as noise reduction, binarization, skew correction, and adjustments to brightness and contrast, making the text in the long contract documents easier for OCR to recognize, thus improving the accuracy of OCR recognition. The system then calls a pre-trained OCR model to perform OCR recognition on the long contract documents. OCR recognition includes layout analysis and text recognition. Layout recognition mainly uses the OCR model to identify the positional relationships of elements such as titles, paragraphs, and lists; text recognition uses the OCR model to convert the text in the document into an editable text format. After recognition, the OCR model returns the text content (plain text or formatted text) and layout structure information (such as paragraphs, titles, tables, lists, numbering, font size, etc.). The system then performs structured processing on the returned text content and layout structure information, preserving the typesetting features to obtain contract information in a structured data format (such as XML).
[0032] Because the context window of a Large LLM (Limited Model) is limited, the text length of long contracts may exceed the model's processing capacity, leading to the omission of key information or logical breaks. Before using the Large LLM for information extraction, the contract information is first segmented and subjected to specific feature engineering to obtain at least two contract information vectors, thus completing the segmentation process. The Large LLM is then invoked to extract information from each contract information vector one by one, obtaining all the element information in the contract.
[0033] This invention utilizes a large model to extract element information from long contracts, eliminating the need for data annotation and model training, thus significantly reducing the workload of manual annotation. Furthermore, based on the massive amount of information in long contracts, the contract information is segmented as necessary before extraction to ensure that the large model can successfully extract element information. This improves information extraction efficiency while ensuring that the extracted element information is sufficiently accurate.
[0034] Based on the above embodiments, optionally, before segmenting the contract information to obtain at least two contract information vectors, the method further includes:
[0035] The contract information is preprocessed; the preprocessing includes text cleaning and format adjustment.
[0036] Specifically, before inputting contract information into the large language model, necessary text preprocessing is required. This preprocessing includes text cleaning, such as removing garbled characters and duplicates generated during the OCR process, and formatting adjustments, converting the text into a format suitable for the large language model, such as segmenting and sentence-by-sentence processing. This necessary preprocessing enhances the reference value of the contract information.
[0037] Figure 2 This is a flowchart illustrating a method for extracting information from long contracts, provided as another embodiment of the present invention. This embodiment is an optimization and improvement upon the above embodiments. Figure 2 As shown, the method includes:
[0038] S210. Perform text recognition on the long contract document to obtain the contract information in the long contract document.
[0039] S220. Use a text segmenter to segment the contract information to obtain at least two contract information blocks that retain the corresponding page labels; convert each contract information block into a contract information vector and store the contract information vector in a vector database.
[0040] Specifically, traditional text segmentation methods (such as those based on punctuation, length, and chapter titles) often lack an understanding of the semantics of the text content, easily leading to semantically incoherent segmentation. Semantic-aware text segmenters, however, combine natural language understanding and machine learning techniques to semantically model the text, thus more intelligently identifying semantic boundaries. Using a semantic-aware text segmenter to segment contract information, the corresponding page labels (such as titles and page numbers) are retained during segmentation, resulting in at least two contract information blocks. An embedding model is then used to convert each contract information block into a vector, yielding at least two contract information vectors, which are stored in a vector database. Traditional keyword matching easily misses semantically similar content such as synonyms, near-synonyms, and structural variations. After vectorization, methods such as cosine similarity can be used to find the vector most semantically similar to the query, even if they use different words. By using a text segmenter and vectorizing contract information, the accuracy of subsequent information extraction is further improved.
[0041] S230. Based on the custom contract elements to be extracted, determine the target contract information vector in the vector database that is related to the contract elements to be extracted.
[0042] S240. Input the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted into the large model to obtain the element information in the contract information output by the large model.
[0043] Specifically, users can customize the contract elements to be extracted on the system's front-end interface. For example, if a user wants to know the liquidated damages stipulated in a long-term contract, they can configure the liquidated damages as a contract element to be extracted. Based on the user's configuration on the system interface, one or more contract elements to be extracted can be determined.
[0044] For each contract element to be extracted, keyword retrieval is performed on the contract information vector, and the contract information vector containing the same or synonymous expression as the contract element to be extracted is determined as the target contract vector.
[0045] The contract elements to be extracted are filled into the corresponding prompt template to obtain extraction prompts constructed based on the contract elements to be extracted, such as "Please extract the specific amount of liquidated damages based on the following content". The target contract vector and extraction prompts are input into the large model to obtain the element information in the contract information output by the large model, such as "The specific amount of liquidated damages is 50 million". By allowing users to customize the contract elements to be extracted, the flexibility of element information extraction is improved.
[0046] Based on the above embodiments, optionally, if the target contract information vector is not unique, then the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted are input into the large model, and the element information in the contract information output by the large model includes:
[0047] For each target contract information vector, the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted are input into the large model to obtain the extraction result output by the large model.
[0048] The extraction results corresponding to each target contract information vector are mutually verified, and the extraction results that pass the verification are determined as the element information in the contract information.
[0049] Specifically, there may be multiple target contract information vectors; for example, a long contract may mention the specific amount of liquidated damages in multiple paragraphs. In this case, information is extracted from multiple target contract information vectors sequentially, resulting in multiple extraction results. These extraction results are then cross-validated to check for consistency in fields (e.g., whether multiple amounts conflict). If conflicts are found, a larger model can be invoked for judgment, or manual review can be performed to obtain the verified extraction results. These verified extraction results are then confirmed as element information within the contract information. By cross-validating the extraction results, inconsistencies in element information extracted at different times are avoided.
[0050] S250. Based on a preset data format, summarize at least two elements of information to obtain a summary result in tabular format; display the summary result and the long contract document.
[0051] Specifically, the extracted information includes, but is not limited to, contract number, amount, signing date, and payment method. This information is fragmented and difficult for users to view and verify. Therefore, the information is entered into a pre-stored table template according to a specific data format, such as YYMMDD for dates and Chinese capital letters for amounts, resulting in a summary in a table format. This summary result and the original long contract document are displayed on the system's front-end interface, with the source of each information element in the summary result indicated in the long contract document. By summarizing and displaying the information according to a specific format, the efficiency of users viewing and verifying the information can be improved.
[0052] The embodiments of the present invention can improve the efficiency of users in viewing and verifying element information by summarizing and displaying element information in a specific format.
[0053] Figure 3 This is a schematic diagram of a long contract information extraction device provided in another embodiment of the present invention. Figure 3 As shown, the device includes:
[0054] The recognition module 310 is used to perform text recognition on the long contract document to obtain the contract information in the long contract document;
[0055] Segmentation module 320 is used to segment the contract information to obtain at least two contract information vectors;
[0056] Extraction module 330 is used to call the large model to extract information from the contract information vector to obtain the element information in the contract information.
[0057] The information extraction device for long contracts provided in the embodiments of the present invention can execute the information extraction method for long contracts provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0058] Optionally, the segmentation module 320 includes:
[0059] The segmentation unit is used to segment the contract information using a text segmenter to obtain at least two contract information blocks that retain the corresponding page labels;
[0060] The conversion unit is used to convert each contract information block into a contract information vector and store the contract information vector into a vector database.
[0061] Optionally, the extraction module 330 includes:
[0062] The determining unit is used to determine a target contract information vector in the vector database that is related to the contract element to be extracted, based on the custom contract element to be extracted.
[0063] The extraction unit is used to input the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted into the large model, and obtain the element information in the contract information output by the large model.
[0064] Optionally, if the target contract information vector is not unique, the extraction unit is specifically used to: for each target contract information vector, input the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted into the large model to obtain the extraction result output by the large model; mutually verify the extraction results corresponding to each target contract information vector, and determine the extraction results that pass the verification as the element information in the contract information.
[0065] Optionally, the device further includes a preprocessing module for preprocessing the contract information; the text preprocessing includes text cleaning and format adjustment.
[0066] Optionally, the device further includes:
[0067] The summary module is used to summarize at least two elements of information based on a preset data format to obtain a summary result in tabular format.
[0068] The display module is used to show the summarized results and long contract documents.
[0069] The information extraction device for long contracts, as further explained, can also execute the information extraction method for long contracts provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0070] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0071] like Figure 4As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded into the RAM 43 from storage unit 48. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0072] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0073] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as long-contract information extraction methods.
[0074] In some embodiments, the long contract information extraction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the long contract information extraction method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the long contract information extraction method by any other suitable means (e.g., by means of firmware).
[0075] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0076] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0077] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0078] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0079] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0080] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0081] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.
[0082] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for extracting information from long contracts, characterized in that, The method includes: The contract information in the long contract document is obtained by performing text recognition on the long contract document; The contract information is segmented to obtain at least two contract information vectors; The large model is invoked to extract information from the contract information vector to obtain the element information in the contract information.
2. The method according to claim 1, characterized in that, The step of segmenting the contract information to obtain at least two contract information vectors includes: The contract information is segmented using a text splitter to obtain at least two contract information blocks that retain the corresponding page labels; Each contract information block is converted into a contract information vector, and the contract information vector is stored in a vector database.
3. The method according to claim 2, characterized in that, The process of calling the large model to extract information from the contract information vector to obtain the element information in the contract information includes: Based on the custom contract elements to be extracted, determine the target contract information vector in the vector database that is related to the contract elements to be extracted; The target contract information vector and the extraction prompts constructed based on the contract elements to be extracted are input into the large model to obtain the element information in the contract information output by the large model.
4. The method according to claim 3, characterized in that, If the target contract information vector is not unique, then the target contract information vector and the extraction hints constructed based on the contract elements to be extracted are input into the large model, and the element information in the contract information output by the large model includes: For each target contract information vector, the target contract information vector and the extraction prompts constructed based on the contract elements to be extracted are input into the large model to obtain the extraction result output by the large model. The extraction results corresponding to each target contract information vector are mutually verified, and the extraction results that pass the verification are determined as the element information in the contract information.
5. The method according to claim 1, characterized in that, Before segmenting the contract information to obtain at least two contract information vectors, the method further includes: The contract information is preprocessed; the preprocessing includes text cleaning and format adjustment.
6. The method according to claim 1, characterized in that, After the large model is invoked to extract information from the contract information vector to obtain the element information in the contract information, the method further includes: Based on a preset data format, at least two elements of information are summarized to obtain a summary result in tabular format; The summarized results and long contract documents are then displayed.
7. An information extraction device for long contracts, characterized in that, The device includes: The recognition module is used to perform text recognition on long contract documents to obtain contract information in the long contract documents; The segmentation module is used to segment the contract information to obtain at least two contract information vectors; The extraction module is used to call the large model to extract information from the contract information vector and obtain the element information in the contract information.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the information extraction method for long contracts according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the information extraction method for long contracts as described in any one of claims 1-6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the information extraction method for long contracts as described in any one of claims 1-6.