Contract semantic comparison method and system based on multi-modal model

Through the contract semantic comparison method based on multimodal model, the problems of inconsistent contract formats, poor OCR recognition accuracy and insufficient semantic analysis in the prior art are solved, and the comprehensiveness and accuracy of contract comparison are achieved, thereby reducing legal risks.

CN120087373AInactive Publication Date: 2025-06-03BEIJING POWER LAW INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510146864.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing contract management system has problems such as inconsistent formats, poor OCR technology recognition accuracy, lack of semantic depth comparison and insufficient processing of non-text content when dealing with multi-format and multi-version contracts, resulting in inaccurate and incomplete comparison results.

Method used

The contract semantic comparison method based on multimodal model is adopted, and the in-depth semantic analysis is carried out through the steps of document preprocessing, core element extraction, semantic consistency detection, conflict set generation and result display, combined with text recognition, image processing and data analysis technology, text, tables, charts, signatures and other contents in the contract are processed to conduct in-depth semantic analysis.

Benefits of technology

Accurate comparison of contracts of different formats and versions is achieved, and the problems of traditional technology in format differences, OCR recognition errors and insufficient semantic analysis are overcome, ensuring the comprehensiveness and accuracy of the comparison results, and reducing legal risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087373A_ABST
    Figure CN120087373A_ABST
Patent Text Reader

Abstract

The invention discloses a contract semantic comparison method and system based on a multi-modal model, and the key point of the technical scheme is that the method comprises the following steps: S1, preprocessing a document; s2, extracting core elements; s3, semantic consistency detection is carried out; s4, generating a conflict set; s5, displaying a result; the invention provides a contract semantic comparison method based on a multi-modal model, which solves the key problem in the prior art, unifies the contract format through document preprocessing, and overcomes the influence of format difference. Contents such as texts, tables, charts and signatures in the contract are extracted by adopting a multi-modal model, and the identification error of the OCR technology is solved; performing semantic consistency detection by using a large language model to ensure the consistency of the contract in the aspects of rights and obligations; meanwhile, non-text content is processed, contract differences are comprehensively displayed, the technical innovation improves the accuracy and efficiency of contract comparison, and the method is particularly suitable for complex contract management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of contract management, and particularly relates to a contract semantic comparison method and system based on a multimodal model. Background Art

[0002] In modern enterprise operations, contract management is of great importance. Especially in the context of the increasing globalization of the economy and the growing cooperation among enterprises, contracts not only reflect the rights and obligations of both parties but also serve as the legal basis for transactions. With the complexity of contract content, form, and language, the demand for contract management systems has also increased. In business operations, enterprises often need to compare different versions of contracts, especially when comparing electronic contracts (such as Word documents) with the final printed and stamped paper contracts (usually scanned as PDF files). Contracts are often modified, adjusted, and reviewed during the signing and archiving processes. Therefore, ensuring the accuracy and consistency of the final contract is crucial for avoiding legal risks.

[0003] Currently, most contract management systems rely on OCR technology to convert the images in scanned PDF into text for character-level comparison. However, this character-based comparison method has several technical drawbacks, specifically manifested as follows:

[0004] 1. Contract format and layout differences: Different versions of contracts may have differences in format and layout. Existing methods cannot effectively handle these format changes, resulting in inaccurate comparison results.

[0005] 2. Limitations of OCR technology: When dealing with low-quality scanned documents, complex layouts, and non-text content (such as handwritten signatures, tables, and charts), the recognition accuracy of OCR technology is poor, and errors are likely to occur.

[0006] 3. Lack of in-depth semantic comparison: Existing technologies usually rely on character comparison, ignoring the complex legal terms and professional content in contract clauses, and unable to determine whether the rights and obligations are equal, which is prone to false positives or false negatives.

[0007] 4. Processing problems of non-text content: Non-text content such as tables, charts, and signatures in contracts is difficult to accurately process in traditional OCR and text comparison, resulting in incomplete and inaccurate comparison results.

[0008] Therefore, existing contract comparison technologies have significant limitations in comparing multi-format and multi-version contracts. There is an urgent need for a new comparison method that can transcend the limitations of character comparison, perform in-depth semantic analysis, and handle complex text formats and non-text elements. To solve the above problems, we propose a contract semantic comparison method and system based on a multimodal model. Summary of the Invention

[0009] In view of the deficiencies of the prior art, the present invention provides a contract semantic comparison method and system based on a multi-modal model to solve the problems raised in the background art.

[0010] The above technical objectives of the present invention are achieved through the following technical solutions:

[0011] A contract semantic comparison method based on a multi-modal model, the method comprising the following steps:

[0012] S1. Document preprocessing: Preprocess the input contract file, where the contract file is a Word file or a scanned PDF file. The preprocessing steps include converting the contract file into an image-type PDF file in a unified format, and performing image enhancement, format adjustment, and quality unification processing, while retaining the visual information of the non-text content in the contract file.

[0013] S2. Core element extraction: Use a multi-modal model to extract the core elements from the preprocessed image-type PDF file. The core elements include, but are not limited to, semantic-level information such as terms, rights, and obligations in the contract. The multi-modal model can process various formats of content in the contract, such as text, tables, and charts.

[0014] S3. Semantic consistency detection: Use a large language model to perform semantic consistency comparison on the extracted contract core elements, detect the semantic consistency of each element in contract A with the corresponding element in contract B at the levels of rights, obligations, etc., and collect the inconsistent element items into the conflict set.

[0015] S4. Conflict set generation: Merge the conflict element sets in contract A and contract B respectively to generate the final conflict set, which includes all the rights and obligation conflict items existing between contract A and contract B.

[0016] S5. Result display: Visually display the merged conflict set for the user to view and analyze the differences in the contract content. The visual display shows all the conflict items in the form of graphs or tables, facilitating further legal analysis and decision-making by the user.

[0017] Further, the multi-modal model is a comprehensive model that combines text recognition, image processing, and data parsing, and can process non-text content such as text, tables, graphics, and signatures in the contract.

[0018] Further, the large language model is based on a pre-trained legal domain corpus to semantically understand the legal terms and professional clauses involved in the contract, ensuring the high accuracy of the comparison results.

[0019] Further, the semantic consistency detection step further includes:

[0020] Compare each element in Contract A one by one to detect whether the corresponding element in Contract B is consistent in terms of rights and obligations;

[0021] Compare each element in Contract B one by one to detect whether the corresponding element in Contract A is consistent in terms of rights and obligations;

[0022] Record the inconsistent elements as conflict items and merge all conflict items to generate a final conflict set.

[0023] Furthermore, the core element extraction step adopts an iterative loop algorithm to gradually extract key elements such as rights and obligations in the contract until all relevant semantic elements are fully extracted.

[0024] The present invention also provides a contract semantic comparison system based on a multimodal model, including:

[0025] Document preprocessing module: used to receive the input contract file, convert the contract file into an image-type PDF file in a unified format, and perform preprocessing operations such as image enhancement and format adjustment;

[0026] Core element extraction module: used to extract core elements from the preprocessed image-type PDF file based on a multimodal model, including terms, rights, obligations, etc. in the contract;

[0027] Semantic consistency detection module: used to use a large language model to detect the semantic consistency of the core elements, compare the consistency of rights and obligations of corresponding elements in Contract A and Contract B, and collect the inconsistent elements into a conflict set;

[0028] Conflict set merging module: used to merge the conflict element sets extracted from Contract A and Contract B to generate a final conflict set;

[0029] Result display module: used to visually present the final conflict set, display the differences in rights and obligations between contracts for users to analyze and review.

[0030] Furthermore, the multimodal model can simultaneously process non-text content such as text, tables, graphics, signatures, etc. in the contract and perform comprehensive analysis of contract elements by integrating various data processing technologies.

[0031] Furthermore, the large language model is a pre-trained model optimized for the legal field and can understand and process legal professional terms and complex clauses.

[0032] Furthermore, the conflict set merging module further includes:

[0033] Merge all the detected inconsistencies in Contract A and Contract B to generate a merged conflict set;

[0034] Display the inconsistent clauses in the conflict set in the form of graphs or tables to facilitate users to quickly identify the differences in the contracts.

[0035] In summary, the present invention mainly has the following beneficial effects:

[0036] 1. Through the application of the multimodal model, the present invention can handle the differences in format, typesetting, and paragraphs in different versions of contracts, thus avoiding the format inconsistency problems that cannot be solved by traditional character comparison methods, and ensuring the accuracy and consistency of the comparison results.

[0037] 2. The present invention combines text recognition, image processing, and data parsing technologies to accurately identify the contract content in low-quality scanned documents and complex layouts, and effectively handle non-text elements, overcoming the deficiencies of traditional OCR technologies in this regard.

[0038] 3. Through semantic analysis of the core elements of the contract by the large language model, the present invention can go beyond simple character comparison and deeply judge whether the rights and obligations of legal terms and professional content in the contract clauses are equal, thus avoiding false positives and false negatives in traditional technologies.

[0039] 4. The present invention can comprehensively handle the text and non-text content in the contract through the multimodal model, including tables, charts, signatures, etc., thus ensuring the comprehensiveness and accuracy of the comparison and solving the limitations of traditional methods in dealing with non-text content. Description of the Drawings

[0040] Figure 1 is the overall flowchart of contract comparison of the present invention;

[0041] Figure 2 is the flowchart of semantic consistency comparison of the present invention. Detailed Embodiments

[0042] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] The following embodiments are used to illustrate the present invention, but cannot be used to limit the protection scope of the present invention. The conditions in the embodiments can be further adjusted according to specific conditions. Any simple improvement to the method of the present invention under the premise of the concept of the present invention belongs to the scope of protection required by the present invention.

[0044] Embodiment 1

[0045] A method and system for contract semantic comparison based on a multimodal model, referring to Figure 1 - Figure 2 , the application of the method for contract semantic comparison based on a multimodal model in this embodiment specifically includes the following steps:

[0046] Step 1: Document preprocessing

[0047] In this embodiment, assume that Contract A is a scanned PDF file and Contract B is a Word document. The system first preprocesses these two contract files:

[0048] 1. Since Contract A is a scanned document, the system uses image processing technology to convert it into an image-type PDF in a unified format and enhances the image to ensure that the text, signatures, and tables in the contract can be clearly recognized;

[0049] 2. Contract B is an electronic Word document, and the system converts it into an image-type PDF format for comparison in the same format as Contract A.

[0050] The goal of this process is to unify the formats of the two contracts so that subsequent comparison operations can be carried out on the same basis.

[0051] Step 2: Core element extraction

[0052] 1. After Contract A and Contract B are preprocessed, the system uses a multimodal model to extract core elements. This model combines text recognition, image processing, and data parsing technologies and can simultaneously extract non-text elements such as text, tables, and signatures in the contract;

[0053] 2. For example, the system extracts the terms, rights, obligations, payment arrangements, liability for breach of contract, etc. in the contract. Through the multimodal model, the text information, table information, and image data in the contract are accurately converted into structured data for subsequent analysis and comparison.

[0054] Step 3: Semantic consistency detection

[0055] 1. Use a large language model to perform semantic consistency detection on the core elements in Contract A and Contract B. The large language model is based on a pre-trained corpus in the legal field, can understand legal terms in the contract clauses, and accurately judge the meanings of each clause;

[0056] 2. The system compares each clause in Contract A and Contract B one by one to check for semantic inconsistencies in rights, obligations, responsibilities, etc. For example, the "liability for breach of contract" clause stipulated in Contract A may be different from the definition in Contract B. The system will automatically detect this difference and mark it as a conflict item.

[0057] Step Four: Conflict Set Generation

[0058] 1. The system combines the inconsistent items detected in Contract A and Contract B to generate a conflict set. All inconsistent rights and obligation clauses, payment arrangements, etc. are collected into this set to form a detailed conflict report.

[0059] 2. The conflict set includes a detailed comparison of the clause content and possible legal risks.

[0060] Step Five: Result Display

[0061] 1. Finally, the system displays the conflict set through a visual interface. Users can view all the inconsistent items between Contract A and Contract B. The system supports two display methods: graphics and tables, which facilitate users to quickly analyze the differences in the contract content and make further legal decisions.

[0062] Embodiment 2

[0063] A contract semantic comparison method and system based on a multimodal model, referring to Figure 1 - Figure 2 , the application of the contract semantic comparison method based on the multimodal model in cross-border mergers and acquisitions specifically includes the following steps:

[0064] Step One: Document Preprocessing

[0065] During the cross-border corporate merger and acquisition process, Contract A is a scanned PDF file and Contract B is an electronic Word file. The system first performs unified format conversion and preprocessing on these two contract files:

[0066] 1. Convert Contract A into an image-based PDF file and perform image enhancement and format adjustment.

[0067] 2. Convert Contract B from a Word file into an image-based PDF to ensure consistent contract formats for subsequent comparison.

[0068] Step Two: Core Element Extraction

[0069] 1. Use a multimodal model to extract the core clauses in the contract, especially those related to shareholders' rights and interests, merger and acquisition payment terms, etc. The model identifies and extracts the text information, table content, and chart data in the contract to ensure that all contract elements are accurately extracted.

[0070] Step 3: Semantic Consistency Detection

[0071] 1. The system compares the "shareholder rights and interests" clauses, payment arrangements, etc. in Contract A and Contract B to ensure the consistency in legal effect between the two contracts;

[0072] 2. The system uses a large language model to understand the legal terms in the contracts to ensure semantic consistency. For example, the system will conduct an in-depth analysis of the concept of "shareholder rights and interests" to ensure that there is no ambiguity in the definitions in the two contracts.

[0073] Step 4: Conflict Set Generation

[0074] 1. The system detects the differences in the shareholder rights and interests clauses between Contract A and Contract B, classifies these differences into the conflict set, and generates a detailed report.

[0075] Step 5: Result Display

[0076] 1. The system displays all the differences between Contract A and Contract B in the form of graphs or tables to help the merger and acquisition team identify potential risks and make decisions.

[0077] In summary, the present invention provides a contract semantic comparison method and system based on a multi-modal model, which solves several key defects in contract comparison in the prior art. First, for the problem of differences in contract format and layout, the present invention unifies the contract file format through the "document preprocessing" step to ensure that the comparison process is not affected by format differences, thereby improving the comparison accuracy. Second, for the limitations of OCR technology in low-quality scanned documents and complex layouts, the present invention uses a multi-modal model combined with text recognition, image processing, and data parsing technologies to comprehensively extract the text, tables, charts, signatures, etc. in the contracts, effectively overcoming the recognition errors of traditional OCR technology and ensuring the accuracy of contract content;

[0078] In addition, the background technology points out that the prior art cannot perform in-depth comparison at the semantic level and is prone to false positives or false negatives. The present invention uses a large language model to detect the semantic consistency of contract clauses, deeply analyzes the legal terms and professional clauses involved in the contracts, and ensures the semantic consistency between Contract A and Contract B in terms of rights and obligations. Finally, for the problem of insufficient comparison of non-text content, the multi-modal model of the present invention can process both text and non-text content at the same time, and through the "conflict set generation" step, comprehensively presents the differences between the contracts. These innovative technical solutions make the application of the present invention in contract comparison more accurate and comprehensive, especially suitable for complex contract management scenarios, and greatly improve the accuracy and efficiency of the contract management system.

[0079] Although embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The words such as "comprising" or "including" used in the present invention mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents.

[0080] Although embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A contract semantic comparison method based on a multimodal model, characterized in that: The method comprises the following steps: S1. Document preprocessing: preprocessing the input contract document, the contract document being a Word document or a scanned PDF document, the preprocessing step comprising converting the contract document into an image-type PDF file in a unified format, and performing image enhancement, format adjustment and quality uniformity processing, and retaining visual information of non-text content in the contract document; S2. Extracting core elements: extracting core elements from the preprocessed image-based PDF file using a multimodal model, wherein the core elements include but are not limited to semantic information such as terms, rights, and obligations in the contract, and the multimodal model can process content in multiple formats such as text, tables, and charts in the contract; S3. Semantic consistency detection: Use a large language model to compare the extracted core elements of the contract for semantic consistency, detect the semantic consistency of each element in contract A with the corresponding element in contract B in terms of rights, obligations, etc., and collect inconsistent element items into a conflict set; S4. Conflict set generation: merge the conflict element sets in contract A and contract B respectively to generate a final conflict set, which includes all conflicting rights and obligations between contract A and contract B; S5. Result display: The merged conflict set is visualized for users to view and analyze the differences in the contract content. The visualization displays all conflicting items in a graphical or tabular form to facilitate users to conduct further legal analysis and decision-making.

2. The method for comparing contract semantics based on a multimodal model according to claim 1 is characterized in that: The multimodal model is a comprehensive model that combines text recognition, image processing and data parsing, and can process non-text content such as text, tables, graphics, signatures, etc. in contracts.

3. The method for comparing contract semantics based on a multimodal model according to claim 1 is characterized in that: The large language model is based on a pre-trained legal field corpus to perform semantic understanding of the legal terms and professional clauses involved in the contract, ensuring high accuracy of the comparison results.

4. The method for comparing contract semantics based on a multimodal model according to claim 1 is characterized in that: The semantic consistency detection step further comprises: Compare each element in Contract A one by one to check whether the corresponding element in Contract B is consistent in terms of rights and obligations; Compare each element in Contract B one by one to check whether the corresponding element in Contract A is consistent in terms of rights and obligations; Inconsistent features are recorded as conflicts, and all conflicts are merged to generate a final conflict set.

5. The method for comparing contract semantics based on a multimodal model according to claim 1 is characterized in that: The core element extraction step adopts an iterative loop algorithm to gradually extract key elements such as rights and obligations in the contract until all relevant semantic elements are fully extracted.

6. A contract semantic comparison system based on a multimodal model, characterized in that: include: Document preprocessing module: used to receive the input contract document, convert the contract document into an image-type PDF file in a unified format, and perform preprocessing operations such as image enhancement and format adjustment; Core element extraction module: used for extracting core elements of the preprocessed image-type PDF file based on a multimodal model, including terms, rights, obligations, etc. in the contract; Semantic consistency detection module: used to perform semantic consistency detection on the core elements using a large language model, compare the consistency of rights and obligations of corresponding elements in contract A and contract B, and collect inconsistent elements into a conflict set; Conflict set merging module: used to merge the conflict element sets extracted from contract A and contract B to generate the final conflict set; Result display module: used to present the final conflict set in a visual way, showing the differences in rights and obligations between contracts for users to analyze and review.

7. The multimodal model-based contract semantic comparison system according to claim 6, characterized in that: The multimodal model is capable of simultaneously processing text, tables, graphics, signatures and other non-text content in a contract, and performing comprehensive analysis of contract elements by integrating multiple data processing technologies.

8. The multimodal model-based contract semantic comparison system according to claim 6, characterized in that: The large language model is a pre-trained model optimized for the legal field and can understand and process legal terminology and complex clauses.

9. The multimodal model-based contract semantic comparison system according to claim 6, characterized in that: The conflict set merging module further comprises: Merge all inconsistencies detected in Contract A and Contract B to generate a merged conflict set; Display inconsistent clauses in the conflict set in graphical or tabular form to help users quickly identify differences in contracts.

Citation Information

Patent Citations

  • Intelligent contract document comparison system and intelligent contract document comparison method

    CN117540723A

  • Contract information extraction method and equipment based on multi-modal large language model

    CN118734032A

  • Consistency analysis method for achievements and targets of medical and invasive projects

    CN119250080A

Cited By

  • Intelligent contract text comparison method and system based on multi-modal feature fusion

    CN120579534A

  • Service execution system, method and device based on AI, medium and equipment

    CN121787584A