A method, device, and medium for intelligent contract review based on a multimodal large model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本说明书一个或多个实施例提供了一种基于多模态大模型的智能合同审核方法、设备及介质,用于解决如下技术问题:在存在表格等多种模态的合同文档的审核过程中,因现有技术采用单一模态数据分析和浅层数据拼接的方式,存在模态处理能力割裂与跨模态逻辑关联缺失的问题,导致合同审核结果存在风险隐患
[0013]本说明书实施例采用的上述至少一个技术方案能够达到以下有益效果:通过本说明书实施例,相较于传统合同审核仅能处理单一模态或简单拼接多模态数据的局限性,本方案通过构建文本语义特征、表格关系图谱及图像视觉特征的统一表征空间,解决了异构数据语义断层问题,通过跨模态注意力机制建立双向链接,将文本条款的语义与表格的数值逻辑、图像的版式特征深度绑定,从而实现对合同要素的全维度解析,避免信息割裂导致的审核盲区;通过单模态自洽性验证与跨模态矛盾检测的双重机制,既能发现表格内公式计算错误、图像签名模糊等独立问题,又能捕捉跨模态的显式与隐式矛盾;本方通过自动化特征提取与跨模态推理,大幅减少人工介入,快速聚焦高风险项,避免全文档复核的资源浪费,提升了在批量合同处理场景下的合同审核效率。
Smart Images

Figure CN120524239B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of contract review technology, and in particular to an intelligent contract review method, device and medium based on a multimodal large model. Background Technology
[0002] Currently, existing intelligent contract review systems primarily rely on single-modal data processing technologies, such as text analysis models based on Natural Language Processing (NLP) or seal recognition tools based on Computer Vision (CV). However, real-world contract documents typically contain multiple modalities, including text, tables, and images (such as signatures and seals). While these technologies can perform grammatical checks on text clauses or simple classification of image elements in specific scenarios, they suffer from fragmented modal processing capabilities and a lack of cross-modal logical connections.
[0003] Traditional systems typically process only single data types (such as plain text or standalone images), failing to effectively integrate heterogeneous information from text, tables, and images in contract documents. For example, they cannot effectively handle numerical logic such as formula calculation chains in tables, or critical information like signatures and seal codes in images, leading to the omission of potential risks from a large amount of non-textual data during the review process. While existing solutions attempt to integrate multimodal data, they often employ simple modal splicing or manually defined cross-modal associations, achieving only shallow data splicing. For instance, comparing OCR-recognized table content independently with text clauses fails to capture the deep semantic relationships between text descriptions, table values, and image elements. For example, when text clauses stipulate… " The down payment percentage needs to be confirmed by both parties. ” At that time, it is difficult to automatically verify whether the clause is related to the specific amount in the form and the signature area on the signature page.
[0004] Therefore, in the review process of contract documents with multiple modalities such as tables, the existing technology uses single-modal data analysis and shallow data splicing, which has problems such as fragmented modal processing capabilities and lack of cross-modal logical connections, resulting in potential risks in the contract review results. Summary of the Invention
[0005] This specification provides one or more embodiments of a smart contract review method, device, and medium based on a multimodal large model to solve the following technical problem: In the review process of contract documents with multiple modalities such as tables, the existing technology uses single-modal data analysis and shallow data splicing, which results in fragmented modal processing capabilities and a lack of cross-modal logical connections, leading to potential risks in the contract review results.
[0006] One or more embodiments of this specification employ the following technical solutions:
[0007] This specification provides one or more embodiments of an intelligent contract review method based on a multimodal large model. The method includes: acquiring a target contract document to be reviewed; parsing multimodal data in the target contract document, wherein the multimodal data includes text data, tabular data, and image data; extracting features from the multimodal data to obtain multimodal feature data; mapping the multimodal feature data to a unified dimensional space to calculate the association weights between each modality through a cross-modal attention mechanism, and establishing cross-modal bidirectional link data, wherein the multimodal feature data includes any one or more of text semantic features, tabular relationship graphs, and image visual features; and performing unimodal self-consistency verification and cross-modal contradiction detection on the target contract document based on the multimodal feature data and the cross-modal bidirectional link data to generate review information for the target contract document.
[0008] This specification provides one or more embodiments of an intelligent contract review device based on a multimodal large model, comprising:
[0009] At least one processor; and,
[0010] A memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described method.
[0012] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions configured to perform the above-described method.
[0013] The above-mentioned technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: Compared with the limitations of traditional contract review, which can only handle single-modality or simple splicing of multimodal data, this solution solves the problem of semantic discontinuity of heterogeneous data by constructing a unified representation space of text semantic features, table relationship graphs and image visual features. By establishing bidirectional links through cross-modal attention mechanisms, the semantics of text clauses are deeply bound to the numerical logic of tables and the layout features of images, thereby achieving full-dimensional analysis of contract elements and avoiding review blind spots caused by information fragmentation. Through the dual mechanism of single-modal self-consistency verification and cross-modal contradiction detection, it can not only discover independent problems such as formula calculation errors in tables and blurred image signatures, but also capture explicit and implicit contradictions across modalities. Through automated feature extraction and cross-modal reasoning, this solution greatly reduces manual intervention, quickly focuses on high-risk items, avoids the waste of resources in full document review, and improves the efficiency of contract review in batch contract processing scenarios. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0015] Figure 1 A flowchart illustrating a smart contract review method based on a multimodal large model, provided for embodiments of this specification;
[0016] Figure 2 This is a schematic diagram of the structure of an intelligent contract review device based on a multimodal large model, provided as an embodiment of this specification. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0018] This specification provides a smart contract review method based on a multimodal large model. It should be noted that the execution entity in this specification embodiment can be a server or any device with data processing capabilities. Figure 1 A flowchart illustrating a smart contract review method based on a multimodal large model, as provided in the embodiments of this specification, is shown below. Figure 1 As shown, the main steps include the following:
[0019] Step S101: Obtain the target contract document to be reviewed and parse the multimodal data in the target contract document.
[0020] The multimodal data includes text data, tabular data, and image data;
[0021] In one embodiment of this specification, contract documents uploaded by users are received through a standardized interface, supporting multiple file formats including PDF, Word, and scanned images. For electronic documents (such as PDF), a layout parsing engine is used to extract the original structural information of the document, including page layout, text paragraph coordinates, table border positions, and image embedding areas. For scanned or photographed documents, image preprocessing operations are first performed, including noise reduction and correction, resolution unification, and orientation rotation, to ensure the accuracy of subsequent OCR recognition. After format parsing is completed, multimodal data in the target contract document is parsed, whereby multimodal data includes text data, table data, and image data.
[0022] Step S102: Extract features from the multimodal data to obtain multimodal feature data. Map the multimodal feature data to a unified dimensional space to calculate the association weight between each modality through a cross-modal attention mechanism and establish cross-modal bidirectional linked data.
[0023] The multimodal feature data includes any one or more of the following: text semantic features, table relationship graphs, and image visual features.
[0024] Feature extraction is performed on the multimodal data to obtain multimodal feature data. Specifically, this includes: extracting the embedding vectors of the pre-trained language model in the legal domain from the text data to generate text semantic features; performing structured parsing on the table data to construct the corresponding table relationship graph, which includes cell content and row and column relationships; and performing denoising, orientation correction, and resolution unification processing on the image data, and extracting local visual features and global layout features through a multi-scale convolutional network.
[0025] In one embodiment of this specification, a multimodal large-scale model is used to extract legal domain semantic features from the text data. For contract text data, a pre-trained language model (such as a domain-adapted model based on the BERT architecture) is first loaded. The model is trained on a general corpus and then incrementally pre-trained using professional data such as legal documents and contract cases to enhance its semantic understanding of legal terminology and clause structure. After word segmentation, context-dependent embedding vectors are generated for each word, and sentence-level or paragraph-level semantic features are aggregated through pooling layers. For contract-specific clause types (such as breach of contract and dispute resolution), an attention masking mechanism is additionally introduced to focus on capturing the contextual relevance of key clauses, ultimately outputting a text feature vector with legal semantics. Compared to general pre-trained models, the legal domain-adapted semantic extraction model can effectively distinguish ambiguities in professional terminology within contract clauses. " force majeure ” In the terms and conditions " Natural disasters ”With the insurance terms " Natural disasters ” Although the wording is the same, the domain-pre-trained model can generate differentiated embedding vectors, avoiding misjudgments of clauses caused by semantic confusion in traditional methods. At the same time, the attention masking mechanism strengthens the contextual dependencies of core clauses, making... " Method of calculating liquidated damages ” It can capture the semantics of complex expressions more accurately.
[0026] A standardized preprocessing procedure is performed on the raw image data, employing Gaussian filtering and edge-preserving algorithms to remove noise and moiré interference from the scanned documents. Subsequently, an orientation detection model (such as a convolutional network based on text line tilt angle prediction) automatically corrects the image rotation angle to ensure horizontal alignment of text lines. Finally, the image resolution is uniformly scaled to a preset size to eliminate scale deviations caused by device differences. The preprocessed image is input into a multi-scale convolutional network. Shallow layers extract local detail features (such as the continuity of signature strokes and the sharpness of seal edges) using small-sized convolutional kernels, while deep layers capture global layout features (such as the relative position of the signature area on the page and the spatial distribution of multiple seals) through large receptive field convolutions. Local and global features are then concatenated into a joint visual feature vector after channel attention weighting. Traditional image feature extraction methods typically use only a single-scale convolutional network, making it difficult to simultaneously consider the correlation between local details and global layout. Through hierarchical feature fusion of multi-scale networks, the model can recognize both the microscopic features of signature strokes (such as the uniqueness of pen stroke transitions) and the macroscopic spatial relationship between the signature and the seal (such as the compliance of the signing position). For example, when identifying electronic seals, local features can determine the clarity of the code, while global features can verify whether the seal is located in the designated area for contract signing. The combination of the two significantly improves the comprehensiveness of the review.
[0027] The table data is structured and parsed to construct a corresponding table relationship graph. Specifically, this includes: obtaining the logical row and column coordinates and cell content of each table cell; using the logical row and column coordinates as keys and the cell content as values to construct key-value pairs for each table cell, where the cell content includes the original value and any item from a formula expression; when the cell content is a formula expression, based on the cell reference relationships in the formula expression, using the target cell of the formula expression as the root node and the input parameter cells in the formula expression as leaf nodes to construct the numerical dependency relationship corresponding to the target cell; constructing row and column topology relationships based on the coordinates of each table cell, where the row and column topology relationships include the set of cells in the same row and the set of cells in the same column; and determining the table relationship graph corresponding to the table data based on the key-value pairs, the row and column topology relationships, and the numerical dependency relationships.
[0028] In one embodiment of this specification, the input tabular data is first subjected to physical structure parsing to identify the boundary position of each cell and assign logical coordinates according to row and column order. For example, the first cell in the top left corner of the table is marked as... " R1-C1 ” (Row 1, Column 1), row and column numbers increase sequentially downwards and to the right. The content of each cell is extracted as either a raw value or a formula expression. The content of each cell is then converted into key-value pairs, where the key is the logical row and column coordinates. " Row number R - Column number C ” The value is the text, number, or formula expression that begins with an equal sign after removing irrelevant symbols;
[0029] For cells containing formulas (such as...) " =R1-C1+R1-C2 ” The formula text is preserved intact, and a syntax parser separates the target cell (the cell storing the formula result) from the input parameter cells (other cells referenced in the formula). For each formula cell, a tree-like dependency relationship is established, with the cell as the root node and the referenced input parameter cells as leaf nodes. The specific construction process is as follows: regular expressions are used to match cell references in the formula to extract the logical row and column coordinates of the referenced cells; the formula is decomposed into a sequence of operators and a sequence of operands, where operands include constant values and cell references; a dependency tree node is created for each formula cell, and a child node is created for each referenced input parameter cell; the node relationships and operator types of the dependency tree are stored in the `dependency_graph` field of the table-structured JSON data. For example, if a cell... " R3-C3 ” The formula is " =R1-C1*R2-C2 ” The root node is " R3-C3 ” leaf nodes are " R1-C1 ” and "R2-C2" ” This dependency will be traced layer by layer. If the input parameter cell itself contains a formula (such as "R1-C1=R1-C3+R1-C4"), it will be considered. ” If this is done, the hierarchy of the dependency tree can be further expanded to form a complete numerical computation chain.
[0030] Using the logical coordinates of each table cell, all cells in the same row are grouped into a row-wide set, and all cells in the same column are grouped into a column-wide set. For example, cell "R2-C1" ” The set of peers is " R2-C1, R2-C2, R2-C3…” The set of columns is " R1-C1, R2-C1, R3-C1 …” This topological relationship not only records the physical adjacency of cells, but can also be used for subsequent cross-row and cross-column data validation, such as ensuring the consistency of values between the total row and the sub-item rows.
[0031] This graph structure integrates three types of data: key-value pairs (cell coordinates and content), numerical dependency trees, and row-column topological relationships. Nodes in the graph represent cells and contain attributes such as raw values or formula types. Edges are categorized into three types: numerical dependency edges (root node pointing to leaf nodes), row edges (bidirectional connections between cells in the same row), and column edges (bidirectional connections between cells in the same column). Node attributes include cell key, value, and type. Edge relationships include row_adjacent / col_adjacent type edges representing physical adjacency, and dependency type edges representing numerical dependencies. This graph structure is output as a table relationship graph to subsequent modules for linking text clauses with the table's numerical logic. It's important to note that this graph fully represents the table's numerical logic, physical layout, and data relationships, providing structured input for cross-modal review.
[0032] By employing the above technical solutions, this specification addresses the problem that conventional methods only extract the original values of tables, ignoring formula dependencies and row / column topology, which leads to the inability to detect numerical calculation errors or logical contradictions during subsequent review. The embodiments in this specification construct a numerical dependency tree, enabling the tracing of the complete input chain for formula calculations. For example, when an abnormal value is detected in a cell, the error source of the specific input parameter can be located by tracing back along the dependency tree (e.g., a parameter cell was mistakenly filled with text), significantly improving problem location efficiency. Through row / column topology modeling, it can automatically verify whether the sum of the total value and the sum of the individual items in the same row match, and whether the data types in the same column are consistent. The numerical dependency tree and graph structure in the embodiments of this specification enable text clauses (such as...) to be processed more efficiently. " Interest rate fluctuations are calculated based on Appendix Table 1. ” It can be deeply bound to the formula calculation chain in the table (such as interest rate = benchmark interest rate * floating coefficient). When the floating rule described in the text conflicts with the formula logic, the conflicting node can be quickly located along the dependency tree to achieve accurate risk tracing. Through the joint mapping of logical coordinates and physical layout, the original coverage of merged cells can be accurately restored (such as "R1-C1:R2-C2"). ” It indicates a merged area spanning two rows and two columns, and automatically continues broken rows and columns across pages to ensure data integrity.
[0033] The multimodal feature data is mapped to a unified dimensional space to calculate the association weights between each modality through a cross-modal attention mechanism, establishing cross-modal bidirectional links. Specifically, this includes: projecting the text semantic features, table relationship graphs, and image visual features from the multimodal feature data onto the unified dimensional space using learnable linear transformation matrices to generate text projection vectors, table projection vectors, and image projection vectors; determining the first cosine similarity score between the text projection vector and the table projection vector; and filtering out association pairs with the first cosine similarity score higher than a first preset threshold to establish explicit text-table links. The process involves determining the explicit link weights of text tables; determining the second cosine similarity score between the text projection vector and the image projection vector; filtering out association pairs whose second cosine similarity score is higher than a first preset threshold; establishing explicit text-image links; determining the explicit link weights of text images; determining the explicit link set based on the explicit links in the table relationship graph and the explicit link set; and determining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link set by defining the implicit link weights of the table-image links.
[0034] In one embodiment of this specification, text semantic features, table relationship graphs, and image visual features are first mapped to a unified dimensional space using independent learnable linear transformation matrices. Text semantic features (such as 768-dimensional BERT embedding vectors) are transformed into text projection vectors; cell features in the table relationship graph (such as numerical dependency encoding) are mapped to table projection vectors; and multi-scale visual features of the image (local details and global layout) are mapped to image projection vectors. All projection vectors maintain a consistent dimensionality to ensure comparability in cross-modal similarity calculations. For example, the text feature vector (768-dimensional) is mapped to a target dimension d (e.g., 512-dimensional) using a learnable linear transformation. The table feature vector (including cell values and a 256-dimensional vector encoded with topological relationships) is mapped to dimension d. The concatenated local and global image features (512 + 1024 = 1536-dimensional) are mapped to dimension d.
[0035] Subsequently, the cosine similarity between the text projection vector and the table projection vector is calculated. Association pairs with scores higher than a preset threshold are selected, and explicit text-table links are established. The first preset threshold can be set according to actual needs. For example, through experimental verification, in a contract review scenario, a threshold of 0.7 can balance recall and precision, so it can be set to 0.7. Only association pairs with scores higher than 0.7 are retained to ensure that high-confidence associations are preserved. Explicit weights are used to quantify the direct association strength between cross-modal elements. Their calculation is based on the semantic similarity of cross-modal feature vectors. Explicit weights are obtained by mapping the cosine similarity to the [0,1] interval, such as by linearly mapping the similarity score to the weight range [0,1] using the formula w = (s+1) / 2. Similarly, the cosine similarity between the text and image projection vectors is calculated, high-scoring association pairs are selected, explicit text-image links are established, and explicit weights are calculated. For example, the text clause "The down payment ratio is 30% of the total contract price" is stored in table cells R2-C3. " Total price = 1 million ” R2-C4 storage " Down payment = 300,000 ” The calculated similarity score between the text and R2-C3 was 0.9, with a weight of 0.95; the similarity score between the text and R2-C4 was 0.85, with a weight of 0.925. R2-C4, due to its higher weight, became the primary associated object. For example, text clauses... " The contract must be signed by both parties to be effective. ” The image region contains the bounding boxes of the signatures of both parties, and the similarity score is calculated to be 0.75 with a weight of 0.875.
[0036] Based on the numerical dependencies in the table relationship graph and the set of explicit links, the implicit links between the table and the image are determined, and their weights. Specifically, this includes: traversing the set of explicit links, filtering target explicit links whose target modality is the table modality, to obtain the associated table cells and target explicit weights corresponding to the target explicit links; based on the numerical dependencies, tracing back at least one dependent cell of the associated table cell, recording the dependency level of each dependent cell, and determining the implicit weight of each dependent cell based on the dependency level and the target explicit weight; obtaining the cell layout coordinates of each dependent cell and the image layout coordinates of multiple image regions in the target contract document, calculating the intersection-union ratio (IU) based on the cell layout coordinates and the image layout coordinates, and filtering cell-image region pairs whose IU is greater than a second preset threshold; determining the cross-modal implicit link weight corresponding to the cell-image region pair based on the implicit weight of each dependent cell and the IU, and determining the table image implicit link based on the cross-modal implicit link weight.
[0037] In one embodiment of this specification, the explicit link set is traversed to filter links whose target is a table modality, such as text-table explicit links. The table cells associated with these explicit links, such as the down payment amount cells R2-C3, and their explicit link weights are extracted. Based on the numerical dependencies in the table relationship graph, for each explicitly associated table cell (e.g., R2-C3), the input parameter dependency chain of that cell is traced back, and its dependency chain is traversed in reverse. For example, if the formula for R2-C3 is "=R2-C1+R2-C2",... ” The dependency chain includes direct dependencies on R2-C1 and R2-C2, as well as indirect dependencies, such as dependencies of input parameters, like R2-C1 depending on R1-C1. A dependency level number d is generated. The initial cell for explicit links is at level d = 1. d increments by 1 for each level upstream in the dependency chain. For each dependency cell, an implicit weight is calculated based on the level number, with a decay factor of 0.8. The implicit weight is calculated as: Implicit weight = Explicit weight × 0.8^(d-1). For example, the explicit link weight of R2-C3 is 0.8, and the level d = 1. The level d corresponding to R2-C1 is 2. Therefore, the implicit link weight w corresponding to the dependency chain of R2-C1 is 0.8 × 0.8² - 1 = 0.64. Continuing upstream in the dependency chain of R2-C1 to R1-C1, the corresponding level d = 3, resulting in an implicit link weight w = 0.8 × 0.8³ - 1 = 0.512.
[0038] Extract the layout coordinates of dependent cells, such as the position of R2-C3 in the PDF (x1, y1, x2, y2), and extract the bounding box coordinates of image regions, such as the signature region (x1', y1', x2', y2'). Calculate the Intersection over Union (IoU), where IoU = overlap area / merged area, and retain only cell-image region pairs with an IoU greater than a second preset threshold. This second preset threshold can be set to 0.5. For association pairs that satisfy the IoU condition, use the IoU as the association strength, record the association strength, and use it as the layout association weight.
[0039] For each table cell in the dependency chain (e.g., R2-C1), if the cell contains a layout-related image area (e.g., a signature area below the column containing R2-C1), the cross-modal implicit link weight is calculated by multiplying the implicit link weight of the dependency chain by the layout-related weight. Assume the text clause "Total Price Calculation Method" is used. ” An explicit link is made to table cell R2-C3, with an explicit weight of 0.8. In this example, the implicit link is generated as follows, and the numerical dependency chain of R2-C3 is R2-C3. → R2-C1 (where level d = 2 and weight is 0.64), R2-C3 →R2-C2 (level d=2, weight 0.64). The signature area SIGN_AREA_5 exists below the layout coordinates of column (C1) containing R2-C1.
[0040] (IoU = 0.7), the final implicit link weight is calculated as: 0.64 × 0.7 = 0.448. The generated implicit links are as follows: Text Terms → R2-C1 → The signature area has a weight of 0.448.
[0041] Explicit links (text-table, text-image) and implicit links (table-image) are merged into a complete cross-modal association network. This network records the source-target modality, weights, and path information for each link; for example, explicit link paths are direct associations, while implicit link paths include dependency chain levels. This network provides structured input for subsequent unimodal self-consistency verification and cross-modal contradiction detection.
[0042] The above technical solution, through cross-modal projection and attention mechanisms, can capture deep semantic relationships between text descriptions, table values, and image elements. (Example: Text terms) " Interest rate fluctuations are based on Appendix Table 1 ” It not only links to the benchmark interest rate cell in the table, but also traces the formula chain that depends on that cell through implicit links, enabling multi-hop logic verification; by jointly calculating the numerical dependency chain and the intersection-combination ratio of the layout, it can automatically discover hidden cross-modal contradictions, such as when the text description " Quality acceptance requires signatures from both parties. ” At that time, an explicit link is used to connect to the acceptance date table cell, and an implicit link is used to locate the missing area of the signature image corresponding to that date, thus forming... " text → sheet → image ” The visualized contradictory paths provide clear evidence for manual review; the implicit link generation mechanism automatically updates the associated paths through numerical dependencies. If the down payment calculation formula is adjusted after the contract is revised (e.g., R2-C3 is changed to depend on R2-C4), the implicit links can be automatically reconstructed to avoid audit loopholes caused by outdated rule bases; through joint reasoning of explicit and implicit links, it can cover explicit requirements (such as signature integrity) and implicit risks (such as potential conflicts between formula logic and clause description) in legal clauses, achieving closed-loop auditing from data to knowledge.
[0043] Step S103: Based on multimodal feature data and cross-modal bidirectional link data, perform unimodal self-consistency verification and cross-modal contradiction detection on the target contract document to generate audit information for the target contract document.
[0044] Based on the multimodal feature data and the cross-modal bidirectional link data, the target contract document undergoes unimodal self-consistency verification and cross-modal contradiction detection to generate review information for the target contract document. Specifically, this includes: independently verifying each modality based on the multimodal feature data to determine the unimodal self-consistency verification data corresponding to each modality; performing cross-modal contradiction detection based on the cross-modal bidirectional link data to determine cross-modal contradictory data in the target contract document; and acquiring a preset legal knowledge graph to conduct a risk assessment on the unimodal self-consistency verification data and the cross-modal contradictory data based on the legal knowledge graph, determining the comprehensive risk index of the target contract document, and using the comprehensive risk index to determine the review information of the target contract document.
[0045] In one embodiment of this specification, a pre-constructed legal knowledge graph is obtained. This legal knowledge graph is a pre-constructed structured rule base, including hard rules such as mandatory legal provisions and industry filing standards, as well as soft rules such as industry practice quantification rules for ambiguous provisions. " Reasonable period ” The default number of days for different industries can also include historical risk patterns, such as features of contract dispute cases extracted from the China Judgments Online database. For text modality validation, a mandatory clause library from a legal knowledge graph is loaded. A named entity recognition model is used to extract key entities from the contract text, such as the names of the contracting parties and the amount in dispute, and these are matched one by one with the mandatory clauses, marking missing items. Simultaneously, a semantic similarity model is used to detect the logical consistency between contract clauses, such as identifying… " Force Majeure Exemption ” Terms and conditions " Unconditional performance ” Potential conflicts of clauses.
[0046] For the table modality, iterate through the formula cells in the table relationship graph, extract the input parameter values they depend on, recalculate the formula, verify whether the stored value matches the calculation result, and check the data type validity of numeric cells (such as date format, currency symbol). For example, if the formula in cell R2-C3 is... " =R2-C1*R2-C2 ” If the input parameters R2-C1 = 100 and R2-C2 = 50, the calculation result should be 5000; otherwise, it is marked as a calculation error. For image modalities, a pre-trained signature verification model is called to compare the similarity of the signature image with the handwriting features in the filing database. If the similarity is lower than a threshold (e.g., 0.8), it is marked as abnormal. Simultaneously, the seal area is extracted using image segmentation technology to verify its shape standardization and encoding clarity, ensuring compliance with legal requirements.
[0047] Based on this cross-modal bidirectional link data, cross-modal contradiction detection is performed to identify cross-modal contradictory data in the target contract document. Specifically, this includes: obtaining the explicit link set and implicit link set in the cross-modal bidirectional link data; calculating the absolute deviation between the text description values and the table stored values based on the explicit links in the explicit link set; marking it as an explicit contradiction if the deviation exceeds the dynamic threshold corresponding to the contract type of the target contract document; verifying whether the operational requirements of the text description are consistent with the content of the image elements based on the explicit links in the explicit link set; marking it as an explicit contradiction if the deviation is not consistent; and traversing the implicit link paths in the implicit link set to check whether there are cross-modal logical contradictions in the paths; if so, marking them as implicit contradiction association paths.
[0048] In one embodiment of this specification, text-table association pairs are extracted from an explicit set of links, such as... " 30% down payment ” In the terms and conditions and forms, the down payment cell is parsed to compare the numerical values in the text description with the values stored in the table, and the absolute deviation percentage is calculated. The absolute deviation percentage is calculated as follows: first, the difference between the text value and the table value is calculated; then, the absolute value of the ratio of this difference to the text value is used to determine the absolute deviation percentage. Thresholds are dynamically loaded based on the contract type, such as 1% for financial contracts, 5% for engineering contracts, and 10% for service contracts. If the deviation exceeds the threshold, it is marked as an explicit numerical contradiction. For text-image pairs, the operational requirements described in the text are verified to be consistent with the image content. For example, the text requires… " Both parties signed ” At that time, check whether the associated image area contains the signatures of both Party A and Party B. If the number of signatures or the identity of the signatory does not match, mark it as an explicit operational contradiction.
[0049] Traverse the associated paths (such as text terms) in the implicit link set. → Table Cell → The image region is examined to check for logical breaks in the path. For example, if a clause is explicitly linked to an acceptance date cell in a table, and that cell is implicitly linked to a signature area, but the signature for the corresponding date is missing in the image, this is marked as an implicit logical contradiction. The path weight decay is also analyzed; if the implicit link weight is below a preset threshold (e.g., 0.3), it is determined to be a low-confidence association path, indicating potential risk. Furthermore, based on the numerical dependency chains in the table's relationship graph, cross-modal chain contradictions caused by formula errors are detected, such as a contradiction between the associated text clause and the payment plan image caused by an error in the total price calculation.
[0050] Based on this legal knowledge graph, a risk assessment is conducted on the unimodal self-consistency verification data and the cross-modal contradictory data to determine the comprehensive risk index of the target contract document. Specifically, this includes: identifying table calculation errors, signature anomalies, and missing clause information in the unimodal self-consistency verification results; mapping these information to target rules in the legal knowledge graph to determine the unimodal risk index; obtaining explicit contradiction numerical deviations and implicit contradiction association paths in the cross-modal contradictory data; calculating the percentage deviation corresponding to the explicit contradiction numerical deviation to determine the explicit contradiction risk index; determining the average weight and path length corresponding to the implicit contradiction association path to assess its path vulnerability and determine the implicit contradiction risk index; and finally, using the unimodal risk index, the explicit contradiction risk index, and the implicit contradiction risk index, determining the comprehensive risk index of the target contract document.
[0051] In one embodiment of this specification, the results of unimodal self-consistency verification (such as table calculation errors, abnormal signature images, and missing text clauses) are mapped to rules in a legal knowledge graph. For example, in the table... " Total price = Unit price × Quantity ” The calculation error corresponds to "the method of performance needs to be clearly defined". ” Rules, marked as " Risks associated with monetary terms ” Risk weights are assigned based on error type; for example, a weight of 0.8 can be set for calculation errors, and 0.5 for formatting errors. Signature anomalies (such as insufficient clarity) are mapped to... " Reliable electronic signatures ” The rules stipulate that risk weights are calculated linearly based on the degree of anomaly; for example, a clarity score of 0.6 corresponds to a weight of 0.6. Missing clauses are directly matched against terms in the knowledge graph. " Essential elements of a contract ” The rules are set with the highest risk weight, such as 1.0. All unimodal risk indicators are weighted and aggregated according to the rule priority to generate a unimodal comprehensive risk score.
[0052] For explicitly contradictory data, such as text descriptions " 30% down payment ” With table storage values " The percentage deviation of 20% is calculated as Δ = 33.3%, and a threshold parameter is dynamically loaded based on the contract type. For example, if the deviation threshold for a financial contract is set to 1%, then the risk coefficient for the portion exceeding the threshold is Δ / threshold, i.e., 33.3% / 1% = 33.3. When the risk coefficient is greater than 1, it is always set to 1.0. This process yields explicit contradiction risk indicators. For implicit linking paths, such as text clauses... → Table Cell →For image regions, analyze the mean weight and length of the path. For example, if a path has a mean weight of 0.3 (out of 1.0) and a length of 4 layers, its vulnerability score is calculated as: Vulnerability = (1 - mean weight) × path length, i.e., (1 - 0.3) × 4 = 2.8. It should be noted that if the path contains high-risk nodes marked with legal knowledge graphs (such as...), the vulnerability score is calculated as follows: " liquidated damages clause ” The risk weight is increased by an additional factor, such as +0.2.
[0053] The unimodal risk score, explicit contradiction risk score, and implicit contradiction risk score are weighted and summed according to preset weights. The corresponding weight combination here could be 40% for unimodal, 30% for explicit contradiction, and 30% for implicit contradiction, generating a comprehensive risk index. For example, unimodal risk 0.7, explicit contradiction 0.9, and implicit contradiction 0.8. Based on preset thresholds (e.g., 0-0.4 low risk, 0.4-0.7 medium risk, 0.7-1.0 high risk), the final risk level is output, and a structured report containing legal basis, contradiction paths, and correction suggestions is generated.
[0054] Through the above technical solution, and by combining explicit and implicit link analysis, deep interactive verification of multimodal data is achieved. The implicit link path analysis in the embodiments of this specification can generate a visual chain of evidence; for example, when an anomaly in a payment plan is detected, it can be traced back to... " Contract text, payment terms, form formulas, signature date ” The complete correlation path clearly indicates the root cause of the contradiction (such as incorrect formula input parameters or missing signatures), and the source tracing mechanism significantly reduces the cost of manual review, especially suitable for evidence presentation scenarios in legal disputes; it dynamically loads threshold parameters according to contract type (such as strict thresholds for financial contracts and lenient thresholds for service contracts), and supports self-optimization of thresholds based on historical review data; for example, in engineering contract scenarios, the system can set a higher tolerance for numerical deviations in the bill of materials table, while in financial loan contracts, the threshold for interest rate calculation deviations approaches zero, thereby balancing review efficiency and risk control; the embodiments in this specification, through the joint detection of explicit and implicit contradictions, can simultaneously cover It covers both explicit legal requirements (such as signature integrity) and implicit industry rules (such as semantic consistency between formula logic and clause description); through the rule mapping mechanism of legal knowledge graph, it directly links single-modal errors (such as ambiguous signatures) with specific legal provisions, giving risk assessment a clear legal basis; by integrating three-dimensional indicators of single-modal, explicit, and implicit contradictions, it can cover the entire chain of risks from data anomalies to logical breaks, and identify compound risks through weighted calculations to avoid the one-sidedness of single-dimensional assessment; by dynamically loading thresholds and weight parameters by contract type (such as strict for financial contracts and lenient for engineering contracts), it avoids the over- or under-detection problems caused by traditional fixed thresholds.
[0055] Through the embodiments described in this specification, compared to the limitations of traditional contract review which can only handle single-modality or simply spliced multimodal data, this solution solves the problem of semantic fragmentation in heterogeneous data by constructing a unified representation space of textual semantic features, table relationship graphs, and image visual features. It establishes bidirectional links through a cross-modal attention mechanism, deeply binding the semantics of textual clauses with the numerical logic of tables and the layout features of images, thereby achieving full-dimensional analysis of contract elements and avoiding review blind spots caused by information fragmentation. Through a dual mechanism of single-modal self-consistency verification and cross-modal contradiction detection, it can discover independent problems such as formula calculation errors within tables and blurred image signatures, as well as capture explicit and implicit contradictions across modalities. This solution significantly reduces manual intervention through automated feature extraction and cross-modal reasoning, quickly focusing on high-risk items, avoiding the waste of resources in full document review, and improving contract review efficiency in batch contract processing scenarios.
[0056] This specification also provides an intelligent contract review device based on a multimodal large model, such as... Figure 2 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method.
[0057] This specification also provides a non-volatile computer storage medium storing computer-executable instructions configured to perform the above-described method.
[0058] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0059] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0060] The devices, media, and methods provided in the embodiments of this specification are one-to-one correspondences. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0061] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A smart contract review method based on a multimodal large model, characterized in that, The method includes: Obtain the target contract document to be reviewed, and parse the multimodal data in the target contract document, wherein the multimodal data includes text data, tabular data and image data; Feature extraction is performed on the multimodal data to obtain multimodal feature data. The multimodal feature data is then mapped to a unified dimensional space to calculate the association weights between each modality through a cross-modal attention mechanism, thereby establishing cross-modal bidirectional linked data. The multimodal feature data includes any one or more of the following: text semantic features, table relationship graphs, and image visual features. Based on the multimodal feature data and the cross-modal bidirectional link data, the target contract document is subjected to unimodal self-consistency verification and cross-modal contradiction detection to generate the review information of the target contract document, specifically including: Based on the multimodal feature data, each modality is independently verified to determine the single-modal self-consistency verification data corresponding to each modality; Based on the cross-modal bidirectional link data, cross-modal contradiction detection is performed to identify cross-modal contradiction data in the target contract document; The multimodal feature data is mapped to a unified dimensional space to calculate the association weights between each modality through a cross-modal attention mechanism, establishing cross-modal bidirectional data links, specifically including: The text semantic features, table relationship graphs, and image visual features in the multimodal feature data are projected onto a unified dimensional space through learnable linear transformation matrices to generate text projection vectors, table projection vectors, and image projection vectors. Determine the first cosine similarity score between the text projection vector and the table projection vector to filter association pairs whose first cosine similarity score is higher than a first preset threshold, establish explicit text-table links, and determine the text-table explicit link weights. Determine the second cosine similarity score between the text projection vector and the image projection vector, filter the association pairs whose second cosine similarity score is higher than the first preset threshold, establish explicit text-image links, and determine the text-image explicit link weights. The set of explicit links is determined by the explicit links in the text table, the explicit link weights in the text table, the explicit links in the text image, and the explicit link weights in the text image. Based on the numerical dependencies in the table relationship graph and the explicit link set, the implicit links and implicit link weights of the table image are determined to determine the implicit link set. Based on the numerical dependencies in the table relationship graph and the explicit link set, the implicit links and weights of the table images are determined, specifically including: Traverse the set of explicit links, filter the target explicit links whose target modality is a table modality, and obtain the associated table cell and target explicit weight corresponding to the target explicit link; Based on the numerical dependency relationship, at least one dependent cell of the associated table cell is traced backward, and the dependency level of each dependent cell is recorded, so as to determine the implicit weight corresponding to each dependent cell based on the dependency level and the target explicit weight. Obtain the cell layout coordinate data of each dependent cell and the image layout coordinate data of multiple image regions in the target contract document, and calculate the intersection-union ratio based on the cell layout coordinate data and the image layout coordinate data, and filter cell-image region pairs whose intersection-union ratio is greater than a second preset threshold; Based on the implicit weight of each dependent cell and the intersection-union ratio, the cross-modal implicit link weights corresponding to the cell-image region pairs are determined, and the implicit links of the table images are determined based on the cross-modal implicit link weights.
2. The intelligent contract review method based on a multimodal large model according to claim 1, characterized in that, Feature extraction is performed on the multimodal data to obtain multimodal feature data, specifically including: Embedsion vectors of a pre-trained language model in the legal domain are extracted from the text data to generate text semantic features; The table data is subjected to structured parsing to construct a corresponding table relationship graph, wherein the table relationship graph includes cell content and row and column relationships; The image data is subjected to denoising, orientation correction and resolution unification processing, and local visual features and global layout features are extracted through a multi-scale convolutional network.
3. The intelligent contract review method based on a multimodal large model according to claim 2, characterized in that, The table data is structured and parsed to construct a corresponding table relationship graph, specifically including: Obtain the logical row and column coordinates and cell content of each table cell in the table data. Use the logical row and column coordinates as keys and the cell content as values to construct the key-value pair content of each table cell. The cell content includes any one of the original value and the formula expression. When the cell content is a formula expression, based on the cell reference relationship in the formula expression, the target cell of the formula expression is used as the root node and the input parameter cells in the formula expression are used as leaf nodes to construct the numerical dependency relationship corresponding to the target cell. A row-column topology is constructed based on the coordinates of each table cell, wherein the row-column topology includes a set of cells in the same row and a set of cells in the same column; Based on the key-value pairs, the row-column topology, and the numerical dependencies, determine the table relationship graph corresponding to the table data.
4. The intelligent contract review method based on a multimodal large model according to claim 1, characterized in that, Based on the multimodal feature data and the cross-modal bidirectional link data, the target contract document is subjected to unimodal self-consistency verification and cross-modal contradiction detection to generate the review information of the target contract document, specifically including: A preset legal knowledge graph is obtained, and based on the legal knowledge graph, a risk assessment is performed on the single-modal self-consistency verification data and the cross-modal contradictory data to determine the comprehensive risk index of the target contract document, so as to determine the review information of the target contract document through the comprehensive risk index.
5. The intelligent contract review method based on a multimodal large model according to claim 4, characterized in that, Based on the aforementioned cross-modal bidirectional link data, cross-modal contradiction detection is performed to identify cross-modal contradiction data in the target contract document, specifically including: Obtain the explicit link set and implicit link set from the cross-modal bidirectional link data; Based on the explicit links in the text table in the explicit link set, calculate the absolute deviation between the text description value and the table stored value. If it exceeds the dynamic threshold corresponding to the contract type of the target contract document, it is marked as an explicit contradiction. Based on the explicit links between text and images in the set of explicit links, verify whether the operational requirements of the text description are consistent with the content of the image elements. If not, mark it as an explicit contradiction. Traverse the implicit link paths in the implicit link set and check if there are cross-modal logical contradictions in the paths. If so, mark them as implicit contradiction association paths.
6. The intelligent contract review method based on a multimodal large model according to claim 4, characterized in that, Based on the legal knowledge graph, a risk assessment is performed on the unimodal self-consistency verification data and the cross-modal contradictory data to determine the comprehensive risk index of the target contract document, specifically including: The table calculation error information, signature anomaly information, and clause missing information in the unimodal self-consistency verification result are identified, and the table calculation error information, signature anomaly information, and clause missing information are mapped to the target rules in the legal knowledge graph to determine the unimodal risk indicator; Obtain the explicit contradiction numerical deviation and implicit contradiction association path in the cross-modal contradiction data, calculate the deviation percentage corresponding to the explicit contradiction numerical deviation, and determine the explicit contradiction risk index. The mean weight and path length corresponding to the implicit contradiction association path are determined in order to assess the path vulnerability of the implicit contradiction association path and determine the implicit contradiction risk index. The comprehensive risk index of the target contract document is determined by using the single-modal risk index, the explicit contradiction risk index, and the implicit contradiction risk index.
7. A smart contract review device based on a multimodal large model, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-6.
8. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are configured to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Form information acquisition method and device and server
CN113011144A
Intelligent contract image recognition and contract element extraction method and device
CN114758341A