Intelligent contract auditing method and device based on multi-modal large model and medium
Through multimodal large-scale models, cross-modal correlations between text, tables and images are constructed, which solves the problem of splitting modal processing capabilities in smart contract audits, and realizes full-dimensional contract audits, improving audit efficiency and accuracy.
Patent Information
- Application Number
- CN202510635437.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing intelligent contract audit system has risks and hidden dangers in the contract audit results due to the fragmented modal processing capabilities and the lack of correlation between cross-modal logic.
A multimodal large model is used for feature extraction, and the correlation between text, tables and images is established through the cross-modal attention mechanism, cross-modal bidirectional link data is constructed, and single-modal self-consistent verification and cross-modal contradiction detection are performed.
It realizes full-dimensional analysis of contract documents, automatically discovers independent problems such as formula calculation errors in the table, blurred image signatures, and captures explicit and implicit contradictions across modalities, reduces manual intervention, and improves contract review efficiency.
Smart Images

Figure CN120524239A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of contract audit technology, and in particular to a smart contract audit method, device, and medium based on a multimodal large model. Background Art
[0002] Currently, existing technologies for smart contract review systems primarily rely on single-modal data processing techniques, such as text analysis models based on natural language processing (NLP) or seal recognition tools based on computer vision (CV). However, real-world contract documents typically contain multimodal data, including text, tables, and images (such as signatures and seals). While these technologies can perform grammatical checking of text clauses or simple classification of image elements in specific scenarios, they suffer from fragmented modal processing capabilities and a lack of cross-modal logical connections.
[0003] Traditional systems usually only process a single data type (such as plain text or independent images) and are unable to effectively integrate heterogeneous information such as text, tables, and images in contract documents. For example, they are unable to effectively process numerical logic such as formula calculation chains in tables and key information such as signature handwriting and seal codes in images, resulting in the omission of a large amount of potential risks of non-text data during the review process. Although existing solutions attempt to integrate multimodal data, they mostly use simple modal splicing or manual rules to define cross-modal associations, and only achieve shallow data splicing. For example, after identifying the content of a table through OCR and comparing it with the text terms independently, they are unable to capture the deep semantic association between text descriptions, table values, and image elements. For example, when the text terms stipulate " The down payment ratio needs to be signed and confirmed by both parties ” It is difficult to automatically verify whether the clause is related to the specific amount in the form and the signature area on the signature page.
[0004] Therefore, in the review process of contract documents with multiple modalities such as tables, the existing technology uses single-modal data analysis and shallow data splicing, which leads to the problems of fragmented modal processing capabilities and lack of cross-modal logical associations, resulting in risks in the contract review results. Summary of the Invention
[0005] One or more embodiments of this specification provide a smart contract review method, device and medium based on a multimodal large model, which is used to solve the following technical problems: in the review process of contract documents with multiple modalities such as tables, because the existing technology adopts single modal data analysis and shallow data splicing, there are problems of modal processing capability fragmentation and lack of cross-modal logical association, which leads to risks in the contract review results.
[0006] One or more embodiments of this specification adopt the following technical solutions:
[0007] One or more embodiments of the present specification provide a smart contract audit method based on a multimodal large model, the method comprising: obtaining a target contract document to be audited, parsing multimodal data in the target contract document, wherein the multimodal data comprises text data, table data, and image data; performing feature extraction on the multimodal data to obtain multimodal feature data, mapping the multimodal feature data to a unified dimensional space, calculating the association weight between each modality through a cross-modal attention mechanism, and establishing cross-modal bidirectional link data, wherein the multimodal feature data comprises any one or more of text semantic features, table relationship graphs, and image visual features; performing unimodal self-consistency verification and cross-modal contradiction detection on the target contract document based on the multimodal feature data and the cross-modal bidirectional link data to generate audit information for the target contract document.
[0008] One or more embodiments of this specification provide a smart contract auditing device based on a multimodal large model, including:
[0009] at least one processor; and,
[0010] a memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0012] One or more embodiments of this specification provide a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above method.
[0013] At least one of the above-mentioned technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: Through the embodiments of this specification, compared with the limitations of traditional contract review that can only process a single modality or simply splice multimodal data, this solution solves the semantic fault problem of heterogeneous data by constructing a unified representation space of text semantic features, table relationship maps and image visual features, establishes a two-way link through a cross-modal attention mechanism, and deeply binds the semantics of text terms with the numerical logic of the table and the layout features of the image, thereby realizing a full-dimensional analysis of contract elements and avoiding audit blind spots caused by information fragmentation; through the dual mechanisms of single-modal self-consistency verification and cross-modal contradiction detection, it can not only discover independent problems such as formula calculation errors in the table and blurred image signatures, but also capture explicit and implicit contradictions across modalities; through automated feature extraction and cross-modal reasoning, this party greatly reduces manual intervention, quickly focuses on high-risk items, avoids the waste of resources in full-document review, and improves the efficiency of contract review in batch contract processing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:
[0015] Figure 1 A flowchart of a smart contract audit method based on a multimodal large model provided in an embodiment of this specification;
[0016] Figure 2 A schematic diagram of the structure of a smart contract audit device based on a multimodal large model provided in an embodiment of this specification. DETAILED DESCRIPTION
[0017] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification without creative work should fall within the scope of protection of this specification.
[0018] The embodiments of this specification provide a smart contract audit method based on a multimodal large model. It should be noted that the execution entity in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 A flowchart of a smart contract audit method based on a multimodal large model is provided in the embodiment of this specification, such as Figure 1 As shown, it mainly includes the following steps:
[0019] Step S101: Obtain a target contract document to be reviewed and parse the multimodal data in the target contract document.
[0020] Wherein, the multimodal data includes text data, table data and image data;
[0021] In one embodiment of the present specification, contract documents uploaded by users are received through a standardized interface, and multiple file formats including PDF, Word, and scanned images are supported. For electronic documents (such as PDF), a layout parsing engine is used to extract the original structural information of the document, including page layout, text paragraph coordinates, table border position, and image embedding area. For scanned or photographed files, image preprocessing operations are first performed, including denoising correction, resolution unification, and direction rotation to ensure the accuracy of subsequent OCR recognition. After completing the format parsing, the multimodal data in the target contract document is parsed, where the multimodal data includes text data, table data, and image data.
[0022] In step S102, feature extraction is performed on the multimodal data to obtain multimodal feature data, and the multimodal feature data is mapped to a unified dimensional space to calculate the association weight between each modality through a cross-modal attention mechanism to establish cross-modal bidirectional link data.
[0023] The multimodal feature data includes any one or more of text semantic features, table relationship graphs, and image visual features;
[0024] Feature extraction is performed on the multimodal data to obtain multimodal feature data, specifically including: extracting embedding vectors of a pre-trained language model in the legal field on the text data to generate text semantic features; performing structured parsing on the table data to construct a corresponding table relationship map, wherein the table relationship map includes cell content and row and column relationships; performing denoising, direction correction, and resolution unification processing on the image data, and extracting local visual features and global layout features through a multi-scale convolutional network.
[0025] In one embodiment of the present specification, the text data is subjected to legal field text semantic feature extraction through a multimodal large model. For the contract text data, a language model pre-trained in the legal field (such as a domain adaptation model based on the BERT architecture) is first loaded. Based on the general corpus training, the model is incrementally pre-trained through professional data such as legal documents and contract cases to enhance the semantic understanding of legal terms and clause structures. After the input text is segmented, a context-related embedding vector is generated for each token, and sentence-level or paragraph-level semantic features are aggregated through a pooling layer. For contract-specific clause types (such as liability for breach of contract, dispute resolution), the model additionally introduces an attention mask mechanism to focus on capturing the contextual relevance of key clauses, and ultimately outputs a text feature vector with legal semantics. Compared with the general pre-training model, the semantic extraction model adapted to the legal field can effectively distinguish the ambiguity of professional terms in contract terms. As " force majeure ” In the terms " Natural disasters ”and insurance terms " Natural disasters ” Although the words are the same, the domain pre-trained model can generate differentiated embedding vectors, avoiding the misjudgment of terms caused by semantic confusion in traditional methods. At the same time, the attention mask mechanism strengthens the context dependency of the core terms, making " Calculation method of liquidated damages ” The semantics of complex expressions such as .
[0026] A standardized preprocessing process is performed on the raw image data, using Gaussian filtering and edge-preserving algorithms to remove noise and moiré artifacts from the scanned document. Subsequently, an orientation detection model (such as a convolutional network based on text line tilt prediction) automatically corrects image rotation to ensure horizontal alignment of text lines. Finally, the image resolution is uniformly scaled to a preset size to eliminate scale deviations caused by device differences. The preprocessed image is then fed into a multi-scale convolutional network. The shallow network uses small convolution kernels to extract local detail features (such as the continuity of the signature and the sharpness of the seal edge), while the deep network uses large receptive field convolutions to capture global layout features (such as the relative position of the signature area on the page and the spatial distribution of multiple seals). The local and global features are weighted by channel attention and concatenated into a joint visual feature vector. Traditional image feature extraction methods typically use only a single-scale convolutional network, which struggles to balance the correlation between local details and global layout. By integrating hierarchical features in a multi-scale network, the model can both recognize microscopic signature features (such as the uniqueness of the strokes) and perceive the macroscopic spatial relationship between the signature and seal (such as the compliance of the signature position). For example, when identifying electronic seals, local features can determine the clarity of the code, and global features can verify whether the seal is located in the designated area for contract signing. The combination of the two significantly improves the comprehensiveness of the review.
[0027] Performing structured parsing on the table data and constructing a corresponding table relationship graph, specifically including: obtaining the logical row and column coordinates and cell content of each table cell in the table data, using the logical row and column coordinates as the key and the cell content as the value to construct a key-value pair content of each table cell, wherein the cell content includes any one of the original value and the formula expression; when the cell content is a formula expression, based on the cell reference relationship in the formula expression, using the target cell of the formula expression as the root node and the input parameter cell in the formula expression as the leaf node, constructing a numerical dependency relationship corresponding to the target cell; constructing a row and column topological relationship based on the coordinates of each table cell, wherein the row and column topological relationship includes a set of cells in the same row and a set of cells in the same column; determining the table relationship graph corresponding to the table data based on the key-value pair, the row and column topological relationship and the numerical dependency relationship.
[0028] In one embodiment of the present specification, first, the physical structure of the input table data is parsed to identify the boundary position of each cell and assign logical coordinates in row and column order. For example, the first cell in the upper left corner of the table is marked as " R1-C1 ” (first row, first column), with row and column numbers increasing downwards and to the right. The contents of each cell are extracted as a raw value or a formula expression. The contents of each cell are converted to a key-value pair format, where the key is the logical row and column coordinates. " R row number - C column number ” , the value is text, numeric value or formula expression starting with equal sign after removing irrelevant symbols;
[0029] For cells containing formulas such as " =R1-C1+R1-C2 ” ), retain the formula text intact, and separate the target cell (cell that stores the formula result) and the input parameter cell (other cells referenced in the formula) through the syntax parser. For the formula cell, take the cell as the root node and the input parameter cell it references as the leaf node to establish a tree-like dependency relationship. The specific construction process is as follows: use regular expressions to match cell references in the formula, extract the logical row and column coordinates of the referenced cell; decompose the formula into an operator sequence and an operand sequence, the operands include constant values and cell references; create a dependency tree node for each formula cell, and create a child node for each referenced input parameter cell; store the node relationship and operator type of the dependency tree in the dependency_graph field of the table structured JSON data. For example, if the cell " R3-C3 ” The formula is " =R1-C1*R2-C2 ” , then the root node is " R3-C3 ” , the leaf nodes are " R1-C1 ” and "R2-C2 ” The dependency will be traced layer by layer. If the input parameter cell itself contains a formula (such as "R1-C1=R1-C3+R1-C4 ” ), the hierarchy of the dependency tree is further expanded to form a complete numerical calculation chain.
[0030] Using the logical coordinates of each table cell, all cells in the same row are merged into a row cell set, and all cells in the same column are merged into a column cell set. For example, the cell "R2-C1 ” The peer set of " R2-C1, R2-C2, R2-C3…” , the same column set is " R1-C1, R2-C1, R3-C1 …” This topological relationship not only records the physical adjacency of cells, but can also be used for subsequent cross-row and cross-column data verification, such as the numerical consistency between the total row and the item row.
[0031] The three types of data, key-value pairs (cell coordinates and content), numerical dependency trees, and row and column topological relationships, are integrated into a graph structure. Nodes in the graph represent cells, containing original values or formula type attributes; edges are divided into three categories: numerical dependency edges (root nodes point to leaf nodes), peer edges (cells in the same row are bidirectionally connected), and column edges (cells in the same column are bidirectionally connected). Node attributes include cell keys, values, and types; edge relationships include row_adjacent / col_adjacent type edges that represent physical adjacency, and dependency type edges that represent numerical dependencies. The graph structure is output as a table relationship graph to subsequent modules for associating text clauses with table numerical logic. It should be noted that this graph fully represents the numerical logic, physical layout, and data relevance of the table, providing structured input for cross-modal review.
[0032] Through the above technical solution, conventional methods only extract the original values of the table, ignore the formula dependency and row and column topology, resulting in the inability to detect numerical calculation errors or logical contradictions in subsequent audits. The embodiment of this specification can trace the complete input chain of formula calculations by constructing a numerical dependency tree; for example, when an abnormal value in a cell is detected, the error source of the specific input parameter can be located in reverse along the dependency tree (such as a parameter cell is mistakenly filled with text), which significantly improves the efficiency of problem locating; through row and column topology relationship modeling, it can automatically check whether the total value of the same row matches the sum of the item values, whether the data type of the same column is consistent, etc.; the numerical dependency tree and graph structure of the embodiment of this specification make text terms (such as " Interest rate fluctuation is calculated according to Appendix Table 1 ” ) can be deeply bound to the formula calculation chain in the table (such as interest rate = base rate * floating coefficient). When the floating rule described in the text conflicts with the formula logic, the conflicting node can be quickly located along the dependency tree to achieve accurate risk tracing; through the joint mapping of logical coordinates and physical layout, the original coverage of the merged cells (such as "R1-C1:R2-C2 ” Indicates a merged area spanning two rows and two columns) and automatically continues broken rows and columns of a table that spans multiple pages to ensure data integrity.
[0033] The multimodal feature data is mapped to a unified dimensional space to calculate the association weight between each modality through a cross-modal attention mechanism, and establish cross-modal bidirectional link data, specifically including: projecting the text semantic features, table relationship map and image visual features in the multimodal feature data to a unified dimensional space through a learnable linear transformation matrix, generating a text projection vector, a table projection vector and an image projection vector; determining the first cosine similarity score between the text projection vector and the table projection vector to screen association pairs whose first cosine similarity score is higher than a first preset threshold, and establishing a text-table explicit link. And determine the text table explicit link weight; determine the second cosine similarity score between the text projection vector and the image projection vector, screen the association pairs whose second cosine similarity score is higher than the first preset threshold, establish the text image explicit link, and determine the text image explicit link weight; determine the explicit link set through the text table explicit link, the text table explicit link weight, the text image explicit link and the text image explicit link weight; based on the numerical dependency in the table relationship map and the explicit link set, determine the table image implicit link and the table image implicit link weight to determine the implicit link set.
[0034] In one embodiment of the present specification, the text semantic features, table relationship map and image visual features are first mapped to a unified dimensional space through an independent learnable linear transformation matrix. The text semantic features (such as 768-dimensional BERT embedding vector) are transformed into a text projection vector through a matrix, the cell features in the table relationship map (such as numerical dependency encoding) are mapped to a table projection vector, and the multi-scale visual features of the image (local details and global layout) are mapped to an image projection vector. The dimensions of all projection vectors remain consistent to ensure the comparability of cross-modal similarity calculations. For example, the text feature vector (768 dimensions) is mapped to the target dimension d (such as 512 dimensions) through a learnable linear transformation. The table feature vector (a 256-dimensional vector including cell values and topological relationship encoding) is mapped to the d dimension. The local and global image features are spliced (512+1024=1536 dimensions) and mapped to the d dimension.
[0035] Subsequently, the cosine similarity between the text projection vector and the table projection vector is calculated, and the association pairs with scores higher than the preset threshold are screened to establish a text-table explicit link. The first preset threshold here can be set according to actual needs. For example, through experimental verification, in the contract review scenario, a threshold of 0.7 can balance the recall rate and the precision rate, so it can be set to 0.7. Only association pairs with scores higher than 0.7 are retained to ensure that high-confidence associations are retained. Explicit weights are used to quantify the direct association strength between cross-modal elements. The calculation is based on the semantic similarity of the cross-modal feature vectors. The explicit weight is obtained by mapping the cosine similarity to the [0,1] interval, such as linearly mapping the similarity score to the weight range [0,1] by the formula w=(s+1) / 2. Similarly, the cosine similarity between the text and image projection vectors is calculated, high-scoring association pairs are screened, a text-image explicit link is established, and the explicit weight is calculated. For example, the text clause "the down payment ratio is 30% of the total contract price" is stored in table cells R2-C3. " Total price = 1 million ” , R2-C4 storage " Down payment = 300,000 ” The calculated similarity score between the text and R2-C3 is 0.9, with a weight of 0.95; the similarity score between the text and R2-C4 is 0.85, with a weight of 0.925. R2-C4 becomes the main associated object due to its higher weight. For example, the text clause " The contract must be signed by both parties to take effect ” , the image area contains the bounding box of the signatures of both parties A and B, and the similarity score is calculated to be 0.75 with a weight of 0.875.
[0036] Based on the numerical dependency relationship in the table relationship graph and the explicit link set, the table image implicit link and the table image implicit link weight are determined, specifically including: traversing the explicit link set, screening the target explicit link whose target modality is the table modality, to obtain the associated table cell and the target explicit weight corresponding to the target explicit link; based on the numerical dependency relationship, tracing back at least one dependent cell of the associated table cell, recording the number of dependency levels of each dependent cell, and determining the implicit weight corresponding to each dependent cell based on the number of dependency levels and the target explicit weight; obtaining the cell layout coordinate data of each dependent cell and the image layout coordinate data of multiple image areas in the target contract document, calculating the intersection-and-union ratio based on the cell layout coordinate data and the image layout coordinate data, and screening the cell-image area pairs whose intersection-and-union ratio is greater than a second preset threshold; determining the cross-modal implicit link weight corresponding to the cell-image area pair based on the implicit weight of each dependent cell and the intersection-and-union ratio, and determining the table image implicit link based on the cross-modal implicit link weight.
[0037] In one embodiment of the present specification, a traversal is performed in the explicit link set to filter out links with table mode as the target, such as text-table explicit links, and the table cells associated with the explicit link, such as the down payment amount cell R2-C3, and their display link weights are extracted. According to the numerical dependency relationship in the table relationship graph, for each explicitly associated table cell (such as R2-C3), the input parameter dependency chain of the cell is traced back and its dependency chain is traversed in reverse. For example, if the formula of R2-C3 is "=R2-C1+R2-C2 ” , the dependency chain includes direct dependencies on R2-C1 and R2-C2, as well as indirect dependencies, such as the dependent cells of input parameters, for example, R2-C1 depends on R1-C1. Generate the number of dependency levels d. The initial cell of the explicit link is level d = 1. Each level up the dependency chain increases by 1. For each dependent cell, an implicit weight is calculated based on the number of levels, with a decay factor of 0.8: implicit weight = explicit weight × 0.8^(d-1). For example, if R2-C3 has an explicit link weight of 0.8 and level d = 1, and R2-C1 has a corresponding level d = 2, then the implicit link weight of the R2-C1 dependency chain is w = 0.8 × 0.82 - 1 = 0.64. If R2-C1's dependency chain continues up to R1-C1, corresponding to level d = 3, the resulting implicit link weight is w = 0.8 × 0.83 - 1 = 0.512.
[0038] Extract the layout coordinates of the dependent cells, such as the position of R2-C3 in the PDF (x1, y1, x2, y2). Also extract the bounding box coordinates of the image region, such as the signature region's (x1', y1', x2', y2'). Calculate the Intersection over Union (IoU), where IoU = overlapping area / combined area. Only cell-image region pairs with an IoU greater than a second preset threshold are retained. This second preset threshold can be set to 0.5. For association pairs that meet the IoU condition, the IoU is used as the association strength, and the association strength is recorded as the layout association weight.
[0039] For each table cell in the dependency chain (such as R2-C1), if the cell has a layout-associated image area (such as a signature area under the column where R2-C1 is located), the cross-modal implicit link weight is calculated by multiplying the implicit link weight corresponding to the dependency chain and the layout-associated weight. ” Explicitly link to table cell R2-C3, where the explicit weight is 0.8. In this example, the implicit link generation process is as follows, and the numerical dependency chain of R2-C3 is R2-C3 → R2-C1 (where level d=2 and weight is 0.64), R2-C3 →R2-C2 (level d=2, weight 0.64). There is a signature area SIGN_AREA_5 below the layout coordinates of the column (C1) where R2-C1 is located.
[0040] (IoU=0.7), calculate the final implicit link weight: 0.64×0.7=0.448. Generate implicit links as follows: Text terms → R2-C1 → Signature area, weight 0.448.
[0041] Explicit links (text-table, text-image) and implicit links (table-image) are combined into a complete cross-modal association network. This network records the source-target modality, weight, and path information of each link. For example, explicit link paths are direct associations, while implicit link paths include dependency chain hierarchies. This network provides structured input for subsequent unimodal consistency verification and cross-modal contradiction detection.
[0042] Through the above technical solution, through cross-modal projection and attention mechanism, it is possible to capture the deep semantic associations between text descriptions, table values and image elements. " Interest rate fluctuation is based on Appendix Table 1 ” It not only associates the base interest rate cell in the table, but also traces back to the formula chain that depends on the cell through implicit links, realizing multi-hop logic verification; through the joint calculation of the numerical dependency chain and the layout intersection and union ratio, it can automatically discover hidden cross-modal contradictions, such as when the text description " Quality acceptance requires signatures from both parties ” When the signature image is missing, it is associated with the acceptance date table cell through an explicit link, and then located to the signature image missing area corresponding to the date through an implicit link, forming " text → sheet → image ” The visual contradiction path provides clear evidence for manual review; the implicit link generation mechanism automatically updates the associated path through numerical dependencies. If the down payment calculation formula is adjusted after the contract is revised (such as R2-C3 is changed to rely on R2-C4), the implicit link can be automatically rebuilt to avoid audit loopholes caused by the failure to update the rule base; through the joint reasoning of explicit and implicit links, it can cover the explicit requirements (such as signature integrity) and implicit risks (such as potential conflicts between formula logic and clause descriptions) in legal terms, and realize closed-loop audit from data to knowledge.
[0043] Step S103 : Based on the multimodal feature data and the cross-modal bidirectional link data, single-modal self-consistency verification and cross-modal contradiction detection are performed on the target contract document to generate audit information of the target contract document.
[0044] Based on the multimodal feature data and the cross-modal bidirectional link data, the target contract document is subjected to single-modal self-consistency verification and cross-modal contradiction detection to generate audit information of the target contract document, specifically including: independently verifying each modality according to the multimodal feature data to determine the single-modal self-consistency verification data corresponding to each modality; performing cross-modal contradiction detection based on the cross-modal bidirectional link data to determine the cross-modal contradiction data in the target contract document; obtaining a preset legal knowledge graph to perform risk assessment on the single-modal self-consistency verification data and the cross-modal contradiction data based on the legal knowledge graph to determine the comprehensive risk index of the target contract document, so as to determine the audit information of the target contract document through the comprehensive risk index.
[0045] In one embodiment of this specification, a pre-built legal knowledge graph is obtained. The legal knowledge graph is a pre-built structured rule base, including hard rules such as legal mandatory clauses and industry filing standards, as well as soft rules such as industry practice quantitative rules for ambiguous clauses, such as " Reasonable period ” The default number of days in different industries can also include historical risk patterns, such as the characteristics of contract dispute cases extracted from the Judgment Documents Network. For text modality verification, the mandatory clause library in the legal knowledge graph is loaded, and the key entities in the contract text, such as the names of the contracting parties and the target amount, are extracted through the named entity recognition model, and matched one by one with the mandatory clauses, marking the missing items. At the same time, the semantic similarity model is used to detect the logical consistency between contract clauses, such as identifying " Force Majeure Exemption ” Terms and " Unconditional performance ” Potential Conflicts of Terms.
[0046] For table mode, traverse the formula cells in the table relationship graph, extract the input parameter values they depend on, recalculate the formula, verify whether the stored value is consistent with the calculation result, and check the data type legitimacy of the value cell (such as date format, currency symbol). For example, if the formula in cell R2-C3 is " =R2-C1*R2-C2 ” , if the input parameters R2-C1=100 and R2-C2=50, the calculated result should be 5000; otherwise, it is marked as a calculation error. For image modalities, a pre-trained signature verification model is called to compare the signature image with the handwriting features in the filing database. If the similarity is below a threshold (such as 0.8), it is marked as an anomaly. At the same time, image segmentation technology is used to extract the seal area, verifying its shape standardization and coding clarity to ensure compliance with legal requirements.
[0047] Based on the cross-modal bidirectional link data, cross-modal contradiction detection is performed to determine the cross-modal contradiction data in the target contract document, specifically including: obtaining the explicit link set and the implicit link set in the cross-modal bidirectional link data; based on the text table explicit link in the explicit link set, calculating the absolute deviation between the text description value and the table storage value, if it exceeds the dynamic threshold corresponding to the contract type of the target contract document, it is marked as an explicit contradiction; based on the text image explicit link in the explicit link set, verifying whether the operation requirements of the text description are consistent with the content of the image element, if not, it is marked as an explicit contradiction; traversing the implicit link path in the implicit link set, checking whether there is a cross-modal logical contradiction in the path, if so, marking it as an implicit contradiction association path.
[0048] In one embodiment of the present specification, text-table association pairs are extracted from the explicit link set, such as " 30% down payment ” For the down payment cell in the terms and tables, parse the numerical value of the text description and the table storage value, and calculate the absolute deviation percentage. The absolute deviation percentage here is calculated as follows: first calculate the difference between the text value and the table value, and determine the absolute deviation percentage based on the absolute value of the ratio of this difference to the text value. Dynamically load the threshold according to the contract type, such as the deviation threshold for financial contracts is 1%, the deviation threshold for engineering contracts is 5%, and the deviation threshold for service contracts is 10%. If the deviation exceeds the threshold, it is marked as an explicit numerical contradiction. For text-image association pairs, verify whether the operation requirements of the text description are consistent with the image content, such as the text requirements " Both parties signed ” When , check whether the associated image area contains the signatures of both parties A and B. If the number of signatures or the identity of the signers do not match, it is marked as an explicit operation contradiction.
[0049] Traverse the associated paths in the implicit link collection (such as text terms → Table cells → Image area), check whether there is a logical fault in the path. For example, a clause is linked to the acceptance date cell in the table through an explicit link, and the cell is linked to the signature area through an implicit link, but the signature of the corresponding date in the image is missing, which is marked as an implicit logical contradiction. At the same time, the path weight decay is analyzed. If the implicit link weight is lower than the preset threshold (such as 0.3), it is determined to be a low-confidence association path, indicating potential risks. In addition, based on the numerical dependency chain in the table relationship graph, cross-modal chain contradictions caused by formula errors are detected, such as the contradiction between the associated text clause and the payment plan image caused by an error in the total price calculation.
[0050] Based on the legal knowledge graph, a risk assessment is performed on the unimodal self-consistency verification data and the cross-modal contradiction data to determine the comprehensive risk index of the target contract document, specifically including: determining the table calculation error information, signature exception information and clause missing information in the unimodal self-consistency verification result, mapping the table calculation error information, the signature exception information and the clause missing information to the target rules in the legal knowledge graph, and determining the unimodal risk index; obtaining the explicit contradiction numerical deviation and the implicit contradiction association path in the cross-modal contradiction data, calculating the deviation percentage corresponding to the explicit contradiction numerical deviation, and determining the explicit contradiction risk index; determining the weight mean and path length corresponding to the implicit contradiction association path, and evaluating the path vulnerability of the implicit contradiction association path to determine the implicit contradiction risk index; determining the comprehensive risk index of the target contract document through the unimodal risk index, the explicit contradiction risk index and the implicit contradiction risk index.
[0051] In one embodiment of this specification, the results of the unimodal self-consistency verification (such as table calculation errors, signature image anomalies, and missing text clauses) are mapped to the rules in the legal knowledge graph. For example, in the table " Total price = unit price × quantity ” The calculation error corresponds to "the method of performance must be clear ” Rules, marked as " Amount clause risk ” , and assign risk weights based on the error type. You can set the weight of calculation errors to 0.8 and the weight of format errors to 0.5. Signature anomalies (such as lack of clarity) are mapped to " Reliable electronic signature ” Rules, risk weight is calculated linearly based on the degree of abnormality, such as a clarity score of 0.6 corresponds to a weight of 0.6. If the terms are missing, they are directly matched in the knowledge graph. " Essential elements of a contract ” For each rule, the risk weight is set to the highest level, such as 1.0. All single-modal risk indicators are weighted and aggregated according to the rule priority to generate a single-modal comprehensive risk score.
[0052] For explicitly contradictory data, such as text descriptions " 30% down payment ” Storing values with tables " 20%", calculate its numerical deviation percentage Δ = 33.3%, and dynamically load the threshold parameter according to the contract type. For example, if the deviation threshold of a financial contract is set to 1%, the risk coefficient of the part of Δ that exceeds the threshold is Δ / threshold, that is, 33.3% / 1% = 33.3. When the risk coefficient is greater than 1, it is taken as 1.0. The explicit contradiction risk index is obtained in turn. For implicit link paths, such as text clauses → Table cells →Image area, analyze the path weight mean and length. For example, a path weight mean is 0.3 (full score 1.0) and the length is 4 layers. Its vulnerability score calculation formula is: vulnerability = (1-weight mean) × path length, that is, (1-0.3) × 4 = 2.8. It should be noted that if there are high-risk nodes marked by the legal knowledge graph in the path (such as " Liquidated damages clause ” ), the risk weight is additionally increased, such as +0.2.
[0053] The unimodal risk score, explicit contradiction risk score, and implicit contradiction risk score are weighted and summed according to preset weights. Here, the corresponding weight combination might be 40% for unimodal risk, 30% for explicit contradiction, and 30% for implicit contradiction, to generate a comprehensive risk index. For example, unimodal risk is 0.7, explicit contradiction is 0.9, and implicit contradiction is 0.8. The final risk level is output based on preset thresholds (e.g., 0-0.4 for low risk, 0.4-0.7 for medium risk, and 0.7-1.0 for high risk). A structured report is also generated, including the legal basis, contradiction paths, and corrective measures.
[0054] Through the above technical solution, through the joint analysis of explicit and implicit links, deep interactive verification of multimodal data is achieved; the implicit link path analysis of the embodiment of this specification can generate a visual evidence chain, for example, when an abnormal payment plan is detected, it can be traced back to " Contract text, payment terms, table formulas, signature date ” The complete association path of the contract is clearly indicated, and the root cause of the contradiction (such as incorrect formula input parameters or missing signatures) is clearly indicated. The traceability mechanism greatly reduces the cost of manual review, which is especially suitable for evidence-gathering scenarios in legal disputes. The threshold parameters are dynamically loaded according to the contract type (such as strict thresholds for financial contracts and loose thresholds for service contracts), and support self-optimization of thresholds based on historical audit data. For example, in the engineering contract scenario, the system can set a higher tolerance for numerical deviations in the bill of materials table, while in financial loan contracts, the threshold for interest rate calculation deviations approaches zero, thereby balancing audit efficiency and risk control. The embodiments of this specification can simultaneously cover It covers explicit legal requirements (such as signature integrity) and implicit industry rules (such as semantic consistency between formula logic and clause descriptions); through the rule mapping mechanism of the legal knowledge graph, single-modal errors (such as signature ambiguity) are directly linked to specific legal provisions, so that risk judgment has a clear legal basis; through the fusion of three-dimensional indicators of single modality, explicit and implicit contradictions, it can cover the entire chain of risks from data anomalies to logical faults, identify complex risks through weighted calculations, and avoid the one-sidedness of single-dimensional assessments; through dynamic loading of thresholds and weight parameters by contract type (such as strict financial contracts and loose engineering contracts), it avoids over-inspection or missed inspections caused by traditional fixed thresholds.
[0055] Through the embodiments of this specification, compared with the limitations of traditional contract review that can only process a single modality or simply splice multimodal data, this solution solves the semantic fault problem of heterogeneous data by constructing a unified representation space of text semantic features, table relationship maps and image visual features. Through the cross-modal attention mechanism, a two-way link is established, and the semantics of the text terms are deeply bound with the numerical logic of the table and the layout features of the image, thereby achieving full-dimensional analysis of contract elements and avoiding audit blind spots caused by information fragmentation; through the dual mechanisms of single-modal self-consistency verification and cross-modal contradiction detection, it can not only discover independent problems such as formula calculation errors in the table and blurred image signatures, but also capture explicit and implicit contradictions across modalities; through automated feature extraction and cross-modal reasoning, this party greatly reduces manual intervention, quickly focuses on high-risk items, avoids the waste of resources in full document review, and improves the efficiency of contract review in batch contract processing scenarios.
[0056] The embodiment of this specification also provides a smart contract audit device based on a multimodal large model, such as Figure 2 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0057] The embodiments of this specification also provide a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method.
[0058] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0059] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0060] The devices and media provided in the embodiments of this specification correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0061] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.
Claims
1. A smart contract audit method based on a multimodal large model, characterized in that: The method comprises: Obtaining a target contract document to be reviewed, and parsing multimodal data in the target contract document, wherein the multimodal data includes text data, table data, and image data; Performing feature extraction on the multimodal data to obtain multimodal feature data, mapping the multimodal feature data to a unified dimensional space, calculating the association weight between each modality through a cross-modal attention mechanism, and establishing cross-modal bidirectional link data, wherein the multimodal feature data includes any one or more of text semantic features, table relationship maps, and image visual features; Based on the multimodal feature data and the cross-modal bidirectional link data, single-modal self-consistency verification and cross-modal contradiction detection are performed on the target contract document to generate review information of the target contract document.
2. The method for smart contract auditing based on a multimodal large model according to claim 1 is characterized in that: Extracting features from the multimodal data to obtain multimodal feature data specifically includes: Extracting embedding vectors of a pre-trained language model in the legal field from the text data to generate text semantic features; Performing structured analysis on the table data to construct a corresponding table relationship map, wherein the table relationship map includes cell content and row and column relationships; De-noising, direction correction and resolution unification are performed on the image data, and local visual features and global layout features are extracted through a multi-scale convolutional network.
3. The method for smart contract auditing based on a multimodal large model according to claim 2 is characterized in that: Performing structured analysis on the table data to construct a corresponding table relationship map, specifically including: Obtaining the logical row and column coordinates and cell content of each table cell in the table data, and constructing a key-value pair for each table cell using the logical row and column coordinates as a key and the cell content as a value, wherein the cell content includes any one of an original value and a formula expression; When the cell content is a formula expression, based on the cell reference relationship in the formula expression, the target cell of the formula expression is used as the root node, and the input parameter cells in the formula expression are used as leaf nodes to construct a value dependency relationship corresponding to the target cell; Constructing a row-column topological relationship based on the coordinates of each table cell, wherein the row-column topological relationship includes a set of cells in the same row and a set of cells in the same column; A table relationship graph corresponding to the table data is determined based on the key-value pairs, the row-column topological relationship, and the numerical dependency relationship.
4. The method for smart contract auditing based on a multimodal large model according to claim 1 is characterized in that: Mapping the multimodal feature data to a unified dimensional space to calculate the association weights between each modality through a cross-modal attention mechanism and establish cross-modal bidirectional link data, specifically including: Projecting the text semantic features, table relationship graph, and image visual features in the multimodal feature data into a unified dimensional space through a learnable linear transformation matrix to generate a text projection vector, a table projection vector, and an image projection vector; Determining a first cosine similarity score between the text projection vector and the table projection vector to screen association pairs whose first cosine similarity score is higher than a first preset threshold, establishing a text-table explicit link, and determining a text-table explicit link weight; determining a second cosine similarity score between the text projection vector and the image projection vector, screening association pairs whose second cosine similarity score is higher than the first preset threshold, establishing a text-image explicit link, and determining a text-image explicit link weight; Determining an explicit link set by using the text table explicit link, the text table explicit link weight, the text image explicit link and the text image explicit link weight; Based on the numerical dependency relationship in the table relationship graph and the explicit link set, the table image implicit link and the table image implicit link weight are determined to determine the implicit link set.
5. The method for smart contract auditing based on a multimodal large model according to claim 4 is characterized in that: Determining implicit links of the table image and implicit link weights of the table image based on the numerical dependency relationship in the table relationship graph and the explicit link set, specifically including: Traversing the explicit link set, screening target explicit links whose target modality is a table modality, to obtain associated table cells and target explicit weights corresponding to the target explicit links; According to the numerical dependency relationship, tracing back at least one dependent cell of the associated table cell, recording the dependency level number of each dependent cell, and determining the implicit weight corresponding to each dependent cell based on the dependency level number and the target explicit weight; Obtaining cell layout coordinate data of each of the dependent cells and image layout coordinate data of multiple image areas in the target contract document, calculating an intersection-over-union (IoU) based on the cell layout coordinate data and the image layout coordinate data, and screening cell-image area pairs whose IoU is greater than a second preset threshold; According to the implicit weight of each dependent cell and the intersection-over-union ratio, the cross-modal implicit link weight corresponding to the cell-image area pair is determined, and based on the cross-modal implicit link weight, the table image implicit link is determined.
6. The method for smart contract auditing based on a multimodal large model according to claim 1, characterized in that: Based on the multimodal feature data and the cross-modal bidirectional link data, single-modal self-consistency verification and cross-modal contradiction detection are performed on the target contract document to generate audit information of the target contract document, specifically including: Based on the multimodal feature data, each modality is independently verified to determine the single-modal self-consistency verification data corresponding to each modality; performing cross-modal contradiction detection based on the cross-modal bidirectional link data to determine cross-modal contradictory data in the target contract document; Obtain a preset legal knowledge graph to perform risk assessment on the unimodal self-consistency verification data and the cross-modal contradiction data based on the legal knowledge graph, determine the comprehensive risk index of the target contract document, and determine the audit information of the target contract document through the comprehensive risk index.
7. The method for smart contract auditing based on a multimodal large model according to claim 6 is characterized in that: Performing cross-modal contradiction detection based on the cross-modal bidirectional link data to determine cross-modal contradictory data in the target contract document specifically includes: Obtaining an explicit link set and an implicit link set in the cross-modal bidirectional link data; Based on the text table explicit links in the explicit link set, calculating the absolute deviation between the text description value and the table storage value, and marking it as an explicit contradiction if it exceeds a dynamic threshold corresponding to the contract type of the target contract document; Based on the text-image explicit links in the explicit link set, verify whether the operation requirements described in the text are consistent with the content of the image element, and if not, mark it as an explicit contradiction; Traverse the implicit link path in the implicit link set and check whether there is a cross-modal logic contradiction in the path. If so, mark it as an implicit contradictory association path.
8. The method for smart contract auditing based on a multimodal large model according to claim 6 is characterized in that: Based on the legal knowledge graph, a risk assessment is performed on the single-modal self-consistency verification data and the cross-modal contradiction data to determine a comprehensive risk index of the target contract document, specifically including: Determining table calculation error information, signature exception information, and clause missing information in the single-modal self-consistency verification result, mapping the table calculation error information, the signature exception information, and the clause missing information to target rules in the legal knowledge graph, and determining a single-modal risk indicator; Obtaining explicit contradiction numerical deviations and implicit contradiction association paths in the cross-modal contradiction data, and calculating the deviation percentage corresponding to the explicit contradiction numerical deviations to determine an explicit contradiction risk indicator; Determining the weight mean and path length corresponding to the implicit contradiction association path to evaluate the path vulnerability of the implicit contradiction association path and determine the implicit contradiction risk index; A comprehensive risk index of the target contract document is determined by using the single-modal risk index, the explicit contradiction risk index, and the implicit contradiction risk index.
9. A smart contract audit device based on a multimodal large model, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Form information acquisition method and device and server
CN113011144A
Intelligent contract image recognition and contract element extraction method and device
CN114758341A
Contract risk detection method and device based on large language model
CN118484529A
Contract information extraction method and equipment based on multi-modal large language model
CN118734032A
Electronic contract auditing method and device, nonvolatile storage medium and electronic equipment
CN118918602A
Cited By
Multi-modal bidding document intention anchoring and rule pluggable auditing method and system
CN120766304A
Bidding field information extraction method and device
CN121074929A
Translation manuscript review-oriented multi-modal AI (Artificial Intelligence) cooperative processing system and method
CN121118920A
Power data dynamic verification method based on large model
CN121328527A
Intelligent document compliance auditing system and method based on multi-modal deep learning
CN121365360A